Skip to main content
Microbial Genomics logoLink to Microbial Genomics
. 2026 Jul 8;12(7):001785. doi: 10.1099/mgen.0.001785

Evolving strategies for virus discovery

Amanda Araujo Serrao de Andrade 1,2,†, Andrea Silverj 1,2,†, Theodore Josephs 1,2,†, Ann C Gregory 1,2,*
PMCID: PMC13344882  PMID: 42418234

Abstract

Viruses interact with all domains of life and play fundamental roles in shaping biological systems from individual hosts to global ecosystems. Yet their identification remains difficult due to a lack of a universal marker gene and the extensive diversity of viral genomes. Despite this, the speed of viral discovery is quickly increasing, driven by the growing number of virome studies, improved sequencing technologies and the decreased cost of sequencing. In this review, we examine the evolution of virus identification approaches from classical and molecular methods to contemporary genome-resolved and computational frameworks. By aggregating genome-resolved virome studies from 2010 to early 2026 that meet defined criteria (n=502), we synthesize the current landscape of virus identification methods, including similarity-based, sequence-based artificial intelligence (AI) and hybrid approaches. We also highlight the key limitations of the current methods, particularly biases in reference databases that contribute to persistent viral ‘dark matter’. Finally, we identify emerging opportunities for the field in structure-based and AI-driven approaches that extend detection beyond sequence similarity and outline how these integrative frameworks are poised to improve virus discovery across ecosystems.

Keywords: metagenomics, phage, viral ecology, virome, virus


Impact Statement.

Over the past 5 years, sequence-based artificial intelligence protein structure prediction has transformed the landscape of biology, opening access to an unprecedented diversity of protein folds. In viromics, we believe this will revolutionize how viruses are identified and classified, as viral protein structures make it possible to model distant homologous relationships that are only detectable at this level. We highlight the potential of these new methods compared to currently available approaches, for which we provide a comprehensive overview by listing and summarizing all major studies in the field from 2010 to the present. For most of the modern viromics era, analyses have relied solely on sequence data, limiting the detection of highly divergent viruses in metagenomic datasets. We anticipate that this emerging framework based on protein structure comparisons will reshape how viruses are identified and classified across ecosystems and provide a path towards a more complete view of the virosphere.

Data Availability

The scripts used to generate the plots and tables presented in the manuscript are available at https://github.com/IntegrativeViromicsLab/micro_gen_review.

Introduction

Viruses have historically been among the most difficult biological entities to identify and classify. While bacteria were first observed in 1676, viruses remained invisible to science for nearly two more centuries [1,3]. This delay reflects not only technological but also conceptual challenges, as viruses are too small to be visualized via light microscopy and they lack shared cellular features that underpin classical definitions of life. Consequently, the development of virus identification approaches has long lagged behind those used for bacteria and archaea.

Beyond their small size, one of the most difficult obstacles to virus discovery is the absence of a universal marker gene. In bacteria and archaea, the 16S rRNA gene provides a conserved genetic marker that enables broad detection, classification and comparative analysis across ecosystems [4]. No equivalent marker exists for viruses. Viral genomes comprise DNA or RNA, single- and double-stranded forms, segmented and non-segmented architectures and highly diverse gene repertoires arising from higher mutation rates, recombination and horizontal gene transfer [5,6]. This lack of universal conservation has made marker-based virus identification inherently biassed towards previously characterized lineages.

As a result, virus identification has progressed through a succession of different methodological frameworks, each shaped by the available technology and prior knowledge. Early experimental approaches inferred the presence of viruses indirectly, relying on observable effects such as disease transmission, host cell lysis or the passage of infectious agents through filters that excluded bacteria, rather than direct visualization of the viral particles [3,7, 8]. Later molecular methods focused on conserved genes within specific viral groups to support detection and classification [9,11]. More recently, high-throughput sequencing and metagenomics have transformed virus discovery by enabling the detection of large numbers of uncultivated viruses directly from the environment and host-associated samples [12,13]. Despite these advances, genomic-centric approaches remain limited by reference database coverage, assumptions of sequence similarity and the persistence of ‘viral dark matter’ [14,15].

In this review, we trace the evolution of virus identification methods from pre-genomic foundations to contemporary homology-based, artificial intelligence (AI)-driven and structure-informed approaches. By integrating historical context with current practices and emerging directions, we examine both the enduring challenges posed by viral diversity and the opportunities created by advances in computation, protein structure prediction and large-scale data integration.

Foundational approaches to virus identification

Before the advent of genomics, virus identification relied on indirect experimental evidence rather than direct observation of viral particles. In the late nineteenth century, Adolph Mayer’s work on tobacco mosaic disease showed that infectious material could pass through bacteria-retaining filters, indicating the existence of a different class of infectious agents distinct from bacteria [16,17]. Building upon Adolph Mayer’s work, Dmitri Ivanovsky (1892) and Martinus Beijerinck (1898) independently showed that filtered plant extracts remained infectious, establishing filterability as a defining property of viruses [2,3, 18]. While these early classical experiments could not visually confirm the presence of viral particles, they did initiate a methodological cascade that has continued to expand our toolkit for understanding viral diversity, many of which are still used today (Fig. 1a).

Fig. 1. Methodological advances and core marker genes in viral classification and phylogeny. (a) Timeline of major viral identification methodologies (1885–present). Methods are grouped by approach: classical (phenotypic assessment and visualization), molecular (sequence-based) and computational (genome-scale and AI). Solid bars represent periods when a methodology was the primary standard for viral discovery. Dash lines indicate reduced reliance, marking the historical point where newer technologies emerged and older methods transitioned to supplementary tools. (b) Core marker genes for viral classification and identification. Representations of major viral groups, including RNA viruses (Riboviria), ssRNA (Retroviridae), dsDNA (Adenoviridae, Caudoviricetes, Nucleocytoviricota and cyanophage) and ssDNA viruses (Microviridae and Cressdnaviricota). Highlighted are key structural proteins (e.g. Hexon and HK97-fold) and replication enzymes (e.g. RdRp, PolB and Rep). The cyanophage panel (bottom right) illustrates the viral hijacking of host metabolism (photosystem II: psbA and psbD) to support viral replication. Bracketed numbers indicate citations. Protein structures were sourced from AlphaFold [118,120, 155] (accessions: A0A481YZH0, A0A286Q6J9, A7IY91, A0A2U7NLV3, C6K7K0, P85987, Q71F16 and X2J3M8 [125]; accessions: Q91EK3 and P11819 [126]).

Timeline and diagrams showing viral identification methods from 1885 to present and marker genes across viral groups including Riboviria, Caudoviricetes, and Cyanophage, with key replication enzymes and structural proteins.

The discovery of bacteriophages by Frederick Twort (1915) and Félix d’Hérelle (1917) introduced plaque assays as a quantitative and observational method for detecting viruses through zones of host lysis on bacterial lawns [7,8, 19]. This approach transformed the field of virology by enabling viruses to be isolated, quantified and experimentally manipulated. However, plaque-based detection inherently favours viruses capable of producing visible lysis under laboratory conditions. Thus, early virus discovery was strongly biassed towards lytic viruses infecting readily cultivable hosts, leaving large portions of viral diversity undiscovered.

The 1930s marked a shift inferring viral presence through host lysis to observations of virus structure with the arrival of electron microscopy. This method provided the first direct visualization of viral particles and enabled viruses to be classified based on morphological features, such as capsid organization, tail structures and the presence or absence of envelopes [20,22]. Morphological classification represented a major conceptual advance, allowing viruses to be grouped into structural families and forming the foundation for early viral taxonomy. Yet this framework, like plaque assays, was constrained by both technological and biological limitations, requiring cultivable viruses that could be propagated at high titres.

Advances in molecular biology further expanded the tools available for virus identification. Coinciding with these early molecular advancements, the International Committee on Taxonomy of Viruses (ICTV) published its first report in 1971, establishing the first initial viral species list. Simultaneously, the development of Sanger sequencing in the 1970s enabled the first direct sequencing of viral genes and genomes, providing a molecular framework to study viral diversity and evolution. However, early sequencing efforts remained relatively low throughput and typically required viral isolation prior to sequencing. The broader molecular revolution of the late twentieth century introduced new strategies for virus identification that no longer relied on particle visualization or host cultivation. In particular, approaches like PCR and Reverse Transcription-Polymerase Chain Reaction (RT-PCR) enabled sensitive and rapid detection of viral nucleic acids directly from biological samples, greatly expanding surveillance capabilities of known viruses without the need for viral cultivation [9]. However, these methods still depended on prior knowledge of viral sequences to design primers, restricting the detection to viruses closely related to previously characterized viruses.

While the use of PCR and RT-PCR expanded molecular surveillance for known viruses, it left the broader problem of viral diversity unaddressed. In order to partially overcome this limitation, researchers in the 1980s began targeting conserved viral genes within key viral groups as phylogenetic markers, which enabled the discovery of viruses carrying homologues of genes found in known viral lineages (see Fig. 1b and Table S1, available in the online Supplementary Material [10,23,38]). For example, in bacteriophages, conserved structural and replication-associated genes such as gp23, encoding the major capsid protein of T4-like phages [39,40], and the terminase large subunit (TerL) conserved broadly in tailed dsDNA phages, are used as a phylogenetic marker to infer similarity. Such markers enabled the classification of viruses into viral lineages. Similarly, in RNA viruses, RNA-dependent RNA polymerase (RdRp) was initially used as a genetic marker [41,42]. Despite the lack of a viral universal marker gene, these marker-gene approaches enabled broader surveys of viral diversity but were limited by their dependence on previously known conserved genes, restricting the detection of novel divergent viruses. This absence reflects both the unique viral diversity and evolutionary history but also highlights the limitations of marker-based classification.

Historical approaches have established the conceptual and methodological foundations of virus identification and classification. However, their limitations in requiring cultivation, conserved markers across species and limited detection on low-abundance viruses have motivated the development of current practices.

Current and integrative approaches to virus identification

Current virus identification approaches have been fuelled by the transition from targeted, marker gene-based and low-throughput culturing and sequencing strategies to untargeted metagenomic and metatranscriptomic approaches, which capture DNA and RNA, respectively, capable of surveying entire viral communities. Advances in next-generation sequencing have made it possible to sequence all the nucleic acids within a community without the need for targeted amplification, freeing virus discovery from its historical dependence on known viral groups or predefined genetic markers. Instead, current virus discovery involves the recovery of viral sequences directly from complex mixtures of DNA and/or RNA derived from diverse biological communities. This conceptual and technical shift has led to the exponential rise in the discovery of novel viral species over the last decade. It has also introduced a computational hurdle. Specifically, using an untargeted approach makes extracting and identifying viral genomes from complex datasets difficult. To understand how the field of viromics is overcoming this limitation, we mapped the current state of virus discovery in the viral metagenomic field.

While the first virome was published in 2002 [13], the field did not expand substantially until the 2010s, when genome- and population-level analyses began to take off. For this review, we surveyed Google Scholar using the search terms (‘viral population’ OR ‘vOTU’) AND (‘viral metagenome’ OR ‘virome’) between January 2010 and March 2026, yielding an initial 2,083 results (see full list in Table S2). We excluded 1,580 studies comprising theses/dissertations, reviews, book chapters or those reporting a low number of viral operational taxonomic units (vOTUs). To ensure robust comparative analysis, the dataset was further restricted to studies reporting more than 1,000 de-replicated vOTUs, as well as targeted inclusions of less than <1,000 vOTUs from unique or underrepresented environments. This resulted in a total of 502 studies (summarized in Table 1) with the detailed metadata extracted from these final studies, focusing on assemblers, virus identification tools and vOTU counts, grouped into 3-year periods to highlight temporal trends.

Table 1. Summary of studies included in the database across time periods (2010–2026).

The table presents the number of studies and predominant environments: host, marine, freshwater, soil, extreme and other environments (agricultural slurry, chaparral wildfire, plastisphere, food and public transit air). This table also shows the commonly used assemblers and virus identification tools for each period, along with the total number of de-replicated vOTUs and their mean±sd per study.

Publishing year No. of studies Environment Main assembler Main virus identification tool Total no. of vOTUs Mean and sd
2010–2013 13 Host (6), marine (5), freshwater (1), extreme (1) CLC Genomics Workbench, Newbler, Velvet blast 5,273 1,054.6±2,126.2
2014–2017 51 Host (34), marine (8), other (5), freshwater (2), extreme (2) MIRA, Velvet, SPAdes, IDBA-UD blast, HMMER, VirSorter 68,681 4,040.1±7,295.0
2018–2021 119 Host (56), other (23), marine (18), soil (16), freshwater (4), extreme (2) MEGAHIT, SPAdes, Velvet blast, VirSorter, VIBRANT, VirFinder, DeepVirFinder 31,843,750 403,085.4±33,288,800.4
2022–2025 303 Host (133), other (51), soil (45), marine (45), freshwater (20), extreme (9) MEGAHIT, SPAdes, Trinity, Flye VirSorter2, VIBRANT, blast, geNomad, DeepVirFinder 12,934,157 48,624.7±224,556.30.3
2026* 16 Host (9), marine (4), other (1), extreme (1), freshwater (1) MEGAHIT, SPAdes, Flye blast, PhaGCN, geNomad, VirSorter2, VIBRANT 744,132 49,608.8±97,671.80.8

*As of March 2026.

Early studies heavily focused on exploring aquatic environments [43,48] and simple clinical snapshots of human faecal or oral samples [49,51]. This initial emphasis on aquatic biomes was driven by early estimates, revealing high viral abundance [52]. Given that the marine environment does contain the overwhelming majority of the Earth’s viruses, researchers have focused their efforts on the oceans as a massive reservoir for discovering viral populations. This aquatic focus culminated in landmark global surveys including the J. Craig Venter Institute Global Ocean Survey [53], Malaspina [54] and the Tara Oceans global dataset [55,57], which the latter two helped establish foundational baselines for marine viral diversity, resulting in ~579,904 vOTUs in the most recent expansion of the Global Ocean Virome 2.0 dataset [58] (Fig. 2a).

Fig. 2. Temporal dynamics of viral detection studies across environments (2000–2026). (a) Density plot illustrating the temporal distribution of studies that sequenced within the six environment groups: extreme, freshwater, marine, soil, other and host (n=502 studies). (b) Stacked area plot showing diversification within the host environment group (2011–2026), with subcategories representing specific host types (n=238 studies). Area proportional to the number of studies published per year.

Stacked area charts show viral detection studies by environment and host type from 2011 to 2026, with counts peaking near 2022. Host-associated studies dominate, with human hosts contributing the largest share, while marine and soil remain secondary.

Between 2014 and 2021, there were 170 out of the 502 total studies. Of those 170, 53.9% (n=90) was linked to human health, wildlife and livestock. This reflects a shift towards host-associated microbiomes and viromes (Fig. 2a, b). Aquatic environments (marine and freshwater) remained a core location for major viral metagenomic studies with 18.8% (n=32) being related to this area. This period of time also marked the emergence of research targeting complex soils with 9.4% (n=16) being dedicated to this area.

The current landscape (2022–2026) has seen an explosive surge in data generation for viral metagenomics, with our literature survey capturing 319 out of the 502 metagenomic studies and databases published during this short timeframe (see Table 1). This era is characterized by the consolidation of sequences into massive databases such as Gut Virome Database [59], Gut Phage Database [60], Metagenomic Gut Virus Catalogue [61], ViromeDB [62], Global Soil Virus (GSV) Atlas [63], Groundwater Virome Catalogue (GWVC) [64], Global Deep-sea Sediment RNA Virome 2.0 (GDSR2.0) [65], Chinese Gut Viral Catalogue (cnGVC) [66] and MetaVirus Resource (MetaVir) [67]. The construction of these large viral databases highlights the rapid increase in efficiency of viral detection from 2000 to 2026 (Fig. 3a). This also coincided with a targeted push into extreme and atypical environments, extending viral sampling into habitats as diverse as hydrothermal vents, wildfires, the plastisphere and the cheese virome (Fig. 2a).

Fig. 3. vOTU detection, dataset characteristics and tool usage in viral identification studies. (a) Heatmap showing the number of de-replicated vOTUs (log10 scale) reported per year (2012–2026), stratified by nucleic acid type (DNA and RNA) and environment (extreme, marine, freshwater, soil, other and host). We normalized the vOTU count based on the number of studies per environment per year to show whether viral discovery has become more efficient over time. (b) Distribution of vOTU counts (log10 scale) across study types (virome-enriched, bulk sequencing and both strategies), comparing primary studies and meta-analyses. Points represent individual studies; violin plots indicate data distribution. Differences among sequencing strategies are shown as descriptive patterns and were not subjected to formal statistical testing. (c) Density distribution of publication years weighted by the number of de-replicated vOTUs reported per study, stratified by viral identification methodology. Sequence-based AI methods correspond to machine learning or deep learning approaches, similarity-based methods correspond to homology or reference-dependent approaches and hybrid methods include studies combining tools from multiple methodological categories. Density values were calculated using raw vOTU counts without log transformation. This panel illustrates temporal shifts in viral identification strategies and their association with increasing vOTU discovery over time. (d) Most frequently cited viral detection tools across studies, shown on a log10 scale. Points represent citation frequency, coloured by methodological category (sequence-based AI, similarity-based and hybrid). Citation counts were retrieved from Google Scholar in April 2026, based on the original publication describing each tool. The tools marked with an asterisk (*) are not just for viral detection.

Heatmap, violin plots, density curves, and dot plot show vOTU counts by environment and year, sequencing strategy, and identification method over time, alongside citation counts for tools like BLAST and HMMER.

As viral ecology has expanded into more unique and diverse environments, the associated bioinformatic pipelines used for virus identification have had to account for several biases introduced during wet lab experiments. One of the most significant biases in early viral metagenomics data arises from the need to amplify low quantities of DNA. In environments where viral biomass is extremely low, researchers frequently rely on multiple displacement amplification (MDA) to generate enough material for sequencing. However, the Phi29 polymerase used in the MDA process exhibits a significant amplification bias towards small ssDNA [68]. Therefore, in studies within our dataset, where MDA was utilized, the abundance of ssDNA viral families will most likely be higher. Interestingly, our analysis reveals that MDA usage was drastically higher in the field’s early years, with 46% of the 13 studies utilizing MDA between 2010 and 2013, compared to only 5% of the 319 studies between 2022 and 2026. This steep decline coincides with the advent of better library preparation, extraction and sequencing methods available as well as rigorous method testing [69].

Traditional ‘bulk’ metagenome sequencing, which captures the entire microbial community, enables the identification of free viruses but also integrated prophages and their bacterial hosts. It also captures signals of actively infecting viruses, which often comprise the majority of recovered viral sequences [70]. The primary limitation of this approach is the overwhelming ‘noise’; because host cellular genomes are larger than viral genomes, the viral signal is often drowned out unless the sequencing depth is exceptionally high.

To circumvent this, researchers often aim to enrich the virome by isolating viral-like particles. This is typically achieved through size fractionation using a 0.22 µm filter or chemical flocculation–precipitation, followed by density gradient separation to isolate virions before extraction [71]. Alternatively, the extraction process itself can enrich viruses by utilizing commercial kits designed with specialized lysis buffers and binding matrices optimized specifically for viral nucleic acids [72]. Our data analysis underscores the advantage of a dual strategy, with those studies utilizing a combination of both methods recovering a higher mean number of vOTUs compared to those relying solely on bulk metagenome or virome approaches (Fig. 3b). Overall, the extraction and sequencing strategies utilized across these studies are only a small part of the current virus discovery pipeline. Genomic analysis alone is not suitable to answer the whole spectrum of questions that researchers come up with and might not be ideal in every situation [73]. Virus isolation, culture-based assays and serology still play a very important role in epidemiological studies and defining host range [74,75]; electron microscopy is still widely used to understand how viruses are organized at the molecular level and how they behave in different tissues and conditions, including infections [76,78]; experimental validation is still the most trustworthy line of evidence, which can confirm or disconfirm bioinformatic predictions, which could return false positives due to the presence of fragmented or residual viral nucleic acids [79]. Because of that, all these methods are still actively used in viral identification and characterization, and it is often necessary to integrate more than one of them together to complement the evidence they offer and put together an explanation that is biologically relevant [80].

Today, the success of large-scale viromics studies relies on computational power driven by assembly and virus identification tools. Within this framework, virus identification is now largely driven by assembled sequencing data. Raw short reads are computationally assembled into longer contigs, after which viral sequences must be distinguished against a background dominated by cellular and other non-viral genetic material. As a result, virus identification has become a computational challenge, one that demands the integration of multiple lines of evidence to accurately classify and interpret the recovered genomic fragments. The pairing of specific assemblers with virus identification tools reflects methodological compatibility as well as the field’s evolving preferences, often shaped by dataset characteristics, sequencing technologies and the desired resolution of viral recovery.

Analysis of assembler usage across 502 studies (Tables 1 and S2) revealed clear trends in their usage over time and impact on vOTU detection. From 2010 to 2017, most studies reference assemblers like Newbler [81], Velvet [82] and CLC Genomics Workbench (QIAGEN digitalinsights.qiagen.com), which were originally designed for isolated genomes and low-complexity datasets. As metagenomics grew more widespread, the field shifted towards tools better suited for complex microbial communities and highly variable sequencing depths, such as SPAdes [83] and Megahit [84]. This transition matches with a noticeable increase in the number and continuity of recovered viral contigs, ultimately enhancing vOTU detection over time (Fig. 1b). It is important to mention the transition from general assemblers to metagenomics-specific assemblers, including MetaSPAdes [85], MetaViralSPAdes [86], MetaVelvet [87] and MetaFlye for long-read sequencing [88]. These specialized implementations incorporate algorithmic adjustments tailored to metagenomic complexity. The choice of assembly tool directly impacts vOTU detection, since different assemblers yield substantially different numbers of viral contigs and affect both the length and completeness of recovered viral genomes, as well as the inferred viral community composition. Therefore, assembler choice is not a neutral technical detail but a key methodological decision that shapes which vOTUs are detected, how well they are resolved and ultimately how viral diversity and community structure are interpreted across studies [89].

Following assembly, the subsequent step in vOTU detection involves virus identification tools, which can be divided into four main groups based on their methodological approach: (i) similarity-based, (ii) sequence-based AI, (iii) structure-aware AI and (iv) hybrid that combine two or more of these approaches, with structure-aware AI methods representing a more recent development with only one reported tool (Tables 2 and S3). Similarity-based approaches rely on reference databases to detect and classify viral sequences through both sequence-level comparisons and gene content–based profiles. In contrast, sequence-based AI approaches extract numerical features from sequences and apply machine or deep learning (DL) models to recognize viral genome patterns, including those from novel or highly divergent viruses lacking close homologues. Structure-aware AI methods predict protein structures to inform virus identification and functional annotation. Hybrid approaches integrate multiple elements, often combining sequence similarity, gene content, numeric features and/or structural information to improve detection sensitivity and classification accuracy.

Table 2. Summary of virus identification tools grouped by methodological approach.

The table lists the number of tools in each category (similarity-based, sequence-based AI, structure-aware AI and hybrid) and the viral targets they are designed to detect (all viruses, DNA and RNA viruses). For each group, the table indicates whether the tool provides taxonomic classification and/or functional annotation of detected viral sequences.

Approach No. of tools Viral target Taxonomy
(yes/no)
Functional annotation
(yes/no)
All viruses DNA
virus
RNA virus
Similarity-based 36 23 8 5 Yes: 29
No: 7
Yes: 23
No: 13
Sequence-based AI 28 14 12 2 Yes: 6
No: 22
Yes: 5
No: 23
Structure-aware AI 1 0 0 1 Yes: 0
No: 1
Yes: 1
No: 0
Hybrid 31 14 13 4 Yes: 18
No: 13
Yes: 17
No: 14

The most commonly used virus identification tools described in Table 1 are blast [90], VirSorter2 [91], VIBRANT [92], geNomad [93] and DeepVirFinder [94]. Based on the aforementioned subdivision by methodological approach, blast is classified as a similarity-based tool, DeepVirFinder as a sequence-based AI tool and VirSorter2, VIBRANT and geNomad as hybrid tools (Table S3). Throughout this review, we examine in detail the methods employed by these prominent tools, highlighting their respective strengths and limitations in accurately characterizing viral sequences from complex metagenomic datasets.

To systematically evaluate these approaches, we compiled a list of 95 publicly available virus identification tools capable of detecting viral sequences from unclassified metagenomic data (Tables 2 and S3). For each tool, we provide information on its primary methodological approach as well as the viral groups it targets: (i) all viruses, (ii) DNA viruses or (iii) RNA viruses. This classification allowed us to cross-reference tool usage across the literature and evaluate how different approaches contribute to vOTU detection over the years. In this review, we discuss each methodological approach, highlighting their principles, strengths, limitations and impact on viral detection in metagenomic studies.

Sequence- and gene content–based virus identification

All similarity-based approaches compare query nucleotide or protein sequences directly to reference databases using alignment methods that match positions across sequences [95]. A sequence alignment represents a hypothesis of shared evolutionary history, assuming that nucleotides or amino acids at aligned positions have diverged from common ancestors [95,96]. These methods can be subdivided into sequence-based and gene content–based virus identification tools. Sequence-based methods directly analyse the nucleotide or amino acid composition of raw contigs using local alignment tools like blast [90], MMseqs [97] and DIAMOND [98,99], which perform exact or near-exact matches against curated repositories such as those maintained by the National Center for Biotechnology Information (NCBI) and the European Bioinformatics Institute (EBI). Local alignment tools aim at identifying conserved or homologous regions, enabling similarity analysis, functional inference and taxonomic assignment. In contrast, gene content–based methods, such as HMMER [96], employ local alignments to detect similarities between query amino acid sequences and hidden Markov model (HMM) profiles. These profiles are built from global alignments of proteins deposited in large HMM databases like Pfam (now integrated into Interpro) [100] and/or viral specific HMM databases such as vFams [101], eFams [102], VOGdb [103], PHROG [104] and The RNA Viruses in Metatranscriptomes (RVMT) [105], offering high functional specificity or longer contigs but potentially missing viruses lacking close homologues and annotated ORFs [101].

The survey of 502 studies reveals preferences for certain bioinformatic tools (Table 1), particularly similarity-based approaches such as blast and HMMER, hybrid approaches like geNomad and VirSorter2 and sequence-based AI approaches such as DeepVirFinder. Fig. 3(c) illustrates this trend by showing the density of de-replicated vOTUs identified from 2010 to March 2026 using similarity-based, sequence-based AI and hybrid tools. This figure illustrates temporal shifts in viral identification methodologies and their association with increases in vOTU discovery over time. Similarity-based tools account for the majority of vOTUs across this period, with a marked increase corresponding to the higher number of studies conducted between 2022 and 2025 (n=303). In the same period, vOTUs identified by sequence-based AI and hybrid tools show a noticeable rise, with a prominent peak of vOTUs identified by sequence-based AI tools around 2023. This shift likely reflects their ability to overcome limitations of similarity-based approaches and to integrate strengths from multiple methodologies.

Although similarity-based approaches have identified the majority of vOTUs across the surveyed studies and garnered the most citations in Fig. 3(d), their effectiveness remains intrinsically tied to the genetic diversity and quality of sequences deposited in reference databases as well as the diversity present within the samples themselves. These databases are subject to biases that affect the representation of viral diversity. It is estimated that only a small fraction of the viral diversity found in prokaryotic hosts is represented in databases, due to limitations associated with the isolation and assembly of viral genomes [15]. Consequently, many vOTUs are represented by few genomes, which are often incomplete or low quality. In the absence of close homologues, a large portion of viral genomic sequences is labelled as unknown, also referred to as ‘viral dark matter’. These unknown sequences are frequently discarded and do not undergo subsequent bioinformatic analyses following alignment against a reference database [106]. Beyond database limitations, the similarity-based approaches assume collinearity (linearly conserved nucleotides), which fails for viruses with high mutations, recombination, horizontal gene transfer and gene duplications [107,108].

Sequence-based AI virus identification

To assess the limitations inherent in similarity-based viral detection, recent studies have increasingly used sequence-based AI approaches, either standalone or hybrid with similarity-based methods. Sequence-based AI approaches leverage machine learning (ML) and DL techniques to learn viral genome patterns from reference databases and identify viruses in unclassified contigs as short as 500 bp [108,109]. In contrast to similarity-based approaches, sequence-based AI approaches rely on alignment-free numerical features derived from sequences. By assessing sequence features without the need for alignment to reference genomes, these tools are able to detect conserved genomic signatures in divergent or novel viruses that are often overlooked by similarity-based approaches. This provides robustness to fragmented contigs, low-complexity regions and recombination events, ultimately enhancing sensitivity and uncovering previously uncharacterized viral diversity [94,108, 110].

Sequence-based AI virus identification tools have been published since 2017 [109] with publication rates accelerating between 2019 and 2022 [94,110,112] according to our survey (Table S3). This surge corresponds to the rise in metagenomic studies and the development of more complex algorithms over time [108]. The earliest sequence-based AI tools relied on relatively simple ML models, such as random forest and logistic regression that were trained on computationally simple alignment-free sequence-derived features including k-mer frequencies, codon usage bias and GC content. These features are effective for viral detection since viral genomes often exhibit distinct compositional patterns compared to their host or other microbial sequences, reflecting varying nucleotide usage, coding strategies and genome organization [92,109]. Early implementations primarily focused on simple classification tasks, such as binary classification of viral and non-viral contigs. The development of these tools represented a paradigm shift in virus identification; these models laid the foundation for more sophisticated algorithms, enabling the application of advanced methods from other fields to the challenges of metagenomic virus discovery.

The most highly cited sequence-based AI tools on Google Scholar as of March 2026 are VirFinder [109], DeepVirFinder [94], PPR-Meta [112], ViraMiner [110] and PhaGCN [111], while the leading hybrid tools are VIBRANT [92], VirSorter2 [91] and geNomad [93] (Fig. 3d). This pattern aligns with the compilation of studies outlined in Table 1. In most of these studies, these sequence-based AI tools were used either exclusively or in combination with similarity-based tools within multi-tool pipelines. To further illustrate the diversity of sequence-based AI virus identification approaches, these tools can be divided into ML and DL techniques, which differ in the way they learn (numeric features) and interpret viral sequence features (models used).

The most cited ML virus identification tools VirFinder [109], VIBRANT [92] and VirSorter2 [91] present notable differences in both feature selection and model architecture that influence their performance and overall applicability. For instance, VirFinder is an interpretable and fast tool that uses k-mer-based features and a logistic regression model to classify short contigs as viral or non-viral. However, this tool tends to underperform compared to both VirSorter2 and VIBRANT by producing more false positives, which makes its predictions less reliable than those of VirSorter2 and VIBRANT [113]. In contrast, hybrid tools such as VIBRANT and VirSorter2 combine similarity-based information with ML classifications to improve the detection of fragmented or divergent viral sequences. Specifically, VIBRANT employs a neural network with a ‘v-score’ metric to estimate the probability that a contig contains virus-like protein annotation signatures, thereby enabling viral detection [92]. In contrast, VirSorter2 applies multiclassifier random forest models based on hallmark protein annotations [91].

VirSorter2 features multi-classifier and broad-spectrum viral detection, moving away from a single cohesive model towards a modular framework [91]. Contigs first undergo functional annotation and HMM-based hallmark gene detection, after which VirSorter2 feeds a fixed set of 27 genomic features extracted from the annotated contigs directly into its five group-specific random forest classifiers. This architecture uses five specialized classifiers to detect diverse viral groups: dsDNA phages, ssDNA viruses, RNA viruses, large Nucleocytoviricota and underrepresented groups. Its modularity enables updating individual classifiers or adding new models to classify viral groups as they are characterized without overhauling the entire algorithm [91]. Benchmarking has shown that VirSorter2 [91] outperformed VIBRANT on seawater and gut biomes, while VIBRANT outperformed VirSorter2 on soil biome [113,114]. The main limitations of VirSorter2 are related to reduced sensitivity on short contigs (<3 kb) and non-Caudovirales due to limited hallmark genes [91,113, 114]. Additionally, according to recent benchmarking studies, VirSorter2 presents high false positives on eukaryotic sequences, heavy computational demands and decreased accuracy. This further highlights the need for thorough output analysis and potential combination with other tools to reduce false positives. Notably, of the 107 studies that used VirSorter2 listed in Table 1, 81.3% (n=87) combined VirSorter2 with other virus identification tools, including similarity-based approaches such as HMMER and blast, sequence-based AI approaches such as DeepVirFinder and PHAGCN, and hybrid approaches such as VIBRANT and geNomad.

DL tools are able to complement ML tools by improving the detection of novel or highly divergent viruses that ML tools might miss [94]. ML classifiers, like the random forest in VirSorter2, depend on predefined features, while DL classifiers automatically learn features directly from the input contigs. ML excels with viruses similar to its training data but struggles with distant viruses, novel patterns or subtle signals beyond those features. In contrast, DL uncovers hierarchical, non-linear patterns, enabling stronger generalization to unknown viruses without manual feature selection [94,108, 110]

Among the most cited DL tools (Fig. 3d), DeepVirFinder uses convolutional neural networks (CNNs), which automatically learn patterns in short DNA subsequences by scanning sequences with small filters that capture local motifs, analogous to how the visual system detects edges and shapes in images [94]. These CNNs are trained on k-mer-encoded sequences to distinguish viral from non-viral contigs, demonstrating higher sensitivity than traditional ML tools at shorter contig lengths [108]. ViraMiner extends this approach with dual-branch CNNs that integrate nucleotide input and k-mers, outperforming DeepVirFinder on human microbiomes [110]. PhaGCN employs a graph convolutional network framework that leverages sequence context and taxonomic relationships for superior taxonomic assignment over conventional classifiers [111].

In contrast with these tools, geNomad [93], one of the most robust virus identification tools, identifies viruses and plasmids through a hybrid pipeline combining a deep neural network, which learns complex patterns in raw nucleotide sequences across multiple processing layers (sequence branch) with marker gene-based classification (marker branch) being considered in this review as a hybrid tool. Specifically, geNomad markers were defined by building a comprehensive database of protein profiles from viruses and plasmids, with alignments retrieved from verified and highly curated sources, including VOGdb [103], PHROG [104] and RVMT [105], to identify the profiles that are informative for sequence classification. These markers were then assigned to ICTV taxa using alignment and a majority vote function. This approach allows the programme to provide a reliable and relatively robust taxonomic classification (which can go up to the species level in the latest version of the programme) for the putative viral sequences identified. Other than markers and sequence-based approaches, geNomad is also equipped with a DL module consisting of a gene-based and a sequence-based classifiers, of which the respective outputs are aggregated by a neural network which is able to weight the contribution of the models, returning a final score.

Despite advances in genome-based virus identification, all current approaches remain constrained by their reliance on sequence similarity and protein family homology. Even hybrid frameworks that integrate gene content, sequence composition and ML continue to depend, either directly or indirectly, on patterns derived from known viral genomes. As a result, highly divergent viruses that lack detectable sequence or protein cluster homology remain difficult to identify, contributing to the persistence of viral ‘dark matter’ [15,106, 108].

The future of virus identification

One emerging direction to address these limitations is the incorporation of protein structure as a complementary layer of information for virus discovery. Unlike primary sequence, protein structures are often conserved across deep evolutionary timescale even among proteins that share little to no detectable sequence similarity [115] (Fig. 4a). This property is particularly relevant for viruses, where structural features, such as capsid and nucleocapsid folds, can remain conserved despite extensive sequence divergence [116,117].

Fig. 4. Virus identification is enhanced through protein structural comparison. (a) The use of structural databases enables the identification of viruses at lower sequence identity thresholds, increasing the likelihood of detecting viral sequences across greater evolutionary distances. (b) Plot showing the increase in the number of publicly available viral protein structures following the release of structural databases based on AlphaFold predictions. (c) Different representations of predicted ORFs, ranging from the nucleotide sequence to translated protein sequence, HMM profile and predicted protein structure, capture progressively deeper levels of virus signal, each suited to distinct scales of virus identification and phylogenetic reconstruction (adapted from [139].

Diagram and line graph showing structural databases detect viral distant relatives beyond 30% SeqID, while sequence methods detect only above 70% SeqID. Foldseek extends phylogenetic signal furthest along evolutionary divergence.

Recent advances in AI-driven protein structure prediction, most notably AlphaFold [118], combined with decreasing computational costs, have made it feasible to incorporate structure-based information into large-scale virus discovery workflows. In parallel, rapid structural comparison tools such as DALI [119], Foldseek [120] and Reseek [121] enable efficient searches across vast protein structure space. Together, these developments suggest that virus identification can extend beyond sequence similarity into structural homology, providing a means to detect viruses that are invisible to traditional sequence-based approaches.

Early applications of this idea have shown that viral proteins lacking detectable sequence homology can still be annotated through structural similarity searches, recovering functional and evolutionary signals that are otherwise missed [122]. Pipelines such as Phold have formalized this concept by incorporating structure-aware homology searches into phage genome annotation [123]

In parallel, structural information is increasingly being integrated directly into AI models for virus identification. For RNA viruses in particular, the conserved RdRp has served as a key entry point for detecting highly divergent lineages. While RdRp has long been used as a phylogenetic marker, traditional approaches based on sequence similarity alone can fail when divergence is high. Newer DL models address this limitation by incorporating structural representations alongside sequence features, enabling the detection of RdRp proteins that fall below conventional similarity thresholds. Recent large-scale applications of these approaches, such as the LucaProt method, a DL algorithm which integrates sequence and structural features of a large and highly curated set of viral RdRp proteins (n=5,979), as well as a negative one of non-viral proteins (n=229,434), have substantially expanded the known RNA virosphere, identifying hundreds of thousands of putative viral sequences and uncovering previously unrecognized diversity across global ecosystems [124].

Taken together, these developments point to a broader conceptual shift in virus identification. Rather than relying on a single source of signal, future approaches are likely to integrate multiple layers of information, including sequence, structure and learnt representations derived from deep neural networks. Structure-aware models are particularly promising for resolving viral dark matter, as they capture conserved functional and evolutionary relationships that are not apparent at the sequence level.

The recent development of large-scale viral structural databases now supports this shift. These databases include the Big Fantastic Virus Database (BFVD) [125], Viro3D [126], the Nomburg’s virome database at ModelArchive [127], Viral AlphaFold Database (VAD) [128] and Meta-virus resource (MetaVR) [129]. Taken together, these resources offer access to 1,279,884 folded viral protein sequences (Fig. 4b). This is a staggering increase, considering that the total number of protein structures available before AlphaFold came out was ~180,000 [130], with less than 10% (~18,000) being viral proteins [126,131] (Fig. 4b).

Most of these databases use AlphaFold2 as a predictive model, which compares well with its newer versions (e.g. AlphaFold3) [132], while offering greater reliability and stability in terms of predictions [133]. This model is currently required by the AlphaFold DB community for uploading contributions to the database (European Bioinformatics Institute 2026). Even though the AlphaFold (either 2 or 3) is currently the most used framework for compiling databases of predicted viral protein structures, other methods showing comparable results are available, such as RoseTTAFold [134], ESMFold [135], Boltz-2 [136], Chai-1 (The Chai discovery team 2024), HelixFold3 [137], Protenix (The ByteDance AML AI4Science Team 2025), OpenFold3 (The OpenFold3 Team 2025) and IntelliFold-2 [138]. This scenario is changing rapidly, and comprehensive benchmarks for these tools on viral proteins are still lacking, so the preference for AlphaFold may change in the future.

This abundance of methods and databases is catalysing a new era of virus identification and classification. New structural phylogenetic methods are emerging [139] to capture the phylogenetic signal that lies in protein structures [140], which will likely have an impact on the way we identify and classify viruses, shedding light on deeper evolutionary relationships that cannot be accessed by relying on nucleotide or amino acid sequence alone, as well as on HMMs (Fig. 4c). A more robust and deeply resolved phylogenetic framework grounded in structural information could directly help define reliable markers for virus identification. Unlike sequence-based markers, which are prone to saturation and homoplasy over long evolutionary distances [141], structurally defined markers would retain phylogenetic informativeness even across highly divergent lineages [142], making them particularly suitable for the identification of viruses that fall outside the known sequence space.

One of the most promising approaches in this field consists of making protein structures easily alignable and therefore usable for efficiently getting matrices that can be employed for inferring phylogenies, a capacity with direct implications for virus identification, since viruses that have diverged beyond the point of detectable sequence similarity may still share conserved structural folds. This is done by using structural alphabets, which consist of converting the 3D structure of a protein into a sequence of characters that represent a simplified version of the original protein structure [143]. Foldseek [120] implements one of the most popular structural alphabets so far, using a set of 20 characters called three-dimensional interaction characters (3Di characters) to represent the tertiary interactions of each amino acid with its spatially closest residue in the space. Another method based on a similar approach is Reseek, with the difference of using a much larger alphabet (a ‘mega-alphabet’ of ∼1011 letters) for representing the structure as a sequence, which gives to the method the capacity of capturing more subtle relationships between structures, improving sensitivity to remote homologs [121], a key property when trying to assign taxonomic identity to highly divergent or novel viruses. Other alphabets for the detection of remote homologues, such as TEA [144], are under development and have yet to be tested extensively.

On the other hand, Foldseek’s 3Di characters have already been used to develop methods to efficiently get multiple sequence alignments of protein structures, such as FoldMason [145], as well as to identify structural core genes, like Unicore [146], which uses the ProstT5 [147] protein language model to directly obtain sequences of 3Di characters from amino acid sequences, facilitating and speeding up large-scale phylogenetic analysis with protein structural information. The ability to define and align structural core genes is particularly valuable for virus identification, as it enables the construction of robust reference phylogenies against which unclassified viral sequences, including those recovered from metagenomic datasets, can be placed, even when only predicted structural information is available.

Another method which builds on the same 3Di alphabet is Foldtree [142], which uses matrices obtained from comparing sequences of 3Di characters to reconstruct distance-based trees. These characters can be used alone or in combination with amino acids under neighbour joining or maximum likelihood frameworks [142,148], and the best combination for different cases, as well as the utility and performance of the 3Di approach itself [149], is still under debate [139,148]. Furthermore, combining sequence and structural information has also been shown as a possible way to improve the reliability of support values, such as the multistrap approach for bootstrapping [150], an important consideration when attempting to confidently assign novel viruses to known clades. Despite these advances, substitution matrices and evolutionary models specifically designed for 3Di characters remain underdeveloped, with only a limited number currently available [120,151], which may limit the accuracy of phylogenetic placements for highly divergent viral lineages.

Building on these methodological developments, recent efforts have also extended beyond distance-based and maximum likelihood approaches to incorporate structural information within a Bayesian framework, with new packages like FoldBeast already available [152]. For virus identification, Bayesian methods, which combine a prior probability with a likelihood (derived from data and evolutionary models) to compute the posterior probability of hypotheses [153], are especially relevant as they allow the integration of prior knowledge about viral diversity and can provide probabilistic estimates of clade membership for uncharacterized viruses. However, the application of clock models to structural data is still in its early stages, and further studies are required to determine the extent to which structural information captures a reliable temporal signal [139,142], a question that bears directly on whether structural phylogenetics can eventually support not just classification but also the inference of viral emergence timescales and evolutionary origins.

An example of the power of these new structural approaches in virology is well represented by the recent study by Mifsud et al. [154], in which the evolutionary history of the Flaviviridae family was reconstructed with unprecedented detail, with protein structure predictions employed to define the relationships among distant groups, define major clades, discover new members and define novel and acquired proteins across different genera. Importantly, this kind of structural phylogenetic resolution, reaching evolutionary relationships invisible to sequence-based methods, directly informs virus identification, providing a framework within which highly divergent or novel viruses can be placed and recognized even in the absence of meaningful sequence similarity. Moreover, the integration of structural phylogenetics into viral classification schemes could contribute to a more universal and evolutionarily coherent taxonomy, in which the boundaries of viral families and higher-order groupings are defined not by arbitrary sequence similarity thresholds, but by genuine shared evolutionary history reflected in protein architecture.

Looking forward, the next generation of virus discovery tools will likely combine structure-based homology, DL and large-scale metagenomic data (including further expanded protein structure databases) within unified frameworks. Equally important will be the incorporation of experimental validation to confirm computational predictions, together with ecological and environmental context to better understand viral distributions and functions. This integration would greatly enhance our ability to detect and classify viruses across many different scenarios, from low-diversity population-based studies in which closely related strains need to be tracked to high-divergence phylogenomic analyses of distantly related taxonomic groups. Such approaches have the potential not only to improve detection of highly divergent viruses but also to enhance functional annotation, evolutionary inference and ecological interpretation. As these methods mature, they may substantially reshape our view of the virosphere and provide a more complete understanding of viral diversity across ecosystems.

Supplementary material

Table S1.
mgen-12-01785-s001.xlsx (10.6KB, xlsx)
DOI: 10.1099/mgen.0.001785
Table S2.
mgen-12-01785-s002.xlsx (66.5KB, xlsx)
DOI: 10.1099/mgen.0.001785
Table S3.
mgen-12-01785-s003.xlsx (21.6KB, xlsx)
DOI: 10.1099/mgen.0.001785

Abbreviations

AI

artificial intelligence

CNN

convolutional neural network

3Di characters

three-dimensional interaction characters

DL

deep learning

HMM

hidden Markov model

ICTV

International Committee on Taxonomy of Viruses

MDA

multiple displacement amplification

ML

machine learning

RdRp

RNA-dependent RNA polymerase

vOTU

viral operational taxonomic unit

Footnotes

Funding: This review reflects numerous discussions among the authors within our lab and is primarily supported by the Natural Sciences and Engineering Research Council of Canada (grants CRC 2023-00102 and RGPIN-2024-05779), Alberta Innovates (grant 252608357), Burroughs Wellcome Fund (grant 1468853), Canadian Institute for Advanced Research (grants CF-0471, CF-0600, CF-0608 and CF-0611) to A.C.G. The review is also supported by the Alberta Innovates Postdoctoral Fellowship to A.A.S.d.A. and the University of Calgary VPR Postdoctoral Match-Funding Program and Snyder Institute Beverley Phillips Rising Star Postdoctoral Fellowship to A.S.

Author contributions: All authors contributed to the outline and structure of the review, as well as data compilation and analyses. A.A.S.d.A., A.S. and T.J. contributed to the writing, figure generation and editing. A.C.G. conceived the review, contributed to writing and oversaw the editing process. All authors approved the final version.

Contributor Information

Amanda Araujo Serrao de Andrade, Email: amanda.andrade@ucalgary.ca.

Andrea Silverj, Email: andrea.silverj@ucalgary.ca.

Theodore Josephs, Email: theodore.josephs@ucalgary.ca.

Ann C. Gregory, Email: ann.gregory@ucalgary.ca.

References

  • 1.Leeuwenhoek AV. Observations, communicated to the publisher by Mr. Antony van Leewenhoeck, in a dutch letter of the 9th Octob. 1676. here English’d: concerning little animals by him observed in rain-well-sea- and snow water; as also in water wherein pepper had lain infused. Philos Trans R Soc Lond. 1677;12:821–831. doi: 10.1098/rstl.1677.0003. [DOI] [Google Scholar]
  • 2.Ivanowski D. Ueber die mosaikkrankheit der tabakspflanze. st. petersb. Acad Imp Sci Bul. 1892;35:67–70. [Google Scholar]
  • 3.Bos L. Beijerinck’s work on tobacco mosaic virus: historical context and legacy. Philos Trans R Soc Lond B Biol Sci . 1999;354:675–685. doi: 10.1098/rstb.1999.0420. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 4.Woese CR, Fox GE. Phylogenetic structure of the prokaryotic domain: the primary kingdoms. Proc Natl Acad Sci U S A. 1977;74:5088–5090. doi: 10.1073/pnas.74.11.5088. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 5.Koonin EV, Dolja VV, Krupovic M, Varsani A, Wolf YI, et al. Global organization and proposed megataxonomy of the virus world. Microbiol Mol Biol Rev. 2020;84:e00061-19. doi: 10.1128/MMBR.00061-19. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6.Sanjuán R, Domingo-Calap P. Encyclopedia of Virology. Elsevier; Genetic Diversity and Evolution of Viral Populations; pp. 53–61. [DOI] [Google Scholar]
  • 7.Abedon ST, Yin J. Bacteriophage plaques: theory and analysis. Methods Mol Biol Clifton NJ. 2009;501:161–174. doi: 10.1007/978-1-60327-164-6_17. [DOI] [PubMed] [Google Scholar]
  • 8.d’Herelle F. Sur un microbe invisible antagoniste des bacilles dysentérique. Acad Sci Paris. 1917:373–375. [Google Scholar]
  • 9.Gouvea V, Allen JR, Glass RI, Fang ZY, Bremont M, et al. Detection of group B and C rotaviruses by polymerase chain reaction. J Clin Microbiol. 1991;29:519–523. doi: 10.1128/jcm.29.3.519-523.1991. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 10.Koonin EV, Gorbalenya AE, Chumakov KM. Tentative identification of RNA‐dependent RNA polymerases of dsRNA viruses and their relationship to positive strand RNA viral polymerases. FEBS Letters. 1989;252:42–46. doi: 10.1016/0014-5793(89)80886-5. [DOI] [PubMed] [Google Scholar]
  • 11.Kamer G, Argos P. Primary structural comparison of RNA-dependent polymerases from plant, animal and bacterial viruses. Nucl Acids Res. 1984;12:7269–7282. doi: 10.1093/nar/12.18.7269. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12.Shi M, Lin X-D, Tian J-H, Chen L-J, Chen X, et al. Redefining the invertebrate RNA virosphere. Nature. 2016;540:539–543. doi: 10.1038/nature20167. [DOI] [PubMed] [Google Scholar]
  • 13.Breitbart M, Salamon P, Andresen B, Mahaffy JM, Segall AM, et al. Genomic analysis of uncultured marine viral communities. Proc Natl Acad Sci USA. 2002;99:14250–14255. doi: 10.1073/pnas.202488399. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 14.Krishnamurthy SR, Wang D. Origins and challenges of viral dark matter. Virus Research. 2017;239:136–142. doi: 10.1016/j.virusres.2017.02.002. [DOI] [PubMed] [Google Scholar]
  • 15.Roux S, Hallam SJ, Woyke T, Sullivan MB. Viral dark matter and virus-host interactions resolved from publicly available microbial genomes. Elife. 2015;4:e08490. doi: 10.7554/eLife.08490. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16.Mettenleiter TC. The first ‘Virus hunters. Adv Virus Res. 2017;99:1–16. doi: 10.1016/bs.aivir.2017.07.005. [DOI] [PubMed] [Google Scholar]
  • 17.Burrell CJ, Howard CR, Murphy FA. History and impact of virology. Fenner Whites Med Virol. 2017:3–14. [Google Scholar]
  • 18.Mayer A. Concerning the Mosaic Disease of Tobacco. Am Phytopathol Soc Press; [Google Scholar]
  • 19.Twort FW. An investigation on the nature of ultra-microscopic viruses. The Lancet. 1915;186:1241–1243. doi: 10.1016/S0140-6736(01)20383-3. [DOI] [Google Scholar]
  • 20.Kruger DH, Schneck P, Gelderblom HR. Helmut Ruska and the visualisation of viruses. The Lancet. 2000;355:1713–1717. doi: 10.1016/S0140-6736(00)02250-9. [DOI] [PubMed] [Google Scholar]
  • 21.Ackermann H-W, Ackermann H-W. The first phage electron micrographs. Bacteriophage. 2011;1:225–227. doi: 10.4161/bact.1.4.17280. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22.Haguenau F, Hawkes PW, Hutchison JL, Satiat-Jeunemaître B, Simon GT, et al. Key events in the history of electron microscopy. Microsc Microanal. 2003;9:96–138. doi: 10.1017/S1431927603030113. [DOI] [PubMed] [Google Scholar]
  • 23.Koonin EV. The phylogeny of RNA-dependent RNA polymerases of positive-strand RNA viruses. Journal of General Virology. 1991;72:2197–2206. doi: 10.1099/0022-1317-72-9-2197. [DOI] [PubMed] [Google Scholar]
  • 24.Iyer LM, Aravind L, Koonin EV. Common origin of four diverse families of large eukaryotic DNA viruses. J Virol. 2001;75:11720–11734. doi: 10.1128/JVI.75.23.11720-11734.2001. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 25.Raoult D, Audic S, Robert C, Abergel C, Renesto P, et al. The 1.2-megabase genome sequence of Mimivirus. Science. 2004;306:1344–1350. doi: 10.1126/science.1101485. [DOI] [PubMed] [Google Scholar]
  • 26.Koonin EV, Ilyina TV. Computer-assisted dissection of rolling circle DNA replication. Biosystems. 1993;30:241–268. doi: 10.1016/0303-2647(93)90074-m. [DOI] [PubMed] [Google Scholar]
  • 27.Desiere F, Lucchini S, Brüssow H. Comparative sequence analysis of the DNA packaging, head, and tail morphogenesis modules in the temperate cos-site Streptococcus thermophilus bacteriophage Sfi21. Virology. 1999;260:244–253. doi: 10.1006/viro.1999.9830. [DOI] [PubMed] [Google Scholar]
  • 28.Duda RL, Martincic K, Hendrix RW. Genetic basis of bacteriophage HK97 prohead assembly. J Mol Biol. 1995;247:636–647. doi: 10.1006/jmbi.1994.0169. [DOI] [PubMed] [Google Scholar]
  • 29.Hendrix RW, Smith MCM, Burns RN, Ford ME, Hatfull GF. Evolutionary relationships among diverse bacteriophages and prophages: all the world’s a phage. Proc Natl Acad Sci U S A . 1999;96:2192–2197. doi: 10.1073/pnas.96.5.2192. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 30.Brüssow H, Desiere F. Comparative phage genomics and the evolution of Siphoviridae: insights from dairy phages. Mol Microbiol. 2001;39:213–222. doi: 10.1046/j.1365-2958.2001.02228.x. [DOI] [PubMed] [Google Scholar]
  • 31.McKenna R, Xia D, Willingmann P, Ilag LL, Krishnaswamy S, et al. Atomic structure of single-stranded DNA bacteriophage phi X174 and its functional implications. Nature. 1992;355:137–143. doi: 10.1038/355137a0. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 32.Ebner K, Suda M, Watzinger F, Lion T. Molecular detection and quantitative analysis of the entire spectrum of human adenoviruses by a two-reaction real-time PCR assay. J Clin Microbiol. 2005;43:3049–3053. doi: 10.1128/JCM.43.7.3049-3053.2005. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 33.Hu SL, Hays WW, Potts DE. Sequence homology between bovine and human adenoviruses. J Virol. 1984;49:604–608. doi: 10.1128/jvi.49.2.604-608.1984. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 34.McClure MA, Johnson MS, Feng DF, Doolittle RF. Sequence comparisons of retroviral proteins: relative rates of change and general phylogeny. Proc Natl Acad Sci USA. 1988;85:2469–2473. doi: 10.1073/pnas.85.8.2469. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 35.Mann NH, Cook A, Millard A, Bailey S, Clokie M. Marine ecosystems: bacterial photosynthesis genes in a virus. Nature. 2003;424:741. doi: 10.1038/424741a. [DOI] [PubMed] [Google Scholar]
  • 36.Sullivan MB, Lindell D, Lee JA, Thompson LR, Bielawski JP, et al. Prevalence and evolution of core photosystem II genes in marine cyanobacterial viruses and their hosts. PLoS Biol. 2006;4:e234. doi: 10.1371/journal.pbio.0040234. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 37.Zhong Y, Chen F, Wilhelm SW, Poorvin L, Hodson RE. Phylogenetic diversity of marine cyanophage isolates and natural virus communities as revealed by sequences of viral capsid assembly protein gene g20. Appl Environ Microbiol . 2002;68:1576–1584. doi: 10.1128/AEM.68.4.1576-1584.2002. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 38.Fuller NJ, Wilson WH, Joint IR, Mann NH. Occurrence of a sequence in marine cyanophages similar to that of T4 g20 and its application to PCR-based detection and quantification techniques. Appl Environ Microbiol. 1998;64:2051–2060. doi: 10.1128/AEM.64.6.2051-2060.1998. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 39.Rao VB, Black LW. Evidence that a phage T4 DNA packaging enzyme is a processed form of the major capsid gene product. Cell. 1985;42:967–977. doi: 10.1016/0092-8674(85)90293-4. [DOI] [PubMed] [Google Scholar]
  • 40.Tétart F, Desplats C, Kutateladze M, Monod C, Ackermann HW, et al. Phylogeny of the major head and tail genes of the wide-ranging T4-type bacteriophages. J Bacteriol . 2001;183:358–366. doi: 10.1128/JB.183.1.358-366.2001. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 41.Koonin EV, Gorbalenya AE, Chumakov KM. Tentative identification of RNA‐dependent RNA polymerases of dsRNA viruses and their relationship to positive strand RNA viral polymerases. FEBS Letters. 1989;252:42–46. doi: 10.1016/0014-5793(89)80886-5. [DOI] [PubMed] [Google Scholar]
  • 42.Kamer G, Argos P. Primary structural comparison of RNA-dependent polymerases from plant, animal and bacterial viruses. Nucl Acids Res. 1984;12:7269–7282. doi: 10.1093/nar/12.18.7269. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 43.Ray J, Dondrup M, Modha S, Steen IH, Sandaa R-A, et al. Finding a needle in the virus metagenome haystack--micro-metagenome analysis captures a snapshot of the diversity of a bacteriophage armoire. PLOS One. 2012;7:e34238. doi: 10.1371/journal.pone.0034238. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 44.Rodriguez-Brito B, Li L, Wegley L, Furlan M, Angly F, et al. Viral and microbial community dynamics in four aquatic environments. ISME J. 2010;4:739–751. doi: 10.1038/ismej.2010.1. [DOI] [PubMed] [Google Scholar]
  • 45.Emerson JB, Thomas BC, Andrade K, Allen EE, Heidelberg KB, et al. Dynamic viral populations in hypersaline systems as revealed by metagenomic assembly. Appl Environ Microbiol. 2012;78:6309–6320. doi: 10.1128/AEM.01212-12. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 46.Mizuno CM, Rodriguez-Valera F, Kimes NE, Ghai R. Expanding the marine virosphere using metagenomics. PLoS Genet. 2013;9:e1003987. doi: 10.1371/journal.pgen.1003987. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 47.Yoshida M, Takaki Y, Eitoku M, Nunoura T, Takai K. Metagenomic analysis of viral communities in (Hado)Pelagic Sediments. PLOS ONE. 2013;8:e57271. doi: 10.1371/journal.pone.0057271. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 48.Garcia-Heredia I, Martin-Cuadrado A-B, Mojica FJM, Santos F, Mira A, et al. Reconstructing viral genomes from the environment using fosmid clones: the case of haloviruses. PLoS One. 2012;7:e33802. doi: 10.1371/journal.pone.0033802. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 49.Pride DT, Salzman J, Haynes M, Rohwer F, Davis-Long C, et al. Evidence of a robust resident bacteriophage population revealed through analysis of the human salivary virome. ISME J. 2012;6:915–926. doi: 10.1038/ismej.2011.169. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 50.Reyes A, Haynes M, Hanson N, Angly FE, Heath AC, et al. Viruses in the faecal microbiota of monozygotic twins and their mothers. Nature. 2010;466:334–338. doi: 10.1038/nature09199. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 51.Robles-Sikisaka R, Ly M, Boehm T, Naidu M, Salzman J, et al. Association between living environment and human oral viral ecology. ISME J. 2013;7:1710–1724. doi: 10.1038/ismej.2013.63. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 52.Bergh Ø, BØrsheim KY, Bratbak G, Heldal M. High abundance of viruses found in aquatic environments. Nature. 1989;340:467–468. doi: 10.1038/340467a0. [DOI] [PubMed] [Google Scholar]
  • 53.Rusch DB, Halpern AL, Sutton G, Heidelberg KB, Williamson S, et al. The sorcerer II global ocean sampling expedition: northwest Atlantic through eastern tropical Pacific. PLoS Biol. 2007;5:e77. doi: 10.1371/journal.pbio.0050077. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 54.Spanish National Research Council (CSIC) Expedición Malaspina. 2010. [25-March-2026]. https://www.expedicionmalaspina.es/Malaspina/Main.html#content:Home accessed.
  • 55.Guidi L, Chaffron S, Bittner L, Eveillard D, Larhlimi A, et al. Plankton networks driving carbon export in the oligotrophic ocean. Nature. 2016;532:465–470. doi: 10.1038/nature16942. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 56.Brum JR, Ignacio-Espinoza JC, Roux S, Doulcier G, Acinas SG, et al. Patterns and ecological drivers of ocean viral communities. Science. 2015;348:1261498. doi: 10.1126/science.1261498. [DOI] [PubMed] [Google Scholar]
  • 57.Roux S, Brum JR, Dutilh BE, Sunagawa S, Duhaime MB, et al. Ecogenomics and potential biogeochemical impacts of globally abundant ocean viruses. Nature. 2016;537:689–693. doi: 10.1038/nature19366. [DOI] [PubMed] [Google Scholar]
  • 58.Tian F, Wainaina JM, Howard-Varona C, Domínguez-Huerta G, Bolduc B, et al. Prokaryotic-virus-encoded auxiliary metabolic genes throughout the global oceans. Microbiome. 2024;12:159. doi: 10.1186/s40168-024-01876-z. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 59.Gregory AC, Zablocki O, Zayed AA, Howell A, Bolduc B, et al. The gut virome database reveals age-dependent patterns of virome diversity in the human gut. Cell Host Microbe. 2020;28:724–740. doi: 10.1016/j.chom.2020.08.003. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 60.Camarillo-Guerrero LF, Almeida A, Rangel-Pineros G, Finn RD, Lawley TD. Massive expansion of human gut bacteriophage diversity. Cell. 2021;184:1098–1109. doi: 10.1016/j.cell.2021.01.029. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 61.Nayfach S, Páez-Espino D, Call L, Low SJ, Sberro H, et al. Metagenomic compendium of 189,680 DNA viruses from the human gut microbiome. Nat Microbiol. 2021;6:960–970. doi: 10.1038/s41564-021-00928-6. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 62.Zolfo M, Silverj A, Blanco-Miguez A, Manghi P, Rota-Stabelli O, et al. Discovering and exploring the hidden diversity of human gut viruses using highly enriched virome samples. Microbiology. 2024 doi: 10.1101/2024.02.19.580813. [DOI]
  • 63.Graham EB, Camargo AP, Wu R, Neches RY, Nolan M, et al. A global atlas of soil viruses reveals unexplored biodiversity and potential biogeochemical impacts. Nat Microbiol. 2024;9:1873–1883. doi: 10.1038/s41564-024-01686-x. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 64.Wu Z, Liu T, Chen Q, Chen T, Hu J, et al. Unveiling the unknown viral world in groundwater. Nat Commun. 2024;15:6788. doi: 10.1038/s41467-024-51230-y. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 65.Zhang X, Huang L, Zhang X. Ancient deep-sea environmental virome provides insights into the evolution of human pathogenic RNA viruses. Resour Environ Sustain. 2024;18:100175. doi: 10.1016/j.resenv.2024.100175. [DOI] [Google Scholar]
  • 66.Guo R. cnGVC - Chinese gut virus catalogue. 2025 doi: 10.5281/ZENODO.14671176. Epub ahead of print 16 January 2025. [DOI] [PMC free article] [PubMed]
  • 67.Fiamenghi MB, Camargo AP, Chasapi IN, Baltoumas FA, Roux S, et al. Meta-virus resource (metavr): expanding the frontiers of viral diversity with 24 million uncultivated virus genomes. Nucleic Acids Res. 2026;54:D801–D812. doi: 10.1093/nar/gkaf1283. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 68.Kim K-H, Bae J-W. Amplification methods bias metagenomic libraries of uncultured single-stranded and double-stranded DNA Viruses. Appl Environ Microbiol. 2011;77:7663–7668. doi: 10.1128/AEM.00289-11. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 69.Roux S, Solonenko NE, Dang VT, Poulos BT, Schwenck SM, et al. Towards quantitative viromics for both double-stranded and single-stranded DNA viruses. PeerJ. 2016;4:e2777. doi: 10.7717/peerj.2777. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 70.Roux S, Matthijnssens J, Dutilh BE. Encyclopedia of Virology. Elsevier; Metagenomics in Virology; pp. 133–140. [DOI] [Google Scholar]
  • 71.Thurber RV, Haynes M, Breitbart M, Wegley L, Rohwer F. Laboratory procedures to generate viral metagenomes. Nat Protoc. 2009;4:470–483. doi: 10.1038/nprot.2009.10. [DOI] [PubMed] [Google Scholar]
  • 72.Klenner J, Kohl C, Dabrowski PW, Nitsche A. Comparing viral metagenomic extraction methods. Curr Issues Mol Biol. 2017;24:59–70. doi: 10.21775/cimb.024.059. [DOI] [PubMed] [Google Scholar]
  • 73.Slavov SN. Routine detection of viruses through metagenomics: where do we stand? Am J Trop Med Hyg. 2025;112:479–480. doi: 10.4269/ajtmh.24-0652. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 74.Brault AC, Blitvich BJ. Continued need for comprehensive genetic and phenotypic characterization of viruses: benefits of complementing sequence analyses with functional determinations. Am J Trop Med Hyg. 2018;98:1213. doi: 10.4269/ajtmh.18-0128. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 75.Schuettenberg A, Piña A, Metrailer M, Peláez-Sánchez RG, Agudelo-Flórez P, et al. Highly multiplexed serology for nonhuman mammals. Microbiol Spectr. 2022;10:e0287322. doi: 10.1128/spectrum.02873-22. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 76.Goldsmith CS, Miller SE. Modern uses of electron microscopy for detection of viruses. Clin Microbiol Rev. 2009;22:552–563. doi: 10.1128/CMR.00027-09. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 77.Popov VL, Tesh RB, Weaver SC, Vasilakis N. Electron microscopy in discovery of novel and emerging viruses from the collection of the world reference center for emerging viruses and arboviruses (WRCEVA) Viruses. 2019;11:477. doi: 10.3390/v11050477. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 78.Romero-Brey I, Bartenschlager R. Viral infection at high magnification: 3D electron microscopy methods to analyze the architecture of infected cells. Viruses. 2015;7:6316–6345. doi: 10.3390/v7122940. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 79.Porter AF, Cobbin J, Li C-X, Eden J-S, Holmes EC. Metagenomic identification of viral sequences in laboratory reagents. Viruses. 2021;13:2122. doi: 10.3390/v13112122. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 80.Wang T, Chen S, Wang Y, Zhang Y, Song X, et al. From in silico to in vitro: a comprehensive guide to validating bioinformatics findings
  • 81.Margulies M, Egholm M, Altman WE, Attiya S, Bader JS, et al. Genome sequencing in microfabricated high-density picolitre reactors. Nature. 2005;437:376–380. doi: 10.1038/nature03959. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 82.Zerbino DR, Birney E. Velvet: Algorithms for de novo short read assembly using de Bruijn graphs. Genome Res. 2008;18:821–829. doi: 10.1101/gr.074492.107. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 83.Bankevich A, Nurk S, Antipov D, Gurevich AA, Dvorkin M, et al. SPAdes: a new genome assembly algorithm and its applications to single-cell sequencing. J Comput Biol. 2012;19:455–477. doi: 10.1089/cmb.2012.0021. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 84.Li D, Liu C-M, Luo R, Sadakane K, Lam T-W. MEGAHIT: an ultra-fast single-node solution for large and complex metagenomics assembly via succinct de Bruijn graph. Bioinformatics. 2015;31:1674–1676. doi: 10.1093/bioinformatics/btv033. [DOI] [PubMed] [Google Scholar]
  • 85.Nurk S, Meleshko D, Korobeynikov A, Pevzner PA. metaSPAdes: a new versatile metagenomic assembler. Genome Res. 2017;27:824–834. doi: 10.1101/gr.213959.116. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 86.Antipov D, Raiko M, Lapidus A, Pevzner PA. Metaviral SPAdes: assembly of viruses from metagenomic data. Bioinformatics. 2020;36:4126–4129. doi: 10.1093/bioinformatics/btaa490. [DOI] [PubMed] [Google Scholar]
  • 87.Namiki T, Hachiya T, Tanaka H, Sakakibara Y. MetaVelvet: an extension of Velvet assembler to de novo metagenome assembly from short sequence reads. Nucleic Acids Res. 2012;40:e155. doi: 10.1093/nar/gks678. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 88.Kolmogorov M, Bickhart DM, Behsaz B, Gurevich A, Rayko M, et al. metaFlye: scalable long-read metagenome assembly using repeat graphs. Nat Methods. 2020;17:1103–1110. doi: 10.1038/s41592-020-00971-x. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 89.Sutton TDS, Clooney AG, Ryan FJ, Ross RP, Hill C. Choice of assembly software has a critical impact on virome characterisation. Microbiome. 2019;7:12. doi: 10.1186/s40168-019-0626-5. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 90.Altschul SF, Gish W, Miller W, Myers EW, Lipman DJ. Basic local alignment search tool. J Mol Biol. 1990;215:403–410. doi: 10.1016/S0022-2836(05)80360-2. [DOI] [PubMed] [Google Scholar]
  • 91.Guo J, Bolduc B, Zayed AA, Varsani A, Dominguez-Huerta G, et al. VirSorter2: a multi-classifier, expert-guided approach to detect diverse DNA and RNA viruses. Microbiome. 2021;9:37. doi: 10.1186/s40168-020-00990-y. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 92.Kieft K, Zhou Z, Anantharaman K. VIBRANT: automated recovery, annotation and curation of microbial viruses, and evaluation of viral community function from genomic sequences. Microbiome. 2020;8:90. doi: 10.1186/s40168-020-00867-0. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 93.Camargo AP, Roux S, Schulz F, Babinski M, Xu Y, et al. Identification of mobile genetic elements with geNomad. Nat Biotechnol. 2024;42:1303–1312. doi: 10.1038/s41587-023-01953-y. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 94.Ren J, Song K, Deng C, Ahlgren NA, Fuhrman JA, et al. Identifying viruses from metagenomic data using deep learning. Quant Biol. 2020;8:64–77. doi: 10.1007/s40484-019-0187-4. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 95.Pearson WR. An introduction to sequence similarity (“Homology”) searching. CP in Bioinformatics . 2013;42 doi: 10.1002/0471250953.bi0301s42. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 96.Eddy SR. Accelerated profile HMM searches. PLoS Comput Biol. 2011;7:e1002195. doi: 10.1371/journal.pcbi.1002195. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 97.Steinegger M, Söding J. MMseqs2 enables sensitive protein sequence searching for the analysis of massive data sets. Nat Biotechnol. 2017;35:1026–1028. doi: 10.1038/nbt.3988. [DOI] [PubMed] [Google Scholar]
  • 98.Buchfink B, Reuter K, Drost H-G. Sensitive protein alignments at tree-of-life scale using DIAMOND. Nat Methods. 2021;18:366–368. doi: 10.1038/s41592-021-01101-x. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 99.Buchfink B, Xie C, Huson DH. Fast and sensitive protein alignment using DIAMOND. Nat Methods. 2015;12:59–60. doi: 10.1038/nmeth.3176. [DOI] [PubMed] [Google Scholar]
  • 100.Mistry J, Chuguransky S, Williams L, Qureshi M, Salazar GA, et al. Pfam: The protein families database in 2021. Nucleic Acids Res. 2021;49:D412–D419. doi: 10.1093/nar/gkaa913. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 101.Skewes-Cox P, Sharpton TJ, Pollard KS, DeRisi JL. Profile hidden Markov models for the detection of viruses within metagenomic sequence data. PLoS ONE. 2014;9:e105067. doi: 10.1371/journal.pone.0105067. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 102.Zayed AA, Lücking D, Mohssen M, Cronin D, Bolduc B, et al. efam: an expanded, metaproteome-supported HMM profile database of viral protein families. Bioinformatics. 2021;37:4202–4208. doi: 10.1093/bioinformatics/btab451. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 103.Trgovec-Greif L, Hellinger H-J, Mainguy J, Pfundner A, Frishman D, et al. VOGDB-database of virus orthologous groups. Viruses. 2024;16:1191. doi: 10.3390/v16081191. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 104.Terzian P, Olo Ndela E, Galiez C, Lossouarn J, Pérez Bucio RE, et al. PHROG: families of prokaryotic virus proteins clustered using remote homology. NAR Genomics and Bioinformatics. 2021;3:lqab067. doi: 10.1093/nargab/lqab067. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 105.Neri U, Wolf YI, Roux S, Camargo AP, Lee B, et al. Expansion of the global RNA virome reveals diverse clades of bacteriophages. Cell. 2022;185:4023–4037. doi: 10.1016/j.cell.2022.08.023. [DOI] [PubMed] [Google Scholar]
  • 106.Santiago-Rodriguez TM, Hollister EB. Unraveling the viral dark matter through viral metagenomics. Front Immunol. 2022;13:1005107. doi: 10.3389/fimmu.2022.1005107. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 107.Zielezinski A, Vinga S, Almeida J, Karlowski WM. Alignment-free sequence comparison: benefits, applications, and tools. Genome Biol. 2017;18:186. doi: 10.1186/s13059-017-1319-7. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 108.Sinno A, Baghdadi R, Narch R, El Rayes S, Tokajian S, et al. Charting the virosphere: computational synergies of AI and bioinformatics in viral discovery and evolution. J Virol. 2025;99:e0155425. doi: 10.1128/jvi.01554-25. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 109.Ren J, Ahlgren NA, Lu YY, Fuhrman JA, Sun F. VirFinder: a novel k-mer based tool for identifying viral sequences from assembled metagenomic data. Microbiome. 2017;5:69. doi: 10.1186/s40168-017-0283-5. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 110.Tampuu A, Bzhalava Z, Dillner J, Vicente R. ViraMiner: Deep learning on raw DNA sequences for identifying viral genomes in human samples. PLOS ONE. 2019;14:e0222271. doi: 10.1371/journal.pone.0222271. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 111.Shang J, Jiang J, Sun Y. Bacteriophage classification for assembled contigs using graph convolutional network. Bioinformatics. 2021;37:i25–i33. doi: 10.1093/bioinformatics/btab293. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 112.Fang Z, Tan J, Wu S, Li M, Xu C, et al. PPR-Meta: a tool for identifying phages and plasmids from metagenomic fragments using deep learning. GigaScience. 2019;8:giz066. doi: 10.1093/gigascience/giz066. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 113.Wu L-Y, Wijesekara Y, Piedade GJ, Pappas N, Brussaard CPD, et al. Benchmarking bioinformatic virus identification tools using real-world metagenomic data across biomes. Genome Biol. 2024;25:97. doi: 10.1186/s13059-024-03236-4. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 114.Hegarty B, Riddell V J, Bastien E, Langenfeld K, Lindback M, et al. Benchmarking informatics approaches for virus discovery: caution is needed when combining in silico identification methods. mSystems. 2024;9:e01105–23. doi: 10.1128/msystems.01105-23. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 115.Huang IK, Pei J, Grishin NV. Defining and predicting structurally conserved regions in protein superfamilies. Bioinformatics. 2013;29:175–181. doi: 10.1093/bioinformatics/bts682. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 116.Cheng S, Brooks CL. Viral capsid proteins are segregated in structural fold space. PLoS Comput Biol. 2013;9:e1002905. doi: 10.1371/journal.pcbi.1002905. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 117.Krupovic M, Koonin EV. Multiple origins of viral capsid proteins from cellular ancestors. Proc Natl Acad Sci USA. 2017;114:E2401–E2410. doi: 10.1073/pnas.1621061114. Epub ahead of print 21 March 2017. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 118.Jumper J, Evans R, Pritzel A, Green T, Figurnov M, et al. Highly accurate protein structure prediction with AlphaFold. Nature. 2021;596:583–589. doi: 10.1038/s41586-021-03819-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 119.Holm L. DALI and the persistence of protein shape. Protein Sci. 2020;29:128–140. doi: 10.1002/pro.3749. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 120.van Kempen M, Kim SS, Tumescheit C, Mirdita M, Lee J, et al. Fast and accurate protein structure search with Foldseek. Nat Biotechnol. 2024;42:243–246. doi: 10.1038/s41587-023-01773-0. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 121.Edgar RC. Protein structure alignment by Reseek improves sensitivity to remote homologs. Bioinformatics. 2024;40:btae687. doi: 10.1093/bioinformatics/btae687. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 122.Say H, Joris B, Giguere D, Gloor GB. Annotating metagenomically assembled bacteriophage from a unique ecological system using protein structure prediction and structure homology search. Genomics. 2023 doi: 10.1101/2023.04.19.537516. Epub ahead of print 21 April 2023. [DOI]
  • 123.Bouras G, Grigson SR, Mirdita M, Heinzinger M, Papudeshi B, et al. Protein structure-informed bacteriophage genome annotation with Phold. Nucleic Acids Res. 2026;54:gkaf1448. doi: 10.1093/nar/gkaf1448. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 124.Hou X, He Y, Fang P, Mei S-Q, Xu Z, et al. Using artificial intelligence to document the hidden RNA virosphere. Cell. 2024;187:6929–6942. doi: 10.1016/j.cell.2024.09.027. [DOI] [PubMed] [Google Scholar]
  • 125.Kim RS, Levy Karin E, Mirdita M, Chikhi R, Steinegger M. BFVD—a large repository of predicted viral protein structures. Nucleic Acids Res. 2025;53:D340–D347. doi: 10.1093/nar/gkae1119. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 126.Litvin U, Lytras S, Jack A, Robertson DL, Hughes J, et al. Viro3D: a comprehensive database of virus protein structure predictions. Mol Syst Biol. 2025;21:1599–1617. doi: 10.1038/s44320-025-00147-9. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 127.Nomburg J, Doherty EE, Price N, Bellieny-Rabelo D, Zhu YK, et al. Birth of protein folds and functions in the virome. Nature. 2024;633:710–717. doi: 10.1038/s41586-024-07809-y. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 128.Odai R, Leemann M, Al-Murad T, Abdullah M, Shyrokova L, et al. The Viral AlphaFold Database of monomers and homodimers reveals conserved protein folds in viruses of bacteria, archaea, and eukaryotes. Sci Adv. 2025;11:eadz8560. doi: 10.1126/sciadv.adz8560. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 129.Fiamenghi MB, Camargo AP, Chasapi IN, Baltoumas FA, Roux S, et al. Meta-virus resource (MetaVR): expanding the frontiers of viral diversity with 24 million uncultivated virus genomes. Nucleic Acids Res. 2026;54:D801–D812. doi: 10.1093/nar/gkaf1283. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 130.Varadi M, Anyango S, Deshpande M, Nair S, Natassia C, et al. AlphaFold protein structure database: massively expanding the structural coverage of protein-sequence space with high-accuracy models. Nucleic Acids Res. 2022;50:D439–D444. doi: 10.1093/nar/gkab1061. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 131.Berman HM, Westbrook J, Feng Z, Gilliland G, Bhat TN, et al. The protein data bank. Nucleic Acids Res. 2000;28:235–242. doi: 10.1093/nar/28.1.235. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 132.Abramson J, Adler J, Dunger J, Evans R, Green T, et al. Accurate structure prediction of biomolecular interactions with AlphaFold 3. Nature. 2024;630:493–500. doi: 10.1038/s41586-024-07487-w. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 133.Mehdiabadi M, Tosatto SCE, Piovesan D. Modeling intrinsically disordered regions from AlphaFold2 to AlphaFold3. Protein Science. 2026;35:e70426. doi: 10.1002/pro.70426. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 134.Baek M, DiMaio F, Anishchenko I, Dauparas J, Ovchinnikov S, et al. Accurate prediction of protein structures and interactions using a three-track neural network. Science. 2021;373:871–876. doi: 10.1126/science.abj8754. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 135.Lin Z, Akin H, Rao R, Hie B, Zhu Z, et al. Evolutionary-scale prediction of atomic-level protein structure with a language model. Science. 2023;379:1123–1130. doi: 10.1126/science.ade2574. [DOI] [PubMed] [Google Scholar]
  • 136.Passaro S, Corso G, Wohlwend J, Reveiz M, Thaler S, et al. Boltz-2: towards accurate and efficient binding affinity prediction. Molecular Biology. 2025 doi: 10.1101/2025.06.14.659707. [DOI]
  • 137.Liu L, Zhang S, Xue Y, Ye X, Zhu K, et al. Technical report of helixfold3 for biomolecular structure prediction. 2024 doi: 10.48550/ARXIV.2408.16975. Epub ahead of print 2024. [DOI]
  • 138.Qiao L, Yan H, Liu G, Guo G, Sun S. IntelliFold-2: Surpassing AlphaFold 3 via architectural refinement and structural consistency. Bioinformatics . doi: 10.64898/2026.02.09.704787. [DOI]
  • 139.Puente-Lelievre C, Malik A, Douglas J. Protein structural phylogenetics. Genome Biol Evol. 2025;17:evaf139. doi: 10.1093/gbe/evaf139. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 140.Bou Dagher L, Madern D, Malbos P, Brochier-Armanet C. Persistent homology reveals strong phylogenetic signal in 3D protein structures. PNAS Nexus . 2024;3:gae158. doi: 10.1093/pnasnexus/pgae158. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 141.Ghafari M, Simmonds P, Pybus OG, Katzourakis A. A mechanistic evolutionary model explains the time-dependent pattern of substitution rates in viruses. Curr Biol. 2021;31:4689–4696. doi: 10.1016/j.cub.2021.08.020. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 142.Moi D, Bernard C, Steinegger M, Nevers Y, Langleib M, et al. Structural phylogenetics unravels the evolutionary diversification of communication systems in gram-positive bacteria and their viruses. Nat Struct Mol Biol. 2025;32:2492–2502. doi: 10.1038/s41594-025-01649-8. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 143.Guyon F, Camproux A-C, Hochez J, Tufféry P. SA-Search: a web tool for protein structure mining based on a Structural Alphabet. Nucleic Acids Res. 2004;32:W545–8. doi: 10.1093/nar/gkh467. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 144.Pantolini L, Studer G, Engist L, Pudžiuvelytė I, Pommerening F, et al. Rewriting protein alphabets with language models. Bioinformatics. 2025 doi: 10.1101/2025.11.27.690975. [DOI]
  • 145.Gilchrist CLM, Mirdita M, Steinegger M. Multiple protein structure alignment at scale with FoldMason. Science. 2026;391:485–488. doi: 10.1126/science.ads6733. [DOI] [PubMed] [Google Scholar]
  • 146.Kim D, Park S, Steinegger M. Unicore enables scalable and accurate phylogenetic reconstruction with structural core genes. Genome Biol Evol. 2025;17:evaf109. doi: 10.1093/gbe/evaf109. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 147.Heinzinger M, Weissenow K, Sanchez JG, Henkel A, Mirdita M, et al. Bilingual language model for protein sequence and structure. NAR Genomics Bioinforma. 2024;6:lqae150. doi: 10.1093/nargab/lqae150. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 148.Puente-Lelievre C, Malik AJ, Douglas J, Ascher D, Baker M, et al. Tertiary-interaction characters enable fast, model-based structural phylogenetics beyond the twilight zone. Evolutionary Biology. 2023 doi: 10.1101/2023.12.12.571181. [DOI]
  • 149.Mutti G, Ocaña-Pallarès E, Gabaldón T. Newly developed structure-based methods do not outperform standard sequence-based methods for large-scale phylogenomics. Mol Biol Evol. 2025;42:msaf149. doi: 10.1093/molbev/msaf149. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 150.Baltzis A, Santus L, Langer BE, Magis C, de Vienne DM, et al. multistrap: boosting phylogenetic analyses with structural information. Nat Commun. 2025;16:293. doi: 10.1038/s41467-024-55264-0. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 151.Garg SG, Hochberg GKA. A general substitution matrix for structural phylogenetics. Mol Biol Evol. 2025;42:msaf124. doi: 10.1093/molbev/msaf124. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 152.Douglas J, Bromham L. Reconstructing substitution histories on phylogenies, with accuracy, precision, and coverage. Evol Biol. 2025 doi: 10.64898/2025.12.21.695861. Epub ahead of print 23 December 2025. [DOI]
  • 153.Holder M, Lewis PO. Phylogeny estimation: traditional and Bayesian approaches. Nat Rev Genet. 2003;4:275–284. doi: 10.1038/nrg1044. [DOI] [PubMed] [Google Scholar]
  • 154.Mifsud JCO, Lytras S, Oliver MR, Toon K, Costa VA, et al. Mapping glycoprotein structure reveals Flaviviridae evolutionary history. Nature. 2024;633:695–703. doi: 10.1038/s41586-024-07899-8. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 155.Fleming J, Magana P, Nair S, Tsenkov M, Bertoni D, et al. AlphaFold protein structure database and 3D-Beacons: new data and capabilities. J Mol Biol. 2025;437:168967. doi: 10.1016/j.jmb.2025.168967. [DOI] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Table S1.
mgen-12-01785-s001.xlsx (10.6KB, xlsx)
DOI: 10.1099/mgen.0.001785
Table S2.
mgen-12-01785-s002.xlsx (66.5KB, xlsx)
DOI: 10.1099/mgen.0.001785
Table S3.
mgen-12-01785-s003.xlsx (21.6KB, xlsx)
DOI: 10.1099/mgen.0.001785

Data Availability Statement

The scripts used to generate the plots and tables presented in the manuscript are available at https://github.com/IntegrativeViromicsLab/micro_gen_review.


Articles from Microbial Genomics are provided here courtesy of Microbiology Society

RESOURCES