Abstract
Owing to the ineffectiveness of traditional culture techniques for the vast majority of microbial species, culture-independent analyses utilizing next-generation sequencing and bioinformatics have become essential for gaining insight into microbial ecology and function. This mini-review focuses on two essential methods for obtaining genetic information from uncultured prokaryotes, metagenomics and single-cell genomics. We analyzed the registration status of uncultured prokaryotic genome data from major public databases and assessed the advantages and limitations of both the methods. Metagenomics generates a significant quantity of sequence data and multiple prokaryotic genomes using straightforward experimental procedures. However, in ecosystems with high microbial diversity, such as soil, most genes are presented as brief, disconnected contigs, and lack association of highly conserved genes and mobile genetic elements with individual species genomes. Although technically more challenging, single-cell genomics offers valuable insights into complex ecosystems by providing strain-resolved genomes, addressing issues in metagenomics. Recent technological advancements, such as long-read sequencing, machine learning algorithms, and in silico protein structure prediction, in combination with vast genomic data, have the potential to overcome the current technical challenges and facilitate a deeper understanding of uncultured microbial ecosystems and microbial dark matter genes and proteins. In light of this, it is imperative that continued innovation in both methods and technologies take place to create high-quality reference genome databases that will support future microbial research and industrial applications.
Keywords: Metagenomics, Single-cell genomics, Microbiome, Database, Metagenome-assembled genome, Single amplified genome
1. Introduction
Microbial research has historically relied on successful isolation and cultivation of microbial species. However, conventional culture techniques are ineffective for more than 99% of microbial species [1], making culture-independent analyses essential for understanding microbial ecology and functions. Breakthroughs in next-generation sequencing (NGS) and bioinformatics technologies have revolutionized the field, allowing culture-independent genome analysis of environmental microbial communities [2], [3]. The vast amount of sequence data now available in the public domain enables meta-analyses to combine data from multiple studies conducted globally. Microbial genetic data is a valuable resource for understanding microbial ecosystems and functions, as well as for identifying industrially relevant enzymes and antibiotics [4], [5].
Metagenomics is a groundbreaking technique for obtaining genomes from uncultured prokaryotes, bypassing the need to culture them [6]. This approach involves directly sequencing the DNA extracted from microbial communities and then assembling the resulting fragmented sequences into contiguous sequences using computer algorithms. The resulting contigs are then grouped into genomic sequence bins for each microbial species. The extensive genetic data provided by metagenomics provide a thorough understanding of the genomic structures and functions of complex microbial communities. Metagenomic applications include exploring the link between obesity and gut microbiota [7], delineating gut microbiota-specific pathways and metabolic modules in patients with inflammatory bowel disease (IBD) [8], and uncovering the unique functions of individual bacteria in specific environments [9].
Single-cell genomics is a method of obtaining uncultured microbial genomes by physically isolating single cells from individual microbial species, amplifying DNA, and sequencing. [10], [11], [12], [13], [14], [15], [16]. Although it requires more complex techniques than metagenomics because it treats single cells and their tiny DNA, recent advances in technology have resulted in a large amount of single-cell genome sequencing data. It is expected that single-cell genomics will provide new insights into genome resolution at the strain level and confirm the findings of metagenomic studies. [5], [17].
In this review, we have analyzed two primary methods for obtaining uncultured prokaryotic genomes: metagenomics and single-cell genomics. We described these techniques, evaluated the current state of uncultured prokaryotic genome data in public databases, and discussed the quality of gene and genome data in the databases based on the origin of the samples and the ecosystem. We then assessed the advantages and limitations of metagenomics and single-cell genomics and explored potential avenues for expanding the data to improve our understanding of uncultured prokaryotic ecosystems and facilitate the industrial application of prokaryotic genes.
2. Approaches for obtaining uncultured prokaryotic genes and genomes
2.1. Metagenomics
In shotgun metagenomics, DNA fragments extracted from prokaryotic communities are directly sequenced, and sequence reads are then computationally assembled to generate contig sequences as consensus sequences [18], [19]. These contigs, which are composed of sequences from various prokaryotes, are separated into groups to recover the genomes of the individual prokaryotes [20], [21], [22], [23]. This process and recovered genomes are called binning and Metagenome-Assembled Genomes (MAGs), respectively. Various algorithms assign contigs to groups of sequences (bins) based on characteristics, such as GC content, tetranucleotide frequency, and sequence coverage. Because no single binning approach performs well for all metagenomic sequences, bin refinement tools have been developed to consolidate sets of MAGs from different binning predictions [24], [25], [26]. According to our evaluation of the major binning tools [17], CONCOCT [20] and MaxBin 2 [21] tended to put more contigs into the bin, and contamination rates tended to be higher. In contrast, MetaBAT 2 [22] tended to perform conservative binning, and the bin tended to have low contamination and completeness. Bin refinement using DAS_Tool [24] or other tools to extract the reliable MAG from the bin is encouraged.
However, MAGs often contain chimeric sequences from different prokaryotic species [17], [27]. It has been observed that only approximately 7% of MAGs generated from short-read sequencers contain 16S rRNA genes [28], posing challenges in correlating MAGs with 16S rRNA amplicon sequencing. Furthermore, accurately sorting mobile genetic elements, such as plasmids and phages, in MAGs is challenging [29]. Ribosomal protein genes are often not included in MAGs [30]. There are several review articles on metagenomic analysis available; thus, we have not included the details here [31], [32], [33].
2.2. Single-cell genomics
In single-cell genomics, individual cells are first isolated from the prokaryotic community using flow cytometric cell sorting or microfluidics [10], [13]. Cell lysis and whole-genome amplification are then performed to obtain sufficient amounts of DNA for sequencing. Single-cell sequence reads are obtained through indexed sequencing, followed by de novo assembly of sequence reads into Single Amplified Genomes (SAGs). Because single-cell genomic sequences are obtained from individual cells, there is no need for contig binning after assembly to produce SAGs, which offers superior genome recovery of rare prokaryotes from complex prokaryotic communities. Single-cell genomics has an excellent recovery of 16S rRNA genes in SAGs and can link prokaryotic host genomes to mobile genetic elements, such as plasmids and prophages [17], [34]. Although SAGs generally exhibit lower genome completeness than MAGs and often include incorrect assemblies by chimeric sequences or external DNA contamination, these problems can be overcome by co-assembly of SAGs and chimera sequence cleaning [11]. While MAGs are population-representative sequences, SAGs are theoretically strain-resolved sequences; therefore, the quality of genome data is not affected by prokaryotic diversity or the presence of similar or dissimilar prokaryotes. Single-cell genomics applications include the analysis of bacteria visible to the naked eye [35], a comprehensive survey of marine bacteria in surface seawater [12], the identification of secondary metabolite producers from marine sponges [36], [37], the assessment of subspecies and intraspecific recombination in environmental bacterial species [38], [39], and the identification of gut bacteria that degrade soluble dietary fiber [13]. There are some technological review articles on single-cell genomics and its future perspectives [40], [41], [42].
2.3. Quality control for MAGs and SAGs
A method for assessing the quality of MAGs and SAGs [6] was proposed, which involves classifying them into four categories: finished, high-quality, medium-quality, and low-quality. This classification is based on criteria, such as the degree of genome sequence fragmentation (contig numbers), recovery of rRNA genes, number of tRNA genes, genome completeness, and contamination rate. Genome completeness and contamination are determined using single-copy marker genes with tools like CheckM [43]. High- and medium-quality MAGs or SAGs are usually employed to interpret prokaryotic functions. Open reading frames in the metagenome assembly, MAGs, and SAGs are predicted using prokaryotic gene prediction tools [44], [45], and functional analysis is carried out using COG [46], eggnog [47], and KEGG [48].
3. Sequencing data for uncultured prokaryotes in public databases
3.1. Raw sequence data
The Short Read Archive (SRA) [49] is a repository for archiving DNA sequence data generated from NGS, which is operated by the International Nucleotide Sequence Database Collaboration (INSDC) [50], including the DNA Data Bank of Japan (DDBJ) [51], European Molecular Biology Laboratory’s European Bioinformatics Institute (EMBL-EBI) [52], and National Center for Biotechnology Information (NCBI) [53]. As of September 2021, the SRA had approximately 25.6 petabases and 17 petabytes of registered DNA bases and file size, respectively [54]. In the two years between September 2021 and May 2023, the number of DNA bases registered in the SRA more than tripled, reaching over 78 petabases.
These data are publicly available, giving all users unrestricted, permanent, and free access [49]. Cloud-based platforms have been developed owing to the requirement for substantial computer resources and bioinformatics expertise for metagenomic analysis [55], [56], [57], [58]. These services provide the analysis, comparison, and storage of metagenomic data, and users can access these datasets via websites, API, or FTP sites.
3.2. SRA data collections of metagenomics and single-cell genomics
Since 2008, metagenomic data have been accumulating in the SRA (Fig. 1a), with the total number of bases exceeding one petabase by 2022. The rate of accumulation is increasing, with several projects registering terabase pair quantities. Prior to 2013, most data were derived from human-associated samples, but since 2014, there has been a significant increase in data from environmental sources. As of 2022, approximately 293 terabases of human-associated samples have been collected, whereas 323 terabases of environmental samples have been collected, which is a reversal of their earlier proportions.
Fig. 1.
Increase in the SRA data size of metagenomics and single-cell genomics over time. BioProjects with large SRA datasets are shown in boxes. Metadata of the SRA and BioSample information [131] were extracted from SRA_ Accessions.tab and biosample_set.xml.gz, respectively. For the collection of metagenomic data, SRA_Metagenome_Types.tsv provided by PARTIE Github (https://github.com/linsalrob/partie) was used to assign shotgun metagenome sequences. Additional information was extracted from the bioproject.xml and assembly_summary.txt. Single-cell genomic data, including both eukaryotes and prokaryotes, were collected from the SRA if the BioSample package was described as MISAG [6].
Although the acquisition of single-cell genome sequencing data from environmental microbes was reported in 2007 [59], they were not officially recognized as SAGs in the SRA until 2019 (Fig. 1b). The total number of base pairs exceeded seven terabases by 2022, and the rate of data accumulation is rapidly increasing. While the total sequencing effort for single-cell genomics is generally not as high as that for metagenomics, the total amount of SAG data remains small and equivalent to a single metagenomic project. Most samples have been classified as host-associated or unclassified, with only a small number of registered, environmental samples. With future technological advancements and the increased use of single-cell genomics, it is expected that data from environmental microbes will increase, similar to the growth of metagenomic data.
3.3. Web-based analysis platforms for metagenomic data
Shotgun metagenomics requires significant computational resources. According to a report [60], the assembly process requires over 65 GB of memory and 14 h of analysis time depending on the data size and tools used. If a user does not have access to extensive computer resources, they can use analysis platforms such as Integrated Microbial Genome and Microbiomes (IMG/M) [57] or MGnify [58] (previously known as EBI Metagenomics [61]) as alternatives (Table 1).
Table 1.
Data size of the shotgun metagenome dataset based on public metadata.
| Database | Metagenomes | CDSs | Acceptable data | Gene prediction | Functional annotation | Taxonomic annotation | Binning, curation, and assessment | Data accessibility |
|---|---|---|---|---|---|---|---|---|
| IMG/M | 48,866 | 33,797,010,173 | Assembly |
|
|
|
|
|
| MGnify | 33,738 | 4,821,810,124 | Raw reads or Assembly |
|
|
|
|
|
IMG/M (https://img.jgi.doe.gov) is a web-based platform for managing and analyzing metagenomic data. It contains annotated DNA and RNA sequences from various microorganisms, including cultured and uncultured bacteria, archaea, eukaryotes, and viruses. Users can upload their DNA assemblies to run the annotation pipeline, which includes predicting protein-coding sequences (CDSs) using GeneMarkS-2 [62] and Prodigal [44], detecting CRISPRs using CRT [63], and predicting RNA features and tRNAs using infernal [64] and tRNAscan-SE [65], respectively. The platform also performs functional annotations for CDSs using various databases including COG, Pfam [66], TIGRFAM [67], Cath-FumFam [68], SuperFamily [69], SMART [70], KEGG Orthology Terms (KO) [71], and Enzyme Commission (EC) numbers derived from KO terms. LAST [72] is used with the UniRef90 reference database [73] for taxonomic annotation of protein-coding genes. The platform incorporates MetaBAT [74] as a metagenome binning tool with a minimum contig cutoff of 3000 bp. After the quality assessment of MAGs by CheckM [43], taxonomic classifications are assigned using GTDB-Tk [75]. The analyzed datasets can be accessed via the website and API. Additionally, users can analyze MAGs in IMGs; however, only some metadata of the registered MAGs are available for bulk downloads (4.5%).
MGnify (https://www.ebi.ac.uk/metagenomics) is an automated pipeline that provides support for both raw reads and assembly of shotgun metagenomic data. The database currently contains 297 different biomes, with over half of the analyses originating from only nine of them: human-associated samples (fecal, oral, digestive system, skin, and unspecified human), marine, soil, mammalian digestive systems, and mixed biome samples. Users can either submit their own data for analysis or browse all analyzed public datasets available in the repository. In the pipeline, CDSs are predicted using Prodigal and FragGeneScan [76]. Non-coding RNAs are identified and annotated using Infernal, tRNAscan-SE, and Rfam [77]. Predicted genes are annotated using InterPro [78], eggNOG [47], and KEGG orthology [79]. Pathway predictions using KEGG and Genome Property [80], gene ontology term assignment, and biosynthetic gene cluster prediction using antiSMASH [81] are also performed. Taxonomic classification is carried out using MAPseq [82] and SILVA [83]. The predicted genes are compared against the UniRef 90 database using DIAMOND [84]. MAGs are recovered by MetaBAT 2 [22], MaxBin 2 [21], and CONCOCT [20] with a 2500 bp minimum contig cutoff and refined using metaWRAP [85]. Chimeric contigs are removed using GUNC [86] and MAGs are dereplicated using dRep [87]. CheckM and GTDB-Tk are performed to evaluate the quality of MAGs and taxonomy assignment, respectively. MGnify offers a website and API to access the dataset, and the protein database can be downloaded in bulk via the FTP site.
Table 1 and Fig. 2 show the latest data on shotgun metagenome assemblies in IMG/M and MGnify. IMG/M contains approximately 33 billion CDSs, whereas MGnify has approximately 4.8 billion CDSs. In terms of metagenomic assemblies, IMG/M had the most from environmental samples (n = 35,387, 25 billion CDSs), while MGnify had the most from human-associated samples (n = 18,612, 1.7 billion CDSs).
Fig. 2.
Ecosystem distribution of metagenome assemblies and genes in IMG/M and MGnify. The metagenome (n = 82, 604) and CDSs (n = 38, 618, 820, 297) fractions registered in IMG/M and MGnify were plotted for each representative ecosystem. The color indicates the databases in which the data were registered. The data are based on June 2023.
A combined view of the two major databases (Fig. 2) revealed that human-associated samples comprised 25% of the metagenomic assemblies, followed by marine, soil, freshwater, and mammal-associated samples. Conversely, the soil had the highest proportion of CDSs (30%), followed by marine, freshwater, plant, human, and mammal-associated samples. These results suggest that soil and plant microorganisms, as well as marine and freshwater microorganisms, have more diverse and non-redundant environment-specific genes than those in human-associated samples [88], indicating the importance of expanding genetic resources from environmental prokaryotes for understanding microbial function and industrial applications.
4. MAGs and SAGs in public databases
4.1. Uncultured prokaryotic genome databases and catalogs
Table 2 shows publicly available databases and data collection for uncultured prokaryotic genomes. The repositories for MAGs and SAGs were obtained from NCBI (https://www.ncbi.nlm.nih.gov/), IMG/M [57], MGnify [58], and the Genome Taxonomy Database (GTDB) (https://gtdb.ecogenomic.org/) [89]. The GTDB is a database that catalogs MAGs and SAGs to establish a standardized microbial taxonomy based on genome phylogeny using a set of single-copy marker proteins based on GTDB-Tk. Specifically, we discuss case studies of SAG, such as WGA-X at the Single Cell Genomics Center, Bigelow Laboratory for Ocean Sciences (ME, US) [12], [15], and SAG-gel (bit-MAP) by bitBiome, Inc. (Tokyo, Japan) [13], [17], [34], which are provided as analysis services with consistent data acquisition and large data sizes.
Table 2.
MAGs and SAGs in public databases.
| Database/catalog or sample (method) | References | MAGs | SAGs |
|---|---|---|---|
| NCBI | 130,149 | 17,270 | |
| IMG/M | [57] | 232,807 | 4,830 |
| GTDB | [89] | 77,891 | 831 |
| The Genomes from Earth’s Microbiomes (GEM) catalog | [90] | 52,515 | |
| MGnify | [58] | 315,252 | - |
| Unified Human Gastrointestinal Genome (UHGG) | [91] | 289,232 | - |
| GORG-Tropics (WGA-X) | [12] | - | 12,710 |
| Mouse feces (WGA-X) | [15] | - | 698 |
| Rice paddy soil (SAG-gel) | [16] | - | 4,600 |
| Human skin swab (SAG-gel) | [34], [17] | - | 768 |
| Mouse feces (SAG-gel) | [13] | - | 346 |
As of June 2023
With regard to the microbial habitats classified as ecosystems in Fig. 3a, in the IMG/M, approximately 37% of MAGs were derived from aquatic environments, followed by approximately 10% from soil environments. IMG/M includes the Genomes from Earth’s Microbiomes (GEM) catalog (https://portal.nersc.gov/GEM/), which was constructed from 10,450 metagenomes sampled from diverse microbial habitats and geographic locations [90]. Approximately 70% of MAGs in Earth’s microbiomes are derived from the human gut or marine environments. In IMG/M, CDS in the metagenome assembly were more abundant in soil than in marine environments (Fig. 2), but the number of MAGs was reversed, indicating the difficulty of constructing MAGs from soil metagenome assemblies. MGnify has the largest number of MAGs (304,283), with approximately 90% of these MAGs derived from the human gut, which is referred to as the Unified Human Gastrointestinal Genome (UHGG) catalogue [91].
Fig. 3.
Number and quality of MAGs and SAGs in public databases. Number of MAGs (top) and SAGs (bottom) (a). The MAGs (bins) and SAGs obtained from IMG/M metagenome bins, MGnify genome catalogs, and some BioProjects, including PRJEB33281, PRJDB8805, PRJNA692334, PRJNA837408, and DOI: 10.6084/m9.figshare.c.4454150, were analyzed for each ecosystem. Genome completeness of MAGs (left) and SAGs (right) (b), presence rate of the 16S rRNA gene (c), number of phyla (d), and number of CDSs per genome (e) were plotted. Only the medium- and high-quality genomes are shown in (b)–(e). The data are based on June 2023.
Although SAG datasets are one order of magnitude smaller than MAG datasets, the two datasets derived from the ocean [12] and soil environments [16] are larger than other SAG collections, with some projects acquiring thousands to tens of thousands of SAGs. To date, no cross-habitat microbial or large cohort-based genome collection projects, such as the GEM catalog or UHGG, have been undertaken in studies collecting SAGs.
4.2. Qualities of MAGs and SAGs
The statistical data for MAGs and SAGs are shown in Fig. 3. Genome completeness tended to be higher in MAGs (average of 85.3%) than in SAGs (Fig. 3b). MAGs are often selected and registered for bins of medium or higher quality in most projects or catalogs [89], [90], [91], [92], [93], [94]. However, even in human and mammalian samples with a large number of registered MAGs, few MAGs can be classified as high quality. This is due to the low recovery rate of the 16S rRNA genes, which will be discussed later.
The majority of SAGs had a completeness of less than 90%, with an average completeness of 71.0% (Fig. 3b). Conventional single-cell genomics, especially based on flow cytometric cell sorting and the conventional microtube WGA reaction, has very low genome completeness and high contamination rates due to amplification bias, chimera occurrence, and contamination [95], [96], [97]. The number of successfully amplified single cells can vary greatly from sample to sample, and sample-specific experimental optimization is necessary to obtain the best results. Freshwater and marine samples have been reported to have the highest percentage of successfully amplified genomes (up to 40%), whereas soil samples tend to have lower success rates (less than 10%) [95]. However, these shortcomings in genome amplification for single-cell genomics have been addressed by WGA-X, an improved whole-genome amplification enzyme [12], [15], or microfluidic droplets, which use droplets or gel capsules for cell isolation and genome amplification [10], [13], [14], [96], [98], [99]. SAGs derived from human skin using SAG-gel [100] were comparable in genome completeness (85.7%) to those of human-associated MAGs. Depending on the sample type, SAGs with genome completeness similar to that of MAGs were obtained (Fig. 3b). In addition, single-cell genomics has the unique feature of integrating multiple SAGs derived from cells of the same species or strain to improve genome quality [11], [14], [101]. Furthermore, by integrating SAGs with MAGs obtained from the same sample, uncultured prokaryotic genomes with improved accuracy, covering the lack of information in MAG, can be obtained [17], [102], [103], [104].
A major challenge associated with MAGs is the lack of 16S rRNA gene sequences. The presence of skewed species abundance, high 16S rRNA sequence similarity, and dependence on short-read sequencing make it difficult to assemble individual prokaryote-specific 16S rRNA genes from complex prokaryotic communities [105]. It has been reported that only 7% of MAGs have 16S rRNA gene sequences in more than 270,000 human gut MAGs, showing over 95% completeness and less than 5% contamination [28]. As shown in Fig. 3c, this lack of 16S rRNA genes was consistently observed in MAGs, with yields below 10%, particularly in MAGs from human- and mammalian-associated metagenomes. This can be attributed to a significant species bias in symbiotic bacteria and the high similarity of 16S rRNA genes between symbiotic bacteria. Conversely, 16S rRNA gene yields in SAGs were significantly higher than those in MAGs, regardless of ecosystem. The low 16S rRNA gene recovery in MAGs hinders the linking of taxonomy to functional genomic information. To address this methodological gap, the active use of SAGs and the curation of reliable MAGs [89], [90], [91], [93], [94] is crucial in microbiome research.
4.3. Genetic diversities in MAGs and SAGs
There are fewer than 20,000 prokaryotic species with valid published names, representing less than 0.2% of the estimated prokaryotic species diversity [106]. Most prokaryotes are not available as pure cultures and, therefore, cannot be named according to the rules and recommendations of the International Code of Nomenclature of Prokaryotes (ICNP). A code called SeqCode [107] was proposed to effectively publish prokaryotic names based on isolated genomes, MAGs, and SAGs. SeqCode uses genome sequence data as a common currency for typing cultivated and uncultivated microbes, and follows rules similar to those of ICNP for priority.
Prokaryotic genomes, including isolates, MAGs, and SAGs, are typically classified using taxonomic classification tools like GTDB-Tk [75]. The classification system consists of seven major ranks: species, genus, family, order, class, phylum, and domain. While many MAGs were obtained from human and mammalian samples (Fig. 3a), the number of phyla was relatively small (Fig. 3d). This suggests that the genomes of limited microbial lineages are frequently encountered in human- or mammalian-associated samples, and that the genomes of diverse microbial lineages are more likely to be found in environmental samples. For human-related microbiome samples, a thorough understanding of microdiversity is essential. Although not included in this review, various databases have been developed for gut bacteria and other organisms [91], [108], [109], [110], [111], [112], [113], [114]. Single-cell genomics should be utilized to acquire genomes of known species at the strain level from human-associated samples [17], [34], [100] and to identify the genomes of novel species in environmental samples where metagenomic binning is not feasible.
4.4. Exploring prokaryotic genes from public genomes
In terms of the number of CDSs per genome (Fig. 3e), SAG had fewer CDSs per genome (1469 CDSs) than MAG (2108 CDSs), which is consistent with the lower genome completeness of SAG compared to MAGs. SAGs obtained using SAG-gel had a similar number of CDSs (2470, 2329, and 1875 CDSs) to MAGs in humans [17], [100], soil [16], and mammal-associated samples [13], respectively. In mammals, SAGs obtained using WGA-X also had 2025 CDSs [15].
It is essential to identify full-length genes when exploring useful genes such as enzymes. To evaluate this, the average length of CDSs per metagenome assembly and MAG in IMG/M or SAG in WGA-X or SAG-gel was calculated (Fig. 4a). The MAGs and SAGs analyzed here correspond to either high- or medium-quality genomes. The average length of CDSs from MAGs and SAGs was consistently 900–1000 bp across different ecosystems. However, the average length of CDSs from metagenome assemblies was significantly shorter than those of MAGs and SAGs. In soils, the average lengths of CDSs in MAGs and SAGs were 856 bp and 908 bp, respectively, whereas the average length of CDSs in the metagenome assemblies was 481 bp. To analyze this in detail, we examined the contig numbers and sizes of soil metagenome assembly under 200 bp cutoff and revealed that 94.2% of contigs were less than 1 kbp, and these short contigs accounted for 70.8% of the total length (Fig. 4b, c). In contrast, in the SAG, short contigs (<1 kbp) accounted for 85.5% of the contigs, but their total length was only 27.7%, and long contigs (>10 kbp) accounted for over 40% of the total size. In the process of MAG construction, cutoffs of short contigs of less than 3000 bp are generally used. Thus, there were no short contigs in soil MAGs (Fig. 4b, c). The total number of contigs assigned to MAG was quite small, at 0.14% of the total metagenome assemblies, consistent with a previous report [88]. The length of the predicted CDSs in metagenomic assemblies and MAGs depends mainly on the length of the contigs. Although it is desirable to perform gene searches with reference to microbial lineages, most of the contigs on soil metagenomes are short and are discarded during the binning process; therefore, a limited number of CDSs must be used when using MAG as the search source. Therefore, there is a great possibility that soil metagenomes and MAGs may not provide adequate information as a gene discovery resource [88]. Regarding gene prediction, it is also important to consider that partial genes originating from the edges of contigs. Approximately 85% of the genes in the MGnify protein database were partial.
Fig. 4.
Average length of genes in metagenome assemblies, MAGs, and SAGs. The average length of the CDSs per sample was plotted according to ecosystem classification (a). The dataset is same as Fig. 3b. Metagenome assemblies, MAGs, and SAGs are colored green, orange, and blue, respectively. The mean and median values are indicated by yellow circles and bars, respectively. Comparisons of the presence of short contigs are presented in (b) and (c). The number of contigs and total length of soil metagenomes, MAGs, and SAGs were plotted against their abundance ratios separately for each contig length stage (b) (c). Metagenomes and MAGs from BioProject PRJNA375197 and SAGs from BioProject PRJNA869948 were used.
Most genes are specific to a single habitat, and the technical challenge is to efficiently recover rare, habitat-specific, and region-specific genes [88]. Prospects include using MAG and SAG to develop searches for specific enzymes from target prokaryotic species with characteristics such as lack of pathogenicity and industrial accessibility. This will also require improved gene function prediction techniques, including protein structure prediction and search [115], [116], [117], [118], [119], to identify unknown genes.
5. Summary and outlook
The number of uncultured prokaryotic genomes is growing rapidly, and some are publicly available in databases. Metagenomics and single-cell genomics help us identify species and their proteins in prokaryotes from various environments, such as soil, ocean, and even inside the human body. However, little is known about uncultured prokaryotic genes and proteins, beyond their nucleic acid or primary amino acid sequences, making uncultured microbial proteins the 'dark matter' of the protein universe. We are now in an era where this dark matter can be elucidated by adapting state-of-the-art protein structure prediction methods to vast uncultured prokaryotic genome data [117], [119], [120]. Advanced analysis of uncultured prokaryotic genes will help solve evolutionary history mysteries, discover proteins that can cure diseases, clean up the environment, and produce clean energy.
This mini-review provided an overview of uncultured prokaryotic genes and genomes in public databases and evaluated the quality of data available in each ecosystem. This highlighted that while shotgun metagenomics provided a large number of genes, fragmented contigs in ecosystems made it difficult to obtain full-length genes. It is challenging to construct multiple species-resolved MAGs from complex prokaryotic populations because most contigs are unassigned and discarded in the binning process [88]. The use of long-read sequencing technologies, such as PacBio [121], [122], [123] and Oxford Nanopore Technologies [124], [125], [126], [127], can help overcome these issues. Future developments in binning algorithms that leverage machine learning [23], [128] and Hi-C metagenomics [121], [129], [130] may help address these challenges.
However, the quality of SAGs is not affected by specific ecosystems, and SAG can provide complementary information to MAGs, such as 16S rRNA genes and mobile genetic elements. Single-cell genomics is a highly effective method for obtaining unknown species genomes and strain-resolved genomes, especially in environmental samples containing diverse prokaryotes. We suggest using single-cell genomics as a valuable strategy for gaining insight into and conducting a comprehensive analysis of complex ecosystems without the need for complex computing processes, such as metagenomic binning. However, challenges for single-cell genomics include expanding the number of SAGs that can be acquired in a sequencing run, reducing costs, and simplifying the method. We anticipate that continued advancements in this field will lead to the development of an integrated approach between metagenomics and single-cell genomics, resulting in a high-quality prokaryotic genome database.
Funding
This work was partially supported by MEXT/JSPS KAKENHI 21H01733 and JST FOREST JPMJFR210F.
CRediT authorship contribution statement
Koji Arikawa: Conceptualization, Validation, Software, Formal analysis, Data curation, Visualization, Investigation, Writing- Original draft preparation. Masahito Hosokawa: Conceptualization, Supervision, Writing- Reviewing and Editing, Funding acquisition.
Declaration of Competing Interest
K.A. is employed at bitBiome, Inc, which provides single‐cell genomics services using the SAG‐gel workflow as bit‐MAP. M.H. is a founder and shareholder of bitBiome, Inc.
Acknowledgments
We thank Dr. Kazuma Kamata and Dr. Tetsuro Kawano-Sugaya for supporting the survey of public databases and discussion on manuscript preparation.
References
- 1.Pham V.H.T., Kim J. Cultivation of unculturable soil bacteria. Trends Biotechnol. 2012;30:475–484. doi: 10.1016/j.tibtech.2012.05.007. [DOI] [PubMed] [Google Scholar]
- 2.Hugenholtz P., Tyson G.W. Metagenomics. Nat Publ Group UK. 2008 doi: 10.1038/455481a. [DOI] [PubMed] [Google Scholar]
- 3.Sleator R.D., Shortall C., Hill C. Metagenomics. Lett Appl Microbiol. 2008;47:361–366. doi: 10.1111/j.1472-765X.2008.02444.x. [DOI] [PubMed] [Google Scholar]
- 4.Wyman S.K., Avila-Herrera A., Nayfach S., Pollard K.S. A most wanted list of conserved microbial protein families with no known domains. PLoS One. 2018;13 doi: 10.1371/journal.pone.0205749. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 5.Robinson S.L., Piel J., Sunagawa S. A roadmap for metagenomic enzyme discovery. Nat Prod Rep. 2021;38:1994–2023. doi: 10.1039/d1np00006c. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 6.Bowers R.M., Kyrpides N.C., Stepanauskas R., Harmon-Smith M., Doud D., Reddy T.B.K., et al. Minimum information about a single amplified genome (MISAG) and a metagenome-assembled genome (MIMAG) of bacteria and archaea. Nat Biotechnol. 2017;35:725–731. doi: 10.1038/nbt.3893. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 7.Turnbaugh P.J., Ley R.E., Mahowald M.A., Magrini V., Mardis E.R., Gordon J.I. An obesity-associated gut microbiome with increased capacity for energy harvest. Nature. 2006;444:1027–1031. doi: 10.1038/nature05414. [DOI] [PubMed] [Google Scholar]
- 8.Morgan X.C., Tickle T.L., Sokol H., Gevers D., Devaney K.L., Ward D.V., et al. Dysfunction of the intestinal microbiome in inflammatory bowel disease and treatment. Genome Biol. 2012;13:R79. doi: 10.1186/gb-2012-13-9-r79. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 9.Baker B.J., Lazar C.S., Teske A.P., Dick G.J. Genomic resolution of linkages in carbon, nitrogen, and sulfur cycling among widespread estuary sediment bacteria. Microbiome. 2015;3:14. doi: 10.1186/s40168-015-0077-6. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 10.Hosokawa M., Nishikawa Y., Kogawa M., Takeyama H. Massively parallel whole genome amplification for single-cell sequencing using droplet microfluidics. Sci Rep. 2017;7:5199. doi: 10.1038/s41598-017-05436-4. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 11.Kogawa M., Hosokawa M., Nishikawa Y., Mori K., Takeyama H. Obtaining high-quality draft genomes from uncultured microbes by cleaning and co-assembly of single-cell amplified genomes. Sci Rep. 2018;8:2059. doi: 10.1038/s41598-018-20384-3. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 12.Pachiadaki M.G., Brown J.M., Brown J., Bezuidt O., Berube P.M., Biller S.J., et al. Charting the complexity of the marine microbiome through single-cell genomics. Cell. 2019;179:1623–1635. doi: 10.1016/j.cell.2019.11.017. e11. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 13.Chijiiwa R., Hosokawa M., Kogawa M., Nishikawa Y., Ide K., Sakanashi C., et al. Single-cell genomics of uncultured bacteria reveals dietary fiber responders in the mouse gut microbiota. Microbiome. 2020;8:5. doi: 10.1186/s40168-019-0779-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 14.Zheng W., Zhao S., Yin Y., Zhang H., Needham D.M., Evans E.D., et al. High-throughput, single-microbe genomics with strain resolution, applied to a human gut microbiome. Science. 2022;376:eabm1483. doi: 10.1126/science.abm1483. [DOI] [PubMed] [Google Scholar]
- 15.Lyalina S., Stepanauskas R., Wu F., Sanjabi S., Pollard K.S. Single cell genome sequencing of laboratory mouse microbiota improves taxonomic and functional resolution of this model microbial community. PLoS One. 2022;17 doi: 10.1371/journal.pone.0261795. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 16.Aoki W., Kogawa M., Matsuda S., Matsubara K., Hirata S., Nishikawa Y., et al. Massively parallel single-cell genomics of microbiomes in rice paddies. Front Microbiol. 2022;13:1024640. doi: 10.3389/fmicb.2022.1024640. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 17.Arikawa K., Ide K., Kogawa M., Saeki T., Yoda T., Endoh T., et al. Recovery of strain-resolved genomes from human microbiome through an integration framework of single-cell genomics and metagenomics. Microbiome. 2021;9:202. doi: 10.1186/s40168-021-01152-4. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 18.Li D., Liu C.-M., Luo R., Sadakane K., Lam T.-W. MEGAHIT: an ultra-fast single-node solution for large and complex metagenomics assembly via succinct de Bruijn graph. Bioinformatics. 2015;31:1674–1676. doi: 10.1093/bioinformatics/btv033. [DOI] [PubMed] [Google Scholar]
- 19.Nurk S., Meleshko D., Korobeynikov A., Pevzner P.A. metaSPAdes: a new versatile metagenomic assembler. Genome Res. 2017;27:824–834. doi: 10.1101/gr.213959.116. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 20.Alneberg J., Bjarnason B.S., de Bruijn I., Schirmer M., Quick J., Ijaz U.Z., et al. Binning metagenomic contigs by coverage and composition. Nat Methods. 2014;11:1144–1146. doi: 10.1038/nmeth.3103. [DOI] [PubMed] [Google Scholar]
- 21.Wu Y.-W., Simmons B.A., Singer S.W. MaxBin 2.0: an automated binning algorithm to recover genomes from multiple metagenomic datasets. Bioinformatics. 2016;32:605–607. doi: 10.1093/bioinformatics/btv638. [DOI] [PubMed] [Google Scholar]
- 22.Kang D., Li F., Kirton E.S., Thomas A., Egan R.S., An H., et al. MetaBAT 2: an adaptive binning algorithm for robust and efficient genome reconstruction from metagenome assemblies. PeerJ Prepr. 2019 doi: 10.7287/peerj.preprints.27522v1. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 23.Nissen J.N., Johansen J., Allesøe R.L., Sønderby C.K., Armenteros J.J.A., Grønbech C.H., et al. Improved metagenome binning and assembly using deep variational autoencoders. Nat Biotechnol. 2021;39:555–560. doi: 10.1038/s41587-020-00777-4. [DOI] [PubMed] [Google Scholar]
- 24.Sieber C.M.K., Probst A.J., Sharrar A., Thomas B.C., Hess M., Tringe S.G., et al. Recovery of genomes from metagenomes via a dereplication, aggregation and scoring strategy. Nat Microbiol. 2018;3:836–843. doi: 10.1038/s41564-018-0171-1. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 25.Uritskiy G.V., DiRuggiero J., Taylor J. MetaWRAP-a flexible pipeline for genome-resolved metagenomic data analysis. Microbiome. 2018;6:158. doi: 10.1186/s40168-018-0541-1. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 26.Rühlemann M.C., Wacker E.M., Ellinghaus D., Franke A. MAGScoT: a fast, lightweight and accurate bin-refinement tool. Bioinformatics. 2022;38:5430–5433. doi: 10.1093/bioinformatics/btac694. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 27.Shaiber A., Eren A.M. Composite metagenome-assembled genomes reduce the quality of public genome repositories. MBio. 2019:10. doi: 10.1128/mBio.00725-19. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 28.Hiseni P., Snipen L., Wilson R.C., Furu K., Rudi K. Questioning the quality of 16S rRNA gene sequences derived from human gut metagenome-assembled genomes. Front Microbiol. 2021;12 doi: 10.3389/fmicb.2021.822301. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 29.Maguire F., Jia B., Gray K.L., Lau W.Y.V., Beiko R.G., Brinkman F.S.L. Metagenome-assembled genome binning methods with short reads disproportionately fail for plasmids and genomic Islands. Micro Genom. 2020:6. doi: 10.1099/mgen.0.000436. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 30.Mise K., Iwasaki W. Unexpected absence of ribosomal protein genes from metagenome-assembled genomes. ISME Commun. 2022;2:1–9. doi: 10.1038/s43705-022-00204-6. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 31.Thomas T., Gilbert J., Meyer F. Metagenomics - a guide from sampling to data analysis. Micro Inf Exp. 2012;2:3. doi: 10.1186/2042-5783-2-3. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 32.Sharpton T.J. An introduction to the analysis of shotgun metagenomic data. Front Plant Sci. 2014;5:209. doi: 10.3389/fpls.2014.00209. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 33.Breitwieser F.P., Lu J., Salzberg S.L. A review of methods and databases for metagenomic classification and assembly. Brief Bioinform. 2019;20:1125–1136. doi: 10.1093/bib/bbx120. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 34.Hosokawa M., Endoh T., Kamata K., Arikawa K., Nishikawa Y., Kogawa M., et al. Strain-level profiling of viable microbial community by selective single-cell genome sequencing. Sci Rep. 2022;12:4443. doi: 10.1038/s41598-022-08401-y. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 35.Volland J.-M., Gonzalez-Rizzo S., Gros O., Tyml T., Ivanova N., Schulz F., et al. A centimeter-long bacterium with DNA contained in metabolically active, membrane-bound organelles. Science. 2022;376:1453–1458. doi: 10.1126/science.abb3634. [DOI] [PubMed] [Google Scholar]
- 36.Wilson M.C., Mori T., Rückert C., Uria A.R., Helf M.J., Takada K., et al. An environmental bacterial taxon with a large and distinct metabolic repertoire. Nature. 2014;506:58–62. doi: 10.1038/nature12959. [DOI] [PubMed] [Google Scholar]
- 37.Kogawa M., Miyaoka R., Hemmerling F., Ando M., Yura K., Ide K., et al. Single-cell metabolite detection and genomics reveals uncultivated talented producer. PNAS Nexus. 2022;1:gab007. doi: 10.1093/pnasnexus/pgab007. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 38.Zaremba-Niedzwiedzka K., Viklund J., Zhao W., Ast J., Sczyrba A., Woyke T., et al. Single-cell genomics reveal low recombination frequencies in freshwater bacteria of the SAR11 clade. Genome Biol. 2013;14:R130. doi: 10.1186/gb-2013-14-11-r130. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 39.Kashtan N., Roggensack S.E., Rodrigue S., Thompson J.W., Biller S.J., Coe A., et al. Single-cell genomics reveals hundreds of coexisting subpopulations in wild Prochlorococcus. Science. 2014;344:416–420. doi: 10.1126/science.1248575. [DOI] [PubMed] [Google Scholar]
- 40.Gawad C., Koh W., Quake S.R. Single-cell genome sequencing: current state of the science. Nat Rev Genet. 2016;17:175–188. doi: 10.1038/nrg.2015.16. [DOI] [PubMed] [Google Scholar]
- 41.Xu Y., Zhao F. Single-cell metagenomics: challenges and applications. Protein Cell. 2018;9:501–510. doi: 10.1007/s13238-018-0544-5. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 42.Woyke T., Doud D.F.R., Schulz F. The trajectory of microbial single-cell sequencing. Nat Methods. 2017;14:1045–1054. doi: 10.1038/nmeth.4469. [DOI] [PubMed] [Google Scholar]
- 43.Parks D.H., Imelfort M., Skennerton C.T., Hugenholtz P., Tyson G.W. CheckM: assessing the quality of microbial genomes recovered from isolates, single cells, and metagenomes. Genome Res. 2015;25:1043–1055. doi: 10.1101/gr.186072.114. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 44.Hyatt D., Chen G.-L., Locascio P.F., Land M.L., Larimer F.W., Hauser L.J. Prodigal: prokaryotic gene recognition and translation initiation site identification. BMC Bioinforma. 2010;11:119. doi: 10.1186/1471-2105-11-119. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 45.Seemann T. Prokka: rapid prokaryotic genome annotation. Bioinformatics. 2014;30:2068–2069. doi: 10.1093/bioinformatics/btu153. [DOI] [PubMed] [Google Scholar]
- 46.Galperin M.Y., Makarova K.S., Wolf Y.I., Koonin E.V. Expanded microbial genome coverage and improved protein family annotation in the COG database. Nucleic Acids Res. 2015;43:D261–D269. doi: 10.1093/nar/gku1223. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 47.Huerta-Cepas J., Szklarczyk D., Heller D., Hernández-Plaza A., Forslund S.K., Cook H., et al. eggNOG 5.0: a hierarchical, functionally and phylogenetically annotated orthology resource based on 5090 organisms and 2502 viruses. Nucleic Acids Res. 2019;47:D309–D314. doi: 10.1093/nar/gky1085. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 48.Kanehisa M., Araki M., Goto S., Hattori M., Hirakawa M., Itoh M., et al. KEGG for linking genomes to life and the environment. Nucleic Acids Res. 2008;36:D480–D484. doi: 10.1093/nar/gkm882. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 49.Leinonen R., Sugawara H., Shumway M., Collaboration I.N.S.D. The sequence read archive. Nucleic Acids Res. 2010;39:D19–D21. doi: 10.1093/nar/gkq1019. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 50.Cochrane G., Karsch-Mizrachi I., Takagi T. Sequence database collaboration IN. The international nucleotide sequence database collaboration. Nucleic Acids Res. 2016;44:D48–D50. doi: 10.1093/nar/gkv1323. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 51.Kaminuma E., Mashima J., Kodama Y., Gojobori T., Ogasawara O., Okubo K., et al. DDBJ launches a new archive database with analytical tools for next-generation sequence data. Nucleic Acids Res. 2010;38:D33–D38. doi: 10.1093/nar/gkp847. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 52.Silvester N., Alako B., Amid C., Cerdeño-Tarrága A., Clarke L., Cleland I., et al. The European nucleotide archive in 2017. Nucleic Acids Res. 2018;46:D36–D40. doi: 10.1093/nar/gkx1125. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 53.Sayers E.W., Beck J., Bolton E.E., Bourexis D., Brister J.R., Canese K., et al. Database resources of the national center for biotechnology information. Nucleic Acids Res. 2021;49:D10–D17. doi: 10.1093/nar/gkaa892. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 54.Katz K., Shutov O., Lapoint R., Kimelman M., Brister J.R., O’Sullivan C. The sequence read archive: a decade more of explosive growth. Nucleic Acids Res. 2022;50:D387–D390. doi: 10.1093/nar/gkab1053. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 55.Meyer F., Paarmann D., D’Souza M., Olson R., Glass E.M., Kubal M., et al. The metagenomics RAST server - a public resource for the automatic phylogenetic and functional analysis of metagenomes. BMC Bioinforma. 2008;9:386. doi: 10.1186/1471-2105-9-386. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 56.Arkin A.P., Cottingham R.W., Henry C.S., Harris N.L., Stevens R.L., Maslov S., et al. KBase: The United States department of energy systems biology knowledgebase. Nat Biotechnol. 2018;36:566–569. doi: 10.1038/nbt.4163. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 57.Chen I.-M.A., Chu K., Palaniappan K., Ratner A., Huang J., Huntemann M., et al. The IMG/M data management and analysis system v.7: content updates and new features. Nucleic Acids Res. 2023;51:D723–D732. doi: 10.1093/nar/gkac976. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 58.Richardson L., Allen B., Baldi G., Beracochea M., Bileschi M.L., Burdett T., et al. MGnify: the microbiome sequence data analysis resource in 2023. Nucleic Acids Res. 2023;51:D753–D759. doi: 10.1093/nar/gkac1080. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 59.Marcy Y., Ouverney C., Bik E.M., Lösekann T., Ivanova N., Martin H.G., et al. Dissecting biological “dark matter” with single-cell genetic analysis of rare and uncultivated TM7 microbes from the human mouth. Proc Natl Acad Sci USA. 2007;104:11889–11894. doi: 10.1073/pnas.0704662104. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 60.van der Walt A.J., van Goethem M.W., Ramond J.-B., Makhalanyane T.P., Reva O., Cowan D.A. Assembling metagenomes, one community at a time. BMC Genom. 2017:18. doi: 10.1186/s12864-017-3918-9. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 61.Mitchell A.L., Scheremetjew M., Denise H., Potter S., Tarkowska A., Qureshi M., et al. EBI Metagenomics in 2017: enriching the analysis of microbial communities, from sequence reads to assemblies. Nucleic Acids Res. 2018;46:D726–D735. doi: 10.1093/nar/gkx967. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 62.Lomsadze A., Gemayel K., Tang S., Borodovsky M. Modeling leaderless transcription and atypical genes results in more accurate gene prediction in prokaryotes. Genome Res. 2018;28:1079–1089. doi: 10.1101/gr.230615.117. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 63.Bland C., Ramsey T.L., Sabree F., Lowe M., Brown K., Kyrpides N.C., et al. CRISPR recognition tool (CRT): a tool for automatic detection of clustered regularly interspaced palindromic repeats. BMC Bioinforma. 2007;8:209. doi: 10.1186/1471-2105-8-209. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 64.Nawrocki E.P., Eddy S.R. Infernal 1.1: 100-fold faster RNA homology searches. Bioinformatics. 2013;29:2933–2935. doi: 10.1093/bioinformatics/btt509. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 65.Chan P.P., Lin B.Y., Mak A.J., Lowe T.M. tRNAscan-SE 2.0: improved detection and functional classification of transfer RNA genes. Nucleic Acids Res. 2021;49:9077–9096. doi: 10.1093/nar/gkab688. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 66.Mistry J., Chuguransky S., Williams L., Qureshi M., Salazar G.A., Sonnhammer E.L.L., et al. Pfam: the protein families database in 2021. Nucleic Acids Res. 2021;49:D412–D419. doi: 10.1093/nar/gkaa913. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 67.Haft D.H., Selengut J.D., Richter R.A., Harkins D., Basu M.K., Beck E. TIGRFAMs and genome properties in 2013. Nucleic Acids Res. 2013;41:D387–D395. doi: 10.1093/nar/gks1234. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 68.Sillitoe I., Dawson N., Lewis T.E., Das S., Lees J.G., Ashford P., et al. CATH: expanding the horizons of structure-based functional annotations for genome sequences. Nucleic Acids Res. 2019;47:D280–D284. doi: 10.1093/nar/gky1097. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 69.Pandurangan A.P., Stahlhacke J., Oates M.E., Smithers B., Gough J. The superfamily 2.0 database: a significant proteome update and a new webserver. Nucleic Acids Res. 2019;47:D490–D494. doi: 10.1093/nar/gky1130. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 70.Letunic I., Bork P. 20 years of the SMART protein domain annotation resource. Nucleic Acids Res. 2018;46:D493–D496. doi: 10.1093/nar/gkx922. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 71.Kanehisa M., Furumichi M., Sato Y., Ishiguro-Watanabe M., Tanabe M. KEGG: integrating viruses and cellular organisms. Nucleic Acids Res. 2021;49:D545–D551. doi: 10.1093/nar/gkaa970. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 72.Kiełbasa S.M., Wan R., Sato K., Horton P., Frith M.C. Adaptive seeds tame genomic sequence comparison. Genome Res. 2011;21:487–493. doi: 10.1101/gr.113985.110. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 73.Suzek B.E., Wang Y., Huang H., McGarvey P.B., Wu C.H. UniProt Consortium. UniRef clusters: a comprehensive and scalable alternative for improving sequence similarity searches. Bioinformatics. 2015;31:926–932. doi: 10.1093/bioinformatics/btu739. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 74.Kang D.D., Froula J., Egan R., Wang Z. MetaBAT, an efficient tool for accurately reconstructing single genomes from complex microbial communities. PeerJ. 2015;3 doi: 10.7717/peerj.1165. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 75.Chaumeil P.-A., Mussig A.J., Hugenholtz P., Parks D.H. GTDB-Tk: a toolkit to classify genomes with the genome taxonomy database. Bioinformatics. 2019 doi: 10.1093/bioinformatics/btz848. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 76.Rho M., Tang H., Ye Y. FragGeneScan: predicting genes in short and error-prone reads. Nucleic Acids Res. 2010;38 doi: 10.1093/nar/gkq747. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 77.Kalvari I., Argasinska J., Quinones-Olvera N., Nawrocki E.P., Rivas E., Eddy S.R., et al. Rfam 13.0: shifting to a genome-centric resource for non-coding RNA families. Nucleic Acids Res. 2018;46:D335–D342. doi: 10.1093/nar/gkx1038. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 78.Blum M., Chang H.-Y., Chuguransky S., Grego T., Kandasaamy S., Mitchell A., et al. The InterPro protein families and domains database: 20 years on. Nucleic Acids Res. 2021;49:D344–D354. doi: 10.1093/nar/gkaa977. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 79.Kanehisa M., Furumichi M., Tanabe M., Sato Y., Morishima K. KEGG: new perspectives on genomes, pathways, diseases and drugs. Nucleic Acids Res. 2017;45:D353–D361. doi: 10.1093/nar/gkw1092. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 80.Richardson L.J., Rawlings N.D., Salazar G.A., Almeida A., Haft D.R., Ducq G., et al. Genome properties in 2019: a new companion database to InterPro for the inference of complete functional attributes. Nucleic Acids Res. 2019;47:D564–D572. doi: 10.1093/nar/gky1013. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 81.Blin K., Wolf T., Chevrette M.G., Lu X., Schwalen C.J., Kautsar S.A., et al. antiSMASH 4.0-improvements in chemistry prediction and gene cluster boundary identification. Nucleic Acids Res. 2017;45:W36–W41. doi: 10.1093/nar/gkx319. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 82.Matias Rodrigues J.F., Schmidt T.S.B., Tackmann J., von Mering C. MAPseq: highly efficient k-mer search with confidence estimates, for rRNA sequence analysis. Bioinformatics. 2017;33:3808–3810. doi: 10.1093/bioinformatics/btx517. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 83.Quast C., Pruesse E., Yilmaz P., Gerken J., Schweer T., Yarza P., et al. The SILVA ribosomal RNA gene database project: improved data processing and web-based tools. Nucleic Acids Res. 2013;41:D590–D596. doi: 10.1093/nar/gks1219. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 84.Buchfink B., Reuter K., Drost H.-G. Sensitive protein alignments at tree-of-life scale using diamond. Nat Methods. 2021;18:366–368. doi: 10.1038/s41592-021-01101-x. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 85.Uritskiy G.V., DiRuggiero J., Taylor J. MetaWRAP—a flexible pipeline for genome-resolved metagenomic data analysis. Microbiome. 2018;6:1–13. doi: 10.1186/s40168-018-0541-1. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 86.Orakov A., Fullam A., Coelho L.P., Khedkar S., Szklarczyk D., Mende D.R., et al. GUNC: detection of chimerism and contamination in prokaryotic genomes. Genome Biol. 2021;22:178. doi: 10.1186/s13059-021-02393-0. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 87.Olm M.R., Brown C.T., Brooks B., Banfield J.F. dRep: a tool for fast and accurate genomic comparisons that enables improved genome recovery from metagenomes through de-replication. ISME J. 2017;11:2864–2868. doi: 10.1038/ismej.2017.126. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 88.Coelho L.P., Alves R., Del Río Á.R., Myers P.N., Cantalapiedra C.P., Giner-Lamia J., et al. Towards the biogeography of prokaryotic genes. Nature. 2022;601:252–256. doi: 10.1038/s41586-021-04233-4. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 89.Parks D.H., Chuvochina M., Rinke C., Mussig A.J., Chaumeil P.-A., Hugenholtz P. GTDB: an ongoing census of bacterial and archaeal diversity through a phylogenetically consistent, rank normalized and complete genome-based taxonomy. Nucleic Acids Res. 2022;50:D785–D794. doi: 10.1093/nar/gkab776. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 90.Nayfach S., Roux S., Seshadri R., Udwary D., Varghese N., Schulz F., et al. A genomic catalog of Earth’s microbiomes. Nat Biotechnol. 2021;39:499–509. doi: 10.1038/s41587-020-0718-6. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 91.Almeida A., Nayfach S., Boland M., Strozzi F., Beracochea M., Shi Z.J., et al. A unified catalog of 204,938 reference genomes from the human gut microbiome. Nat Biotechnol. 2021;39:105–114. doi: 10.1038/s41587-020-0603-3. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 92.Mende D.R., Letunic I., Maistrenko O.M., Schmidt T.S.B., Milanese A., Paoli L., et al. proGenomes2: an improved database for accurate and consistent habitat, taxonomic and functional annotations of prokaryotic genomes. Nucleic Acids Res. 2020;48:D621–D625. doi: 10.1093/nar/gkz1002. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 93.Nishimura Y., Yoshizawa S. The OceanDNA MAG catalog contains over 50,000 prokaryotic genomes originated from various marine environments. Sci Data. 2022;9:305. doi: 10.1038/s41597-022-01392-5. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 94.Fullam A., Letunic I., Schmidt T.S.B., Ducarmon Q.R., Karcher N., Khedkar S., et al. proGenomes3: approaching one million accurately and consistently annotated high-quality prokaryotic genomes. Nucleic Acids Res. 2023;51:D760–D766. doi: 10.1093/nar/gkac1078. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 95.Rinke C., Lee J., Nath N., Goudeau D., Thompson B., Poulton N., et al. Obtaining genomes from uncultivated environmental microorganisms using FACS-based single-cell genomics. Nat Protoc. 2014;9:1038–1048. doi: 10.1038/nprot.2014.067. [DOI] [PubMed] [Google Scholar]
- 96.Nishikawa Y., Hosokawa M., Maruyama T., Yamagishi K., Mori T., Takeyama H. Monodisperse picoliter droplets for low-bias and contamination-free reactions in single-cell whole genome amplification. PLoS One. 2015;10 doi: 10.1371/journal.pone.0138733. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 97.Lasken R.S., Stockwell T.B. Mechanism of chimera formation during the multiple displacement amplification reaction. BMC Biotechnol. 2007;7:19. doi: 10.1186/1472-6750-7-19. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 98.Ide K., Nishikawa Y., Maruyama T., Tsukada Y., Kogawa M., Takeda H., et al. Targeted single-cell genomics reveals novel host adaptation strategies of the symbiotic bacteria Endozoicomonas in Acropora tenuis coral. Microbiome. 2022;10:220. doi: 10.1186/s40168-022-01395-9. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 99.Nishikawa Y., Kogawa M., Hosokawa M., Wagatsuma R., Mineta K., Takahashi K., et al. Validation of the application of gel beads-based single-cell genome sequencing platform to soil and seawater. ISME Commun. 2022;2:1–11. doi: 10.1038/s43705-022-00179-4. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 100.Ide K., Saeki T., Arikawa K., Yoda T., Endoh T., Matsuhashi A., et al. Exploring strain diversity of dominant human skin bacterial species using single-cell genome sequencing. Front Microbiol. 2022;13 doi: 10.3389/fmicb.2022.955404. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 101.Kogawa M., Nishikawa Y., Saeki T., Yoda T., Arikawa K., Takeyama H., et al. Revealing within-species diversity in uncultured human gut bacteria with single-cell long-read sequencing. Front Microbiol. 2023;14:1133917. doi: 10.3389/fmicb.2023.1133917. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 102.Roux S., Hawley A.K., Torres Beltran M., Scofield M., Schwientek P., Stepanauskas R., et al. Ecology and evolution of viruses infecting uncultivated SUP05 bacteria as revealed by single-cell- and meta-genomics. Elife. 2014;3 doi: 10.7554/eLife.03125. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 103.Nobu M.K., Narihiro T., Rinke C., Kamagata Y., Tringe S.G., Woyke T., et al. Microbial dark matter ecogenomics reveals complex synergistic networks in a methanogenic bioreactor. ISME J. 2015;9:1710–1722. doi: 10.1038/ismej.2014.256. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 104.Mende D.R., Aylward F.O., Eppley J.M., Nielsen T.N., DeLong E.F. Improved environmental genomes via integration of metagenomic and single-cell assemblies. Front Microbiol. 2016;7:143. doi: 10.3389/fmicb.2016.00143. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 105.Yuan C., Lei J., Cole J., Sun Y. Reconstructing 16S rRNA genes in metagenomic data. Bioinformatics. 2015;31:i35–i43. doi: 10.1093/bioinformatics/btv231. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 106.Sutcliffe I.C., Rosselló-Móra R., Trujillo M.E. Addressing the sublime scale of the microbial world: reconciling an appreciation of microbial diversity with the need to describe species. New Microbes New Infect. 2021;43 doi: 10.1016/j.nmni.2021.100931. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 107.Hedlund B.P., Chuvochina M., Hugenholtz P., Konstantinidis K.T., Murray A.E., Palmer M., et al. SeqCode: a nomenclatural code for prokaryotes described from sequence data. Nat Microbiol. 2022;7:1702–1708. doi: 10.1038/s41564-022-01214-9. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 108.Almeida A., Mitchell A.L., Boland M., Forster S.C., Gloor G.B., Tarkowska A., et al. A new genomic blueprint of the human gut microbiota. Nature. 2019;568:499–504. doi: 10.1038/s41586-019-0965-1. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 109.Forster S.C., Kumar N., Anonye B.O., Almeida A., Viciani E., Stares M.D., et al. A human gut bacterial genome and culture collection for improved metagenomic analyses. Nat Biotechnol. 2019;37:186–192. doi: 10.1038/s41587-018-0009-7. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 110.Hiseni P., Rudi K., Wilson R.C., Hegge F.T., Snipen L. HumGut: a comprehensive human gut prokaryotic genomes collection filtered by metagenome data. Microbiome. 2021;9:165. doi: 10.1186/s40168-021-01114-w. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 111.Xie F., Jin W., Si H., Yuan Y., Tao Y., Liu J., et al. An integrated gene catalog and over 10,000 metagenome-assembled genomes from the gastrointestinal microbiome of ruminants. Microbiome. 2021;9:137. doi: 10.1186/s40168-021-01078-x. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 112.Chen C., Zhou Y., Fu H., Xiong X., Fang S., Jiang H., et al. Expanded catalog of microbial genes and metagenome-assembled genomes from the pig gut microbiome. Nat Commun. 2021;12:1106. doi: 10.1038/s41467-021-21295-0. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 113.Dai D., Zhu J., Sun C., Li M., Liu J., Wu S., et al. GMrepo v2: a curated human gut microbiome database with special focus on disease markers and cross-dataset comparison. Nucleic Acids Res. 2022;50:D777–D784. doi: 10.1093/nar/gkab1019. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 114.Zeng S., Patangia D., Almeida A., Zhou Z., Mu D., Paul Ross R., et al. A compendium of 32,277 metagenome-assembled genomes and over 80 million genes from the early-life human gut microbiome. Nat Commun. 2022;13:5139. doi: 10.1038/s41467-022-32805-z. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 115.Hannigan G.D., Prihoda D., Palicka A., Soukup J., Klempir O., Rampula L., et al. A deep learning genome-mining strategy for biosynthetic gene cluster prediction. Nucleic Acids Res. 2019;47 doi: 10.1093/nar/gkz654. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 116.Carroll L.M., Larralde M., Fleck J.S., Ponnudurai R., Milanese A., Cappio E., et al. Accurate de novo identification of biosynthetic gene clusters with GECCO. BioRxiv. 2021 doi: 10.1101/2021.05.03.442509. 2021.05.03.442509. [DOI] [Google Scholar]
- 117.van Kempen M., Kim S.S., Tumescheit C., Mirdita M., Lee J., Gilchrist C.L.M., et al. Fast and accurate protein structure search with Foldseek. Nat Biotechnol. 2023 doi: 10.1038/s41587-023-01773-0. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 118.Yu T., Cui H., Li J.C., Luo Y., Jiang G., Zhao H. Enzyme function prediction using contrastive learning. Science. 2023;379:1358–1363. doi: 10.1126/science.adf2465. [DOI] [PubMed] [Google Scholar]
- 119.Jumper J., Evans R., Pritzel A., Green T., Figurnov M., Ronneberger O., et al. Highly accurate protein structure prediction with AlphaFold. Nature. 2021;596:583–589. doi: 10.1038/s41586-021-03819-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 120.Lin Z., Akin H., Rao R., Hie B., Zhu Z., Lu W., et al. Evolutionary-scale prediction of atomic-level protein structure with a language model. Science. 2023;379:1123–1130. doi: 10.1126/science.ade2574. [DOI] [PubMed] [Google Scholar]
- 121.Bickhart D.M., Kolmogorov M., Tseng E., Portik D.M., Korobeynikov A., Tolstoganov I., et al. Generating lineage-resolved, complete metagenome-assembled genomes from complex microbial communities. Nat Biotechnol. 2022;40:711–719. doi: 10.1038/s41587-021-01130-z. [DOI] [PubMed] [Google Scholar]
- 122.Feng X., Cheng H., Portik D., Li H. Metagenome assembly of high-fidelity long reads with hifiasm-meta. Nat Methods. 2022;19:671–674. doi: 10.1038/s41592-022-01478-3. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 123.Kim C.Y., Ma J., Lee I. HiFi metagenomic sequencing enables assembly of accurate and complete genomes from human gut microbiota. Nat Commun. 2022;13:6367. doi: 10.1038/s41467-022-34149-0. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 124.Moss E.L., Maghini D.G., Bhatt A.S. Complete, closed bacterial genomes from microbiomes using nanopore sequencing. Nat Biotechnol. 2020;38:701–707. doi: 10.1038/s41587-020-0422-6. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 125.Ciuffreda L., Rodríguez-Pérez H., Flores C. Nanopore sequencing and its application to the study of microbial communities. Comput Struct Biotechnol J. 2021;19:1497–1511. doi: 10.1016/j.csbj.2021.02.020. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 126.Liu L., Yang Y., Deng Y., Zhang T. Nanopore long-read-only metagenomics enables complete and high-quality genome reconstruction from mock and complex metagenomes. Microbiome. 2022;10:209. doi: 10.1186/s40168-022-01415-8. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 127.Orellana L.H., Krüger K., Sidhu C., Amann R. Comparing genomes recovered from time-series metagenomes using long- and short-read sequencing technologies. Microbiome. 2023;11:105. doi: 10.1186/s40168-023-01557-3. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 128.Pan S., Zhao X.-M., Coelho L.P. SemiBin2: self-supervised contrastive learning leads to better MAGs for short- and long-read sequencing. BioRxiv. 2023 doi: 10.1101/2023.01.09.523201. 2023.01.09.523201. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 129.Yaffe E., Relman D.A. Tracking microbial evolution in the human gut using Hi-C reveals extensive horizontal gene transfer, persistence and adaptation. Nat Microbiol. 2020;5:343–353. doi: 10.1038/s41564-019-0625-0. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 130.Du Y., Sun F. HiCBin: binning metagenomic contigs and recovering metagenome-assembled genomes using Hi-C contact maps. Genome Biol. 2022;23:63. doi: 10.1186/s13059-022-02626-w. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 131.Barrett T., Clark K., Gevorgyan R., Gorelenkov V., Gribov E., Karsch-Mizrachi I., et al. BioProject and BioSample databases at NCBI: facilitating capture and organization of metadata. Nucleic Acids Res. 2012;40:D57–D63. doi: 10.1093/nar/gkr1163. [DOI] [PMC free article] [PubMed] [Google Scholar]




