Abstract
In 2025, the bacterial diversity database BacDive is the leading database for strain-level bacterial and archaeal information. It has been selected as an ELIXIR Core Data Resource as well as a Global Core Biodata Resource. Since its initial release more than ten years ago, BacDive (https://bacdive.dsmz.de) has grown tremendously in content and functionalities, and is a comprehensive resource covering the phenotypic diversity of prokaryotes with data on taxonomy, morphology, physiology, cultivation, and more. The current release (2023.2) contains 2.6 million data points on 97 334 strains, reflecting an increase by 52% since the previous publication in 2021. This remarkable growth can largely be attributed to the integration of the world-wide largest collection of Analytical Profile Index (API) test results, which are now fully integrated into the database and searchable. A novel BacDive knowledge graph provides powerful search options through a SPARQL endpoint, including the possibility for federated searches across multiple data sources. The high-quality data provided by BacDive is increasingly being used for the training of artificial intelligence models and resulting genome-based predictions with high confidence are now used to fill content gaps in the database.
Graphical Abstract
Graphical Abstract.
Introduction
BacDive is the largest database for standardized information on prokaryotes in the world. It gathers and standardizes phenotypic strain-level research data from diverse sources including internal catalogs of culture collections and primary literature to make them easily accessible. These efforts were recently recognized by the Global Biodata Coalition and the ELIXIR European life science infrastructure, who both awarded BacDive the distinction of a Core Data Resource.
A look at recent publications using BacDive data gives a good overview on the broad research community that relies on this resource. Applications range from biotechnology, using BacDive to find bacteria for the degradation of toxic waste (1) or plastics (2) to the medical research identifying bacteria linked to oral health (3). They also include research investigating the evolutionary metabolic adaptation of specific genera (4) as well as plant health, benefiting from a group of plant-associated flavobacteria (5). A growing trend that can be identified is the use of large data sets in bioinformatic analyses and tools. Datasets that serve specific research applications are developed with the help of BacDive data, like Omnicrobe (6) focusing on habitat information, or MariClus, a dataset dedicated to marine natural products (7). Very exciting are those approaches that take advantage of the great potential of the extensive, standardized data in BacDive to make predictions. TemBerture (8) has predicted protein thermostability based on deep learning techniques and Barnum et al. (9) have predicted microbial growth conditions based on amino acid composition. Surely, well-designed models relying on high-quality data have the potential to close the gaps of knowledge for the already known microbial diversity and to improve the understanding of the so far unknown microbial dark matter. Here, we describe the recent developments in the Core data resource BacDive and how we utilize genome-based models to significantly improve the knowledge about prokaryotic strains and fill data gaps in BacDive.
Primary content
BacDive now contains data on 97 334 strains with a total of 2.6 million data points (Release 2024.1). This represents a growth of 52% since our last report in 2021. In its current release BacDive shows data on 20 060 type strains, covering 98% of the 20 510 validly described species (https://lpsn.dsmz.de/text/numbers). The increase in type strain coverage in BacDive from just 81% in 2021 to 98% in 2024 could be achieved through improved data exchange with one of its sister databases, the List of Prokaryotic names with Standing in Nomenclature (LPSN) (10). Data from LPSN is no longer only used to keep BacDive up to date on correct prokaryotic nomenclature, but also to add all type strains of newly described species to the database in a timely manner. Complementing this approach, extensive data from species descriptions continue to be annotated manually and integrated into BacDive. To date, data on 7893 strains has been added from species descriptions. An overview on all strains providing manually annotated literature data is given under the special collection ‘strains from literature’ (https://bacdive.dsmz.de/collection/literature).
The database was originally conceived to make available the vast amount of data on strains collected and stored in internal files at the German Collection of Microorganisms and Cell Cultures (Leibniz Institute DSMZ). Until today, high-quality data from large culture collections forms the backbone of the database. Apart from regular updates with DSMZ data, a large data set on 13 140 strains from the Biological Resource Center of the Institut Pasteur (CRBIP, France) could recently be integrated. While only 3390 strains were completely new to BacDive, the data set significantly extended the BacDive database with comprehensive information per strain. It contained not only the public catalog data on identity, sampling, and history of the strain, but also internal observations and test results on cultivation, growth conditions, morphology, physiology and metabolism. A notable highlight is the large set of Analytical Profile Index (API) test results for 9783 strains.
API tests are one of the underestimated resources for systematic information about microbial strains. Developed in the 1970s, they are simple sequences of microsize physiological tests that can easily be performed in any laboratory without the need of expensive devices. Until today, these tests are routinely used in medical laboratories as well as in culture collections. In 2017, a first batch of 8977 API test results was mobilized and integrated into the BacDive database, which represented the world's largest publicly available collection of API tests at the time (11). Today the BacDive collection of API test results encompasses a total of 48 130 tests, divided into 17 test types, which provide overall 1 594 078 single data points for 24 112 strains.
Every single data point represents a physiological test, like the enzymatic conversion of a metabolite, the resistance to an antibiotic substance, spore formation, motility or Gram staining. Initially, they were stored as API test data in the database and as such were not easily searchable and accessible. Therefore, all API test data fields were carefully evaluated and for those for which equivalent data fields were available in the BacDive database, data were transformed and incorporated into the database. Thus, the data is now directly accessible side-by-side with the curated data. Unclear results were excluded during the transformation process. In total, the transformations added new datasets on enzymes, motility, hemolysis, culture media, antibiotic sensitivity, metabolite production and utilization, spore formation, and the response to Voges-Proskauer and Indole tests. To distinguish the API test data from manually curated data, every single data point is clearly marked and shows a direct link to the original API test dataset (Figure 1A). Automatically transcribed data can be hidden from the strain detail view through the ‘exclude non-curated data’ function.
Figure 1.
New features on the BacDive strain detail page. (A) Analytical Profile Index (API) test results are now not only shown in specific API tables (bottom), but also integrated alongside manually curated data. (B) Geographic locations and types of very closely related samples with at least 99% 16S rRNA gene sequence identity are shown as provided by Microbeatlas. (C) The sidebar now links to related StrainInfo and PhageDive entries as well as BacDive Special Collections that the strain can be found in. (D) Genome-based predictions with over 90% confidence can be found integrated along experimental data in the respective sections, clearly marked with an AI icon. (E) All predicted data are listed in a new section along with the relevant information. (F) A new literature section lists publications metadata retrieved from PubMed and automatically matched to the strain via culture collection numbers and taxonomy.
What makes this data set a treasure for microbial strain research is the largely unbiased and standardized availability of the data. So far, the world of microbial strain data is dominated by comprehensive information about type strains (buried in species descriptions). As API test data are produced routinely (e.g. for quality checks) these data provide standardized information for a large number of non-type strains for which only little data would be available otherwise. The BacDive API test data set, which includes physiological data on 18 053 non-type strains, is closing an enormous knowledge gap, which is particularly important for the identification of strains for new applications.
Linked content
Besides integrating and standardizing data, a strength of BacDive is to connect strains to related data from other resources by providing web links. On the one hand, this allows users to seamlessly explore strain data throughout different high-quality sources. On the other hand, it allows researchers to connect data via stable identifiers and generate new data sets specific to their needs. This is further facilitated by the knowledge graph described below.
Sequence information is important in many areas of microbial research, including identification, taxonomic classification, comparative and diversity analyses and phenotype–genotype comparisons. Since the beginning, enabling users to easily connect standardized phenotypic data with genomic sequence data to provide genome–phenotype inference has been of high priority to BacDive. BacDive collects accession numbers for genome assemblies and 16S rRNA genes, which are also commonly used for phylogenetic reconstructions and taxonomy, from NCBI GenBank (12), JGI IMG (13) and BV-BRC/PATRIC (14). During the past two years, the sequence data in BacDive has been cleaned up and widely extended. Metadata for all GenBank accession numbers in the database were newly downloaded and compared against the BacDive strain they were matched to, so that wrongly matched sequences could be deleted. In addition, all GenBank nucleotide accessions for contigs, chromosomes or plasmids were consolidated into the respective assemblies. Accessions for GenBank genome assemblies and 16S rRNA gene sequences are now routinely collected by going through current accession lists downloaded from the NCBI FTP server and programmatically matching them to BacDive strains via culture collection numbers and species names, also considering all synonyms known in LPSN and the NCBI taxonomy (15). BacDive now links to 50 588 genome assemblies and 41 458 16S rRNA gene sequences.
Peer-reviewed journal articles are still the most important medium for the communication of research findings. Primary literature therefore stores deep knowledge about prokaryotic strains. While we endeavor to extract standardized information from the literature, the diversity of data is too great to cover everything in a standardized form in the foreseeable future. However, the number and frequency of new publications is constantly increasing, making it very difficult to keep up with new findings. In order to provide an overview about the currently available publications related to a specific strain, we introduced a new literature table within the section ‘External links’ (Figure 1F). This allows researchers to easily find further material on a strain of interest. The literature is gathered automatically by collecting all species names from BacDive, downloading metadata and abstracts for all publications for this species from PubMed (16) and finally screening them for all culture collections numbers linked to BacDive strains of the species. In this way, 59 282 publications for 19 698 prokaryotic strains could be added to the database. For a better overview, the publications are categorized into one of the following topics: Biotechnology (764 publications), Cultivation (419), Enzymology (4532), Genetics (3151), Lipids (5), Metabolism (7139), Pathogenicity (6262), Phenotype (88), Phylogeny (31 007), Physiology (10), Proteome (81), Stress (223) and Transcriptome (161). To assign publications to topics, keywords are extracted from publication titles and MeSH (Medical Subject Headings) terms, excluding stop words and using a lemmatizer. The keywords are then compared against a custom dictionary for assignment to categories. The publications are linked to the original publication via Digital Object Identifier (DOI) and the metadata is available for download in BibTeX format.
For understanding the ecological role of a microbial strain, its isolation source is of major importance. Naturally, a strain is only collected once and therefore only provides a single point of evidence for the potential habitat of the species, which might not be representative. To provide a better overview about occurrences in the environment, we matched the 16S rRNA gene sequence data with data from the MicrobeAtlas database (https://microbeatlas.org/) (17). For 15 655 strains matches with at least 99% 16S rRNA gene sequence identity were found. For each strain, Taxonmaps provide a preview of the global distribution. They were integrated into the strain detail view alongside data about the environmental categorization into aquatic, soil, animal and plant (Figure 1B). By clicking on the map, users are redirected to the respective MicrobeAtlas entry providing a zoomable map, as well as further comprehensive information.
Since January 2023, BacDive is part of the newly built biodata infrastructure DSMZ Digital Diversity (https://hub.dsmz.de), which aims to establish an integrated suite of scientific databases of fundamental relevance for the life sciences. The databases not only include the established Core data resources BRENDA (18), SILVA (19) and LPSN (10), but also newly established databases like StrainInfo (https://straininfo.dsmz.de), MediaDive (20) and PhageDive (21). All resources are developed in a coordinated manner and benefit from the frequent exchange of data. Stable links provide users with easy access to all resources. Linking nomenclatural data to LPSN, enzymatic data to BRENDA and 16S rRNA gene sequence data to SILVA has been established for many years already. Recently, we added links to entries for cultivation media in MediaDive, which provides detailed instructions for cultivation media recipes as well as comfortable functions to find alternative media or to build own media recipes. StrainInfo is a reestablished database, dedicated to strain identity information and provides additional information on the associated cultures and their relations (Figure 1C). PhageDive provides comprehensive information on prokaryotic viruses that are able to infect the respective strain (Figure 1C). While linking is important, linked data are not easily queried and analyzed. Therefore, the goal of DSMZ Digital Diversity is to provide access to integrated data from multiple sources through a central hub to enable scientists to gain new insights that are currently hidden. For this reason, BacDive strains can now also be found via an integrated search through the Digital Diversity Hub (https://hub.dsmz.de/#/search/).
Searching and querying content
As BacDive is a knowledge database, its tools are focused on searching and retrieving data. The most-frequently used tool is the Simple Search (starting page or top of each page) which allows users to easily find a strain based on its name, culture collection number, NCBI Tax-ID or sequence accession number. Lately this search function was supplemented with a Smart Search. When a user starts typing into the search bar, the Smart Search function makes suggestions based on almost 300 000 precalculated Advanced Search queries, for example recognizing search terms like ‘glucose’ and offering various search options. This bridges the gap between the Simple Search and the Advanced Search and therefore provides a low-level entry for users into the more complex search options of BacDive.
Another new query tool in BacDive are so-called Special Collections. These group together strains that belong to specific research projects or collections, and therefore are of special interest to the user. To provide better access and visibility to these otherwise hidden data sets, these strains are displayed on dedicated Special Collections pages and are also query- and retrievable as a whole. An example is the Mouse Microbiome collection, or miBC (mouse intestinal Bacterial Collection), which contains strains that were isolated from the intestines of mice in a joint initiative of the RWTH Aachen and the Leibniz Institute DSMZ and deposited in the DSMZ collection to make them available for research purposes (22). Other examples are the ESA collection of strains isolated from spacecrafts and assembly facilities of the European Space Agency (ESA), and special collections containing all strains that are available through a specific Biological Resource Center. A new Special Collections landing page gives an overview on the available collections and provides links to the individual collection pages, where a short description and overview of the collection can be found (Figure 2).
Figure 2.
Special collection dashboard. A short description, list of strains and selected statistics give an overview of a special collection. The sidebar on the top right allows for easy switching between different collections.
At the top of these pages, the respective special collection is shortly described. Below, tiles contain key statistical values: the total number of strains, the total number of species, and if available, mean values for cell length and width, the cultivation temperature optimum, the cultivation pH optimum, mean incubation time, mean cultivation salt optimum and number of 16S rRNA gene accession numbers connected to the strains. These are followed by a list of all strains in the Special Collection, displayed as a searchable table with BacDive IDs, type strain status, species name, NCBI Taxonomy-ID, and culture collection numbers. For further exploration of the collection, a link to the BacDive Advanced Search with pre-selection for the Special Collection is provided. This enables users to further query the collection and easily download data for all or selected strains of the group. Additional statistics are displayed on the bottom of the page in interactively displayed graphs or charts for different aspects of taxonomy, morphology, cultivation, metabolism, isolation sources, geographical distribution, molecular biology and pathogenicity, and physiological tests. Only graphs for which data are available are displayed. Clicking on a data point in any of the graphs opens the respective search in the Advanced Search, Isolation Source Search or TAXplorer, which allows the user to explore the filtered data set further.
Semantic integration of content: the BacDive knowledge graph
As a knowledge base, the mission of BacDive is to provide access to high-quality, standardized knowledge about prokaryotic strains. The tools provided to search and analyze the data are either limited in power (e.g. Advanced Search) or provide a full set of BacDive data (e.g. Web services) that the user still needs to query with custom tools. Here we present a knowledge graph of the BacDive database that contains >16.5 million triples and provides a powerful SPARQL endpoint to directly search and analyze the knowledge provided through BacDive in a standardized way, supported by a new descriptive ontology. Furthermore, the BacDive knowledge graph will support the integration approaches for the DSMZ Digital Diversity infrastructure previously mentioned.
The Resource Description Framework (RDF) and the RDF Query Language (SPARQL) represent foundational technologies in the realm of semantic web and linked data (23). RDF is a standard model for data interchange on the web, enabling the integration of diverse data sources with a flexible and extensible approach. SPARQL, the query language for RDF, allows for sophisticated queries directly against the data model, enabling more precise and tailored data retrieval.
The new BacDive SPARQL endpoint uses the QLever query engine (24) and can be found at https://sparql.dsmz.de/bacdive. The BacDive knowledge graph includes detailed mapping rules for so far 26 of the most critical entities, ensuring that the most relevant data is accurately represented and connected. These entities include: Strain, Reference, StrainDesignation, CultureCollectionNumber, NutritionType, GramStain, CellMotility, CellSize, CellShape, CultureMedium, CultureTemperature, CulturePH, OxygenTolerance, LocationOfOrigin, IsolationSource, EnrichmentProcedure, ColonyMorphology, SaltTolerance, SporeFormation, RiskAssessment, Pathogenicity, Enzyme, CellPigmentation, 16SSequence, GenomeSequence, GCContent. In total, the graph contains >16.5 million triples. By focusing on these entities, we have laid a robust foundation for the comprehensive integration of BacDive data into broader knowledge frameworks, which can be further expanded in the future.
This extensive RDF dataset not only supports complex queries via SPARQL but also enables federated queries (25), which span multiple data sources. For example, microbial data from BacDive can be seamlessly integrated with protein information from UniProt (26). The potential to conduct such federated queries greatly expands the scope of scientific inquiry, allowing researchers to derive insights from a more holistic dataset than would be possible from isolated sources. A demonstration of this functionality is provided in a sample query, which retrieves all chemolithoautotrophic strains from BacDive along with their corresponding protein sequences from UniProt (Figure 3). A key point of integration is the NCBI Taxonomy-ID, directly linked to BacDive strains with the hasTaxID predicate. To ensure compatibility with other DSMZ databases, the knowledge graph adheres to the DSMZ Digital Diversity Ontology (D3O; https://bioportal.bioontology.org/ontologies/D3O). This allows users to query BacDive data effortlessly and integrate it with other DSMZ endpoints, such as MediaDive.
Figure 3.
Interface to the SPARQL end point with an example of a federated SPARQL query executed on the BacDive database. The query dynamically constructs URLs for taxonomy IDs from BacDive to fetch relevant protein data, showcasing the integration of microbiological and protein sequence data through a federated query approach. The results display strains, their names, taxonomy IDs, associated proteins, and partial amino acid sequences, demonstrating the powerful capability of federated queries in combining data from disparate sources. The visible result is limited to 1000 entries, but the full set can be downloaded.
We have included a variety of example SPARQL queries that demonstrate the powerful capabilities of querying the BacDive knowledge graphs. These will be continually expanded with more complex and varied queries to cater to the diverse needs of researchers. The deployment of the BacDive knowledge graph with its SPARQL endpoint is an important step towards more interconnected microbiological datasets.
Filling content gaps
While the number of data points in BacDive rises constantly, this is mainly due to the integration of new strains. Data availability for each strain varies widely and unfortunately large gaps remain in the phenotypic descriptions of many strains. This even concerns such fundamental information as the Gram staining behavior or oxygen tolerance of the strain. In order to decrease some of these profound gaps, we have recently introduced genome-based predictions from two different projects, DiASPora (https://diaspora-project.de) and deepG (https://deepg.de/), into BacDive.
The DiASPora procedure for producing strain-level phenotype predictions using public genome sequences and high-quality standardized data from BacDive is described in detail by Koblitz et al. (27). In short, machine learning models for several traits were trained on curated BacDive data and Pfam (28) annotated genomes using the Random Forest algorithm (29). Six models performed well enough to be used to generate data for integration into the database. For 15 938 strains with high-quality genome assemblies, new data points could be predicted for flagellated motility, Gram stain, oxygen tolerance (models for aerobe and anaerobe growth), spore-formation and growth temperature range (prediction of thermophilic growth). A second set of machine learning data was integrated into BacDive that was created using a deep learning model trained using the deepG platform (30). It provides predictions for the spore-formation ability of 9023 strains.
All predicted data are listed in a new section on the strain detail page titled ‘Genome-based predictions’ with confidence values and the link to the genome the predictions are based on. Predicted data points with over 90% confidence are additionally integrated into the other sections alongside experimental data, marked with their confidence value and an AI icon. Overall, 104 651 data points were generated by genome-based predictions covering 16 131 strains. 15 977 of these data points have added completely new information where no experimental data is available (Figure 4).
Figure 4.
Data points added by predictions. Number of BacDive strains with data for the traits oxygen tolerance, Gram stain, motility and spore formation. Dark blue sections show numbers of strains for which only experimental values are present in the database, striped sections visualize strains for which predictions were integrated next to existing experimental data and light blue sections represent strains for which integrated predictions added information where none was previously present.
With 16 131 strains for which prediction data is available, only 17% of the strains in BacDive are covered. Nevertheless, 12 718 strains represent type strains, and thereby 62% of the currently validly described prokaryotic species are represented. The major limitation is the availability of high-quality genomes for less well-described strains. With the progress in sequencing technology, the importance of interpreting genome functions by machine learning models will rise over the next years. BacDive is an ideal platform to integrate data resulting from genome-based predictions, as high-quality predictions can fill knowledge gaps for not well-described strains and challenge existing data to improve the data quality. Moreover, by applying published models and integrating the data generated by them, BacDive offers a great way to make them more findable, accessible, interoperable, and reusable (FAIR).
Acknowledgements
We wish to express our thanks to all collaborating scientists involved in the data annotation. Special thanks to Martin Boutroux, Joao Frederico Matias Rodrigues and Christian von Mering for providing comprehensive data sets. We would also like to thank the Collection de l’Institut Pasteur for the friendly permission for publishing internal data. We appreciate the opportunity provided by BioHackathon Japan 2024 to advance our work on the BacDive Knowledge Graph and would like to thank Jerven Bolleman and Daniel Fernández-Álvarez for their guidance on RDF best practices.
Author contributions: I.S.: Data curation, Software, Writing – original draft; J.K.: Methodology, Visualization, Software, Writing – review & editing; J.S.C., C.E., M.S., A.P., R.G., V.I., J.C.: Software; J.O.: Conceptualization, Funding acquisition, Writing – review & editing; L.R.: Conceptualization, Data curation, Writing – original draft.
Contributor Information
Isabel Schober, Leibniz Institute DSMZ - German Collection of Microorganisms and Cell Cultures, Braunschweig, Germany.
Julia Koblitz, Leibniz Institute DSMZ - German Collection of Microorganisms and Cell Cultures, Braunschweig, Germany.
Joaquim Sardà Carbasse, Leibniz Institute DSMZ - German Collection of Microorganisms and Cell Cultures, Braunschweig, Germany.
Christian Ebeling, Leibniz Institute DSMZ - German Collection of Microorganisms and Cell Cultures, Braunschweig, Germany.
Marvin Leon Schmidt, Leibniz Institute DSMZ - German Collection of Microorganisms and Cell Cultures, Braunschweig, Germany.
Adam Podstawka, Leibniz Institute DSMZ - German Collection of Microorganisms and Cell Cultures, Braunschweig, Germany.
Rohit Gupta, German National Library of Science and Technology (TIB) - Leibniz Information Centre for Science and Technology - University Library, Hannover, Germany.
Vinodh Ilangovan, German National Library of Science and Technology (TIB) - Leibniz Information Centre for Science and Technology - University Library, Hannover, Germany.
Javad Chamanara, German National Library of Science and Technology (TIB) - Leibniz Information Centre for Science and Technology - University Library, Hannover, Germany.
Jörg Overmann, Leibniz Institute DSMZ - German Collection of Microorganisms and Cell Cultures, Braunschweig, Germany.
Lorenz Christian Reimer, Leibniz Institute DSMZ - German Collection of Microorganisms and Cell Cultures, Braunschweig, Germany.
Data availability
BacDive data can be freely downloaded in various formats (e.g. CSV, JSON, PDF) without restrictions, except that the origin of the data has to be properly cited when used in other works (CC BY 4.0 license). Registration is necessary to access data through the RESTful API, but registration is free of charge. The AI models from the DiASPora project can be found at https://doi.org/10.5281/zenodo.13757323.
Funding
Federal Ministry of Education and Research (BMBF) [de.NBI 021A539C to J.O.]; Leibniz Association [SAW project DiASPora, Funding No. K280/2019]; German Centre for Infection Research (DZIF) [8005512901, 8005512001]; Deutsche Forschungsgemeinschaft (DFG, German Research Foundation) [NFDI4Biodiversity; project number 442032008; NFDI 5/1, NFDI4Microbiota; project number 460129525; NFDI 28/1]. Funding for open access charge: Institutional funding.
Conflict of interest statement. None declared.
References
- 1. Palma T.L., Costa M.C.. Biodegradation of 17α-ethinylestradiol by strains of Aeromonas genus isolated from acid mine drainage. Clean Technol. 2024; 6:116–139. [Google Scholar]
- 2. Kapoor A., Varshney C.. Microbial degradation of PET plastic sustainably yielding commercially viable products. 2021; Preprints doi:21 June 2021, pre-print: not peer-reviewed 10.20944/preprints202106.0519.v1. [DOI]
- 3. Chopra A., Franco-Duarte R., Rajagopal A., Choowong P., Soares P., Rito T., Eberhard J., Jayasinghe T.N.. Exploring the presence of oral bacteria in non-oral sites of patients with cardiovascular diseases using whole metagenomic data. Sci. Rep. 2024; 14:1476. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 4. da Silva Santos D., Freitas N.S.A., de Morais M.A., Mendonça A.A.. Liquorilactobacillus: a context of the evolutionary history and metabolic adaptation of a bacterial genus from fermentation liquid environments. J. Mol. Evol. 2024; 92:467–487. [DOI] [PubMed] [Google Scholar]
- 5. Seo H., Kim J.H., Lee S.-M., Lee S.-W.. The plant-associated Flavobacterium: a hidden helper for improving plant health. Plant Pathol J. 2024; 40:251–260. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 6. Dérozier S., Bossy R., Deléger L., Ba M., Chaix E., Harlé O., Loux V., Falentin H., Nédellec C.. Omnicrobe, an open-access database of microbial habitats and phenotypes using a comprehensive text mining and data fusion approach. PLoS One. 2023; 18:e0272473. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 7. Hermans C., De Mol M.L., Mispelaere M., De Rop A.-S., Rombaut J., Nusayr T., Creamer R., De Maeseneire S.L., Soetaert W.K., Hulpiau P.. MariClus: your one-stop platform for information on marine natural products, their gene clusters and producing organisms. Mar. Drugs. 2023; 21:449. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 8. Rodella C., Lazaridi S., Lemmin T.. TemBERTure: advancing protein thermostability prediction with deep learning and attention mechanisms. Bioinformatics Advances. 2024; 4:vbae103. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 9. Barnum T.P., Crits-Christoph A., Molla M., Carini P., Lee H.H., Ostrov N.. Predicting microbial growth conditions from amino acid composition. 2024; bioRxiv doi:22 March 2024, pre-print: not peer-reviewed 10.1101/2024.03.22.586313. [DOI]
- 10. Parte A.C., Sardà Carbasse J., Meier-Kolthoff J.P., Reimer L.C., Göker M.. List of Prokaryotic names with Standing in Nomenclature (LPSN) moves to the DSMZ. Int. J. Syst. Evol. Microbiol. 2020; 70:5607–5612. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 11. Reimer L.C., Vetcininova A., Carbasse J.S., Söhngen C., Gleim D., Ebeling C., Overmann J.. BacDive in 2019: bacterial phenotypic data for high-throughput biodiversity analysis. Nucleic Acids Res. 2019; 47:D631–D636. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 12. Sayers E.W., Cavanaugh M., Clark K., Pruitt K.D., Sherry S.T., Yankie L., Karsch-Mizrachi I.. GenBank 2024 update. Nucleic Acids Res. 2024; 52:D134–D137. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 13. Chen I.-M.A., Chu K., Palaniappan K., Ratner A., Huang J., Huntemann M., Hajek P., Ritter S.J., Webb C., Wu D.et al.. The IMG/M data management and analysis system v.7: content updates and new features. Nucleic Acids Res. 2023; 51:D723–D732. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 14. Olson R.D., Assaf R., Brettin T., Conrad N., Cucinell C., Davis J.J., Dempsey D.M., Dickerman A., Dietrich E.M., Kenyon R.W.et al.. Introducing the Bacterial and Viral Bioinformatics Resource Center (BV-BRC): a resource combining PATRIC, IRD and ViPR. Nucleic Acids Res. 2023; 51:D678–D689. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 15. Schoch C.L., Ciufo S., Domrachev M., Hotton C.L., Kannan S., Khovanskaya R., Leipe D., Mcveigh R., O’Neill K., Robbertse B.et al.. NCBI Taxonomy: a comprehensive update on curation, resources and tools. Database. 2020; 2020:baaa062. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 16. Sayers E.W., Beck J., Bolton E.E., Brister J.R., Chan J., Comeau D.C., Connor R., DiCuccio M., Farrell C.M., Feldgarden M.et al.. Database resources of the National Center for Biotechnology Information. Nucleic Acids Res. 2024; 52:D33–D43. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 17. Matias Rodrigues J.F., Schmidt T.S.B., Tackmann J., von Mering C.. MAPseq: highly efficient k-mer search with confidence estimates, for rRNA sequence analysis. Bioinformatics. 2017; 33:3808–3810. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 18. Chang A., Jeske L., Ulbrich S., Hofmann J., Koblitz J., Schomburg I., Neumann-Schaal M., Jahn D., Schomburg D.. BRENDA, the ELIXIR core data resource in 2021: new developments and updates. Nucleic Acids Res. 2021; 49:D498–D508. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 19. Quast C., Pruesse E., Yilmaz P., Gerken J., Schweer T., Yarza P., Peplies J., Glöckner F.O.. The SILVA ribosomal RNA gene database project: improved data processing and web-based tools. Nucleic Acids Res. 2012; 41:D590–D596. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 20. Koblitz J., Halama P., Spring S., Thiel V., Baschien C., Hahnke R.L., Pester M., Overmann J., Reimer L.C.. MediaDive: the expert-curated cultivation media database. Nucleic Acids Res. 2023; 51:D1531–D1538. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 21. Rolland C., Wittmann J., Reimer L.C., Sardà Carbasse J., Schober I., Dudek C.-A., Ebeling C., Koblitz J., Bunk B., Overmann J.. PhageDive: the comprehensive strain database of prokaryotic viral diversity. Nucleic Acids Res. 2024; gkae878. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 22. Lagkouvardos I., Pukall R., Abt B., Foesel B.U., Meier-Kolthoff J.P., Kumar N., Bresciani A., Martínez I., Just S., Ziegler C.et al.. The Mouse Intestinal Bacterial Collection (miBC) provides host-specific insight into cultured diversity and functional potential of the gut microbiota. Nat. Microbiol. 2016; 1:16131. [DOI] [PubMed] [Google Scholar]
- 23. Heath T., Bizer C.. Linked data: evolving the Web into a Global Data Space. Synthesis Lectures on the Semantic Web: Theory and Technology. 2011; Springer Cham. [Google Scholar]
- 24. Bast H., Buchhold B.. QLever: a query engine for efficient SPARQL+text search. International Conference on Information and Knowledge Management, Proceedings. 2017; Part:F131841. [Google Scholar]
- 25. Buil-Aranda C., Arenas M., Corcho O., Polleres A.. Federating queries in SPARQL 1.1: syntax, semantics and evaluation. J. Web Semantics. 2013; 18:1–17. [Google Scholar]
- 26. Bateman A., Martin M.J., Orchard S., Magrane M., Ahmad S., Alpi E., Bowler-Barnett E.H., Britto R., Bye-A-Jee H., Cukura A.et al.. UniProt: the Universal Protein Knowledgebase in 2023. Nucleic Acids Res. 2023; 51:D523–D531. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 27. Koblitz J., Reimer L.C., Pukall R., Overmann J.. Predicting bacterial phenotypic traits through improved machine learning using high-quality, curated datasets. 2024; bioRxiv doi:12 August 2024, pre-print: not peer-reviewed 10.1101/2024.08.12.607695. [DOI]
- 28. Mistry J., Chuguransky S., Williams L., Qureshi M., Salazar G.A., Sonnhammer E.L.L., Tosatto S.C.E., Paladin L., Raj S., Richardson L.J.et al.. Pfam: the protein families database in 2021. Nucleic Acids Res. 2021; 49:D412–D419. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 29. Breiman L. Random forests. Mach Learn. 2001; 45:5–32. [Google Scholar]
- 30. Münch P., Mreches R., To X.-Y., Gündüz H.A., Moosbauer J., Klawitter S., Deng Z.-L., Robertson G., Rezaei M., Asgari E.et al.. A platform for deep learning on (meta)genomic sequences. 2023; ResearchSquare doi:9 February 2023, pre-print: not peer-reviewed 10.21203/rs.3.rs-2527258/v1. [DOI]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Data Availability Statement
BacDive data can be freely downloaded in various formats (e.g. CSV, JSON, PDF) without restrictions, except that the origin of the data has to be properly cited when used in other works (CC BY 4.0 license). Registration is necessary to access data through the RESTful API, but registration is free of charge. The AI models from the DiASPora project can be found at https://doi.org/10.5281/zenodo.13757323.





