Abstract
The National Genomics Data Center (NGDC), as part of the China National Center for Bioinformation (CNCB), provides a suite of database resources for worldwide researchers. As multi-omics big data and artificial intelligence reshape the paradigm of biology research, CNCB–NGDC continuously updates its database resources to enhance data usability, foster knowledge discovery, and support data-driven innovative research. Over the past year, notable progress has been achieved in expanding the scope of high-quality multi-omics datasets, building new database resources, and optimizing extant core resources. Notably, the launch of BIG Search enables cross-database search services for large-scale biological data platforms, including NGDC, National Center for Biotechnology Information (NCBI), and European Bioinformatics Institute (EBI). Additionally, several new resources have been developed, covering genome and variation (Hiland Resource, TOAnnoPriDB), expression (TEDD), single-cell omics (PreDigs, scMultiModalMap, TE-SCALE), radiomics (TonguExpert), health and disease (CAVDdb, IDP, MTB-KB, ResMicroDb), biodiversity and biosynthesis (SugarcaneOmics), as well as research tools (Dingent, miMatch, OmniExtract, RDBSB, xMarkerFinder). All these resources and services are freely accessible at https://ngdc.cncb.ac.cn.
Graphical Abstract
Graphical Abstract.
Introduction
The National Genomics Data Center (NGDC), established in 2019, is administratively affiliated with the China National Center for Bioinformation (CNCB), Beijing Institute of Genomics (BIG), and Chinese Academy of Sciences (CAS) [1]. In collaboration with the Institute of Biophysics and the Shanghai Institute of Nutrition and Health of CAS, CNCB–NGDC has established strategic partnerships with numerous organizations nationwide (https://ngdc.cncb.ac.cn/partners).
Rapid advances in high-throughput sequencing are propelling biology into a multi-omics era, with single-cell and spatial omics techniques further increasing data dimensionality and resolution [2–4]. Worldwide large-scale initiatives, such as All of Us [5], Our Future Health [6], Human Cell Atlas [7], Earth BioGenome Project [8], Single-Cell Expression Atlas [9], UK Biobank [10], and ImmPort [11], have generated extensive multimodal datasets across diverse species, tissues, and populations. In parallel, burgeoning repositories of biomedical images and image-derived phenotypes have emerged. Together, these resources provide unprecedented detail for system-level characterization—from cellular atlases and interactions to immune microenvironments, as well as organ- and tissue-level morphology—thereby enabling a wide range of studies including developmental processes [12–14], immune responses [15–17], aging mechanisms [18, 19], disease etiology [20–22], and potential therapeutic targets [23–25]. At the same time, artificial intelligence (AI) is catalyzing a paradigm shift [26], with landmark models such as AlphaFold [27], Geneformer [28], and scGPT [29]. This, in turn, demands standardized, interoperable, and reusable (“AI-ready”) data foundations and rigorous benchmarking systems [30].
Within this context, over the past year, CNCB–NGDC has further expanded its network of subcenters throughout the country (https://ngdc.cncb.ac.cn/subcenter), covering a diversity of data resources related to biodiversity, medicine, tumor, pathogen, marine, etc. (Table 1). Based on this, CNCB–NGDC has launched several new resources and continuously updated existing ones, committed to providing an increasingly comprehensive and intelligent suite of biological resources (Fig. 1) [31–39]. These data resources, knowledge information, and analytical tools span multiple omics fields, including genome, transcriptome, epigenome, and radiome, and are widely applied in research areas such as precision medicine, embryonic development, immune responses, and aging mechanisms. These core resources are further interlinked, forming an integrated network that facilitates seamless traversal across databases, retrieval of contextually pertinent information, and comprehensive inquiry (Fig. 2). In addition, NGDC has significantly enhanced its cross-platform search capabilities, integrating data from important biological resource platforms such as NCBI and EBI (Fig. 3). This has further facilitated the efficient aggregation and coordinated management of heterogeneous, multi-source datasets, improving data utilization efficiency and accessibility, and ensuring that global researchers can effectively integrate and utilize biological data resources across platforms and databases. CNCB–NGDC adheres to its core mission of providing biological information resources and data analysis services to the global research community, committed to supporting researchers at all levels. Broad biological researchers have access to foundational information on biological literature, codes, and databases, enabling a deeper understanding of the field of bioinformatics. Meanwhile, specialist researchers are provided with high-quality data resources that support the rapid iteration of AI models, facilitating their application in specific research domains. Here, we summarize recent developments at CNCB–NGDC and present its core resources and services (Table 2). All resources and services are publicly accessible through the CNCB–NGDC homepage (https://ngdc.cncb.ac.cn).
Table 1.
The NGDC subcenters network
| Subcenter name | Abbreviation | Affiliation | Location | Research focus | Joined date |
|---|---|---|---|---|---|
| Biodiversity | NGDC-BDV | Kunming Institute of Zoology, Chinese Academy of Sciences | Kunming | Advancing biodiversity data across ecological, species, and genetic levels | December, 2023 |
| Tumor gene diagnosis data | NGDC-TGD | The Biomedical Big Data Center, The First Affiliated Hospital, Zhejiang University School of Medicine | Hangzhou | Managing tumor genetic data to enhance cancer diagnostics and clinical outcomes | April, 2024 |
| Traditional Chinese Medicine | NGDC-TCM | China Academy of Chinese Medical Sciences | Beijing | Standardizing TCM data resources by integrating proteomic, metabolomic, and transcriptomic information | June, 2024 |
| Pathogenic microorganism | NGDC-PMO | Guangdong Provincial Center for Disease Control and Prevention | Guangzhou | Sharing pathogenic microorganism genomic data to improve accessibility | August, 2024 |
| Northeast medical genomics | NGDC-NMG | The First Hospital of Jilin University | Changchun | Building region-specific genomic databases with integrated clinical and multi-omics data | November, 2024 |
| Marine organism genomics | NGDC-MOG | Yantai University | Yantai | Developing marine genomic data systems for blue economy and ecological protection | December, 2024 |
| Metabolomics | NGDC-MET | Dalian Institute of Chemical Physics, Chinese Academy of Sciences | Dalian | Enabling standardized and intelligent analysis of large-scale metabolomics data | May, 2025 |
| Gannan medical genomics | NGDC-GMG | The First Affiliated Hospital of Gannan Medical University and The First People’s Hospital of Nankang District | Ganzhou | Establishing a precision medicine ecosystem in Ganzhou through genomic data and biospecimen integration | May, 2025 |
Figure 1.
Overview of data submissions to CNCB–NGDC. (A) Statistics for BioProject and BioSample. (B) Statistics for Experiments and Runs in GSA. (C) Timeline of data growth in GSA. (D) Statistics for genome assemblies in GWH. All statistics are regularly updated and publicly accessible at https://ngdc.cncb.ac.cn/bioproject, https://ngdc.cncb.ac.cn/biosample, https://ngdc.cncb.ac.cn/gsa, and https://ngdc.cncb.ac.cn/gwh.
Figure 2.
The core database resources of CNCB–NGDC organized into major categories. These resources are publicly accessible and searchable via the CNCB–NGDC homepage (https://ngdc.cncb.ac.cn). A full list of databases is available at https://ngdc.cncb.ac.cn/databases.
Figure 3.
BIG Search, developed by CNCB–NGDC, integrates heterogeneous data resources from global data centers, enabling scalable and cross-domain text retrieval.
Table 2.
The recent updated database resources of CNCB–NGDC
| Resource type | Resource name | URL | Short description | Records | Growth | Audience |
|---|---|---|---|---|---|---|
| Search | BIG Search | https://ngdc.cncb.ac.cn/search | A cross-database search platform for biological resources | 1 699 078 390 | 15.44% | General audience |
| Knowledge | Database Commons | https://ngdc.cncb.ac.cn/databasecommons | A catalog of worldwide biological databases | 7346 | 6.19% | General audience |
| Knowledge | BioCode | https://ngdc.cncb.ac.cn/biocode | A centralized repository dedicated to archiving bioinformatics tool codes | 7536 | 38.62% | General audience |
| Knowledge | OpenLB | https://ngdc.cncb.ac.cn/openlb/home | An open-access platform for searching bioscience literature texts | 39 880 128 | 5.71% | General audience |
| Data—raw data and metadata | BioProject | https://ngdc.cncb.ac.cn/bioproject | A public repository of biological research projects | 32 394 | 55.49% | General audience |
| Data—raw data and metadata | BioSample | https://ngdc.cncb.ac.cn/biosample | A public repository of biological research samples | 3 046 071 | 52.19% | General audience |
| Data—raw data and metadata | GSA | https://ngdc.cncb.ac.cn/gsa | An open-access repository for non-human raw sequence reads | 2 457 662 | 52.13% | General audience |
| Data—raw data and metadata | GSA-Human | https://ngdc.cncb.ac.cn/gsa-human | A GSA sub-database specialized for human genetic omics data | 4 366 361 | 27.88% | General audience |
| Data—raw data and metadata | OMIX | https://ngdc.cncb.ac.cn/omix | A public repository for submitting and sharing diverse life-science data | 57 883 | 160.70% | General audience |
| Data—raw data and metadata | GenBase | https://ngdc.cncb.ac.cn/genbase | An open-access repository for nucleotide sequence archiving and sharing | 1 114 470 | 17.93% | General audience |
| Data—genome | Genome Warehouse | https://ngdc.cncb.ac.cn/gwh | A vital repository for whole-genome sequencing and annotation data | 94 204 | 11.81% | Genomics researcher |
| Data—genome | Hiland Resource | https://ngdc.cncb.ac.cn/hiland/ | A database integrating phenome, genome, and GWAS data of highland populations | 700 342 | – | Highland population genomics researcher |
| Data—genome | SugarcaneOmics | https://ngdc.cncb.ac.cn/scod/ | A comprehensive platform for sugarcane functional genomics and breeding applications | 14 | – | Sugarcane genomics researcher |
| Data—variation | GVM | https://ngdc.cncb.ac.cn/gvm | A global repository of genomic sequence variations across all species | ∼2.09 billion | 2.96% | Genomic variation researcher |
| Data—variation | TOAnnoPriDB | http://bigdata.ibp.ac.cn/TOAnnoPriDB | An integrative database for trans-omic annotations and variant prioritization | ∼220 million | – | Non-coding variant researcher |
| Data—transcriptome | Gene Expression Nebulas | https://ngdc.cncb.ac.cn/gen | A database of transcriptomic profiles analyzed through a unified pipeline | 86 584 | 44.87% | Transcriptomics researcher |
| Data—transcriptome | LncExpDB | https://ngdc.cncb.ac.cn/lncexpdb | A repository of human long non-coding RNA expression profiles | 39 253 | 55.8% | Non-coding rna researcher |
| Data—transcriptome | CROST | https://ngdc.cncb.ac.cn/crost | A platform integrating data, knowledge, and tools for spatial transcriptomics | 1313 | 30.99% | Spatial transcriptomics researcher |
| Data—transcriptome | TWAS Atlas | https://ngdc.cncb.ac.cn/twas | A comprehensive knowledgebase of transcriptome-wide association studies | 274 972 | 68.53% | Human TWAS researcher |
| Data—transcriptome | TEDD | https://ngdc.cncb.ac.cn/tedd | A resource for the dynamics of human translation efficiency (TE, TR, EVI) | 1518 | – | Translation efficiency researcher |
| Data—epigenome | EWAS Open Platform | https://ngdc.cncb.ac.cn/ewas | A comprehensive resource for epigenome-wide association studies | 898 918 | 85.15% | Human EWAS researcher |
| Data—epigenome | MethBank | https://ngdc.cncb.ac.cn/methbank | A comprehensive DNA methylation database across various species | 5297 | 49.08% | Methylation researcher |
| Data—radiome | OBIA | https://ngdc.cncb.ac.cn/obia | A repository for archiving biomedical images and related clinical data | 1094 | 17.01% | General audience |
| Data—radiome | OPIA | https://ngdc.cncb.ac.cn/opia | A platform for archiving and sharing plant image data and i-traits | 573 731 | 1.33% | Plant phenotyping imaging specialist |
| Data—multimodal | IDP | https://ngdc.cncb.ac.cn/idp/ | The Immunity Deciphering Project (IDP) big data platform | 99 662 | – | Immunologist |
| Data—multimodal | PreDigs | https://www.biosino.org/predigs/ | A comprehensive resource of context-specific markers for digestive cell annotation | ∼3.4 million | – | Oncologist |
| Data—multimodal | CAVDdb | https://ngdc.cncb.ac.cn/cavd/ | An integrated multi-omics resource for calcific aortic valve disease (CAVD) | 515 | – | Cardiologist |
| Data—multimodal | MTB-KB | https://ngdc.cncb.ac.cn/mtbkb/ | A literature-curated knowledgebase dedicated to Mycobacterium tuberculosis | 74 408 | – | Microbiologist |
| Data—multimodal | ResMicroDb | https://resmicrodb.cncb.ac.cn/ | A comprehensive database of the respiratory microbiome | 106 464 | – | Microbiologist |
| Data—multimodal | TE-SCALE | https://ngdc.cncb.ac.cn/te-scale/ | A single-cell atlas of transposable element expression across human cancers | 330 | – | Cancer biologist |
| Data—multimodal | scMultiModalMap | https://ngdc.cncb.ac.cn/scmultimodalmap | A resource for integrating and analyzing single-cell multimodal data | 174 | – | Omics researcher |
| Data—multimodal | RDBSB | https://www.biosino.org/rdbsb | An open-access resource addressing catalytic biopart data dispersion in synthetic biology | 83 193 | – | Synthetic biologist |
| Tool—web server | BIT | https://ngdc.cncb.ac.cn/bit/ | A cloud-based bioinformatics platform for online computation and analysis | 120 | 41.18% | General audience |
| Tool—web server | TonguExpert | https://www.biosino.org/TonguExpert | A web server for archiving and analyzing tongue images | 5992 | – | Tongue diagnosis specialist |
| Tool—web server | miMatch | https://www.biosino.org/iMAC/mimatch | A web server designed for microbial metabolic background matching | – | – | Microbiologist |
| Tool—web server | xMarkerFinder | https://www.biosino.org/xmarkerfinder | A web server for identification and validation of cross-cohort biomarkers | – | – | Omics researcher |
| Tool—web server | VISTA | https://ngdc.cncb.ac.cn/vista | A genome-based tool for rapid, scalable virus taxonomy | 22 036 | 21.71% | Virologist |
| Tool—on-premise tool | Dingent | https://ngdc.cncb.ac.cn/biocode/tool/BT008001 | A lightweight LLM agent framework for application development | – | – | General audience |
| Tool—on-premise tool | OmniExtract | https://ngdc.cncb.ac.cn/biocode/tool/BT007992 | An LLM-based extraction tool for data extraction | – | – | General audience |
| Tool—on-premise tool | scVar | https://ngdc.cncb.ac.cn/biocode/tool/BT008000 | A workflow for integrating expression and variation at single-cell resolution | – | – | Single-cell genomics researcher |
Search
BIG Search
BIG Search (https://ngdc.cncb.ac.cn/search) is a distributed and scalable full-text search engine for a large number of biological resources, providing one-stop cross-database search services for the global research community. In its current version, BIG Search integrates both the NGDC internal databases and 64 partner databases (https://ngdc.cncb.ac.cn/partners), resulting in a total of 1.695 billion data entries and over 1.8 terabytes of data. Furthermore, it supports two mechanisms for automated data index updates: scheduled triggering and change data capture. Additionally, it incorporates 35 NCBI biological databases via e-utilities and 165 EBI biological datasets through API. BIG Search is equipped with advanced search functions and cross-database search services for a wide range of data resources, offering users a more convenient and efficient means for data search and retrieval.
Knowledge
Database commons
Database Commons (https://ngdc.cncb.ac.cn/databasecommons) is a curated, categorized catalog of biological databases worldwide, providing impact assessments and valuable statistical insights [40]. Currently, it indexes 7346 biological databases linked to 10 965 publications and 2440 organizations. Based on the average annual citation rate (z-index), the top 10 databases are: DAVID [41], KEGG [42], cBioPortal [43], STRING [44], AlphaFold DB [45], UniProt [46], SILVA [47], ENCODE [48], gnomAD [49], and IGSR [50]; the top 10 institutions are: European Bioinformatics Institute, NCBI, National Cancer Institute, Broad Institute, Swiss Institute of Bioinformatics, Kyoto University, Memorial Sloan Kettering Cancer Center, Stanford University, University of Alberta, and Max Planck Institute for Marine Microbiology. Moreover, Database Commons is equipped with a statistical analysis function that can generate quantitative charts upon keyword search, enabling comparisons by country, institution, category, data type, object, and species.
BioCode
BioCode (https://ngdc.cncb.ac.cn/biocode) is a centralized repository dedicated to archiving bioinformatics tool codes. It features an extensive collection of bioinformatics tools worldwide developed for diverse data analysis purposes, including their names, descriptions, source codes, categories, associated grants, publications, owner information, and organizational affiliations. This year, it has been greatly enhanced by not only allowing any user to submit tools but also retrieving tool information through automated literature curation from leading journals in the bioinformatics field. It has been also further improved by incorporating new features for online tool management and curation. As a result, the current version of BioCode has archived a total of 7520 bioinformatics tools. This centralized archiving has the significant advantage of improving the utility of these bioinformatics tools by making them publicly available, accessible, and usable.
OpenLB
The Open Library of Bioscience (OpenLB; https://ngdc.cncb.ac.cn/openlb/home) is a scalable and distributed platform designed for storing and retrieving biological literature. In its present version, it incorporates a collection of over 39 million accessible literatures from PubMed (https://pubmed.ncbi.nlm.nih.gov/) [51], bioRxiv (https://www.biorxiv.org/), and medRxiv (https://www.medrxiv.org/). The search algorithm prioritizes key attributes such as publication title and author to enhance retrieval precision and efficiency. The search results page has been optimized to include more comprehensive literature information, such as publication summary and associated data tags, and offers a more user-friendly display of results pagination. Furthermore, it is configured to automatically link with datasets from multiple databases, including GSA, GWH, GenBase, BioProject, BioSample, Database Commons, and BioCode among others, enabling users to easily identify corresponding datasets during literature searches.
Data
Raw data and metadata
BioProject and BioSample
BioProject (https://ngdc.cncb.ac.cn/bioproject) and BioSample (https://ngdc.cncb.ac.cn/biosample) are two public repositories of biological research projects and samples, respectively. They collect descriptive metadata on biological projects and samples investigated in experiments, enable centralized access to all public projects and samples, and offer cross-links to associated data resources. Up to August 2025, BioProject and BioSample have gathered a total of 32 394 biological projects and 3 046 071 biological samples submitted by 16 797 users, exhibiting a significant growth in contrast to last year’s 20 833 projects and 2 001 551 samples. Regarding international data exchange, these two repositories have mirrored 794 435 projects and 46 031 341 samples of INSDC (International Nucleotide Sequence Database Collaboration) data from NCBI, and shared 15 880 projects and 722 959 samples to INSDC via DNA Data Bank of Japan (DDBJ).
GSA and GSA-Human
The Genome Sequence Archive (GSA; https://ngdc.cncb.ac.cn/gsa) [52–54] is an open-access repository designed for the archiving, retrieving, and sharing of non-human raw sequence reads. GSA for Human (GSA-Human; https://ngdc.cncb.ac.cn/gsa-human) [52, 53, 55], a sub-database of GSA, functions as a controlled-access repository, concentrating on human genetic omics data. Till August 2025, GSA and GSA-Human have together accumulated 2 645 702 experiments, 2 889 823 runs, and a total of 83.09 PB of data. Additionally, GSA has integrated 35 805 094 experiments, 37 924 684 runs, and 21.2 PB of raw sequence files from the INSDC’s data resources. GSA maintains close and ongoing engagement with INSDC, actively seeking opportunities to become a member of them and participate in the establishment of international data-sharing policies and data standard specifications. In 2025, INSDC announced its criteria for new membership acceptance. In response, GSA is actively collaborating with DDBJ to develop a data-sharing channel between GSA and DRA, ensuring alignment with these requirements.
OMIX
The Open Archive for Miscellaneous Data (OMIX; https://ngdc.cncb.ac.cn/omix) [52, 53], a member of the GSA family, strictly adheres to the FAIR principles and serves as a versatile data repository dedicated to the collection, publication, and sharing of scientific data for biological researchers. As of August 2025, OMIX has archived 8451 datasets comprising 49 432 files, totaling over 154.70 TB of data.
GenBase
GenBase (https://ngdc.cncb.ac.cn/genbase) is a user-friendly platform for the archiving, retrieval, and sharing of sequences [56]. As of August 2025, it has received 1 114 470 sequences from 513 submitters, supporting 99 publications, and released 71 499 SARS-CoV-2 genomes. It integrates ∼650 million sequences from INSDC [57] and 390 million sequences from RefSeq [58] for localized data access. GenBase offers batch quality control and annotation tools for influenza virus genomes. In addition, GenBase has developed a sequence validation tool to ensure data accuracy prior to submission and implemented a controlled submission system for human-related sequences. Furthermore, GenBase provides a version comparison function for the sequences and their updates. In the future, GenBase will continue to enhance its data management capabilities and web-based genome annotation tools.
Genome
Genome Warehouse
The Genome Warehouse (GWH; https://ngdc.cncb.ac.cn/gwh) serves as a vital public repository for genome assembly sequences, annotations, and associated metadata [59]. To improve the standardization of genome submissions, GWH has refined its data organization framework for metagenome-assembled genomes (MAGs) and haplotype genomes, and has significantly enriched the metadata associated with MAGs. To minimize redundancy, a feature has been implemented to detect duplicated genomes. Notably, based on reannotation using the Prokaryotic Genome Annotation Pipeline, GWH now integrates AMRFinderPlus [60] to identify genes associated with antimicrobial resistance, stress response, and virulence. GWH supports both Chinese and English interfaces, enhancing accessibility for a broader user base. These updates improve the standardization, interoperability, and reusability of data, thereby adding substantial value to the global genomics research community.
Hiland Resource
Hiland Resource (HLR; https://ngdc.cncb.ac.cn/hiland/) is a comprehensive database that integrates phenome, genome, and genetic association data of highland populations worldwide [61]. It systematically integrates data from 29 977 highland individuals across 15 publicly available studies, featuring 29 878 206 genetic variants and 700 342 phenotype-genotype associations spanning 185 traits. HLR offers visualizations of phenotypic patterns across different altitudes, populations, and genders. Additionally, it provides dynamic interfaces for exploring the genetic structure and footprints of natural selection among various highland populations. The genotype-phenotype associations based on highland genome-wide association studies are also integrated and visualized, along with the genome-wide variants and genes. Additionally, HLR provides a user-friendly tool for genotype imputation of high-altitude populations based on a high-quality reference panel of 1001 Tibetan genomes.
SugarcaneOmics
SugarcaneOmics (https://ngdc.cncb.ac.cn/scod/) constitutes an indispensable platform addressing critical deficiencies in sugarcane functional genomics and breeding applications, where yield stagnation persists due to the crop’s intricate hybrid polyploid genome [62]. This repository integrates multi-omics datasets spanning 14 sugarcane species and relatives, encompassing genome assemblies, transcriptomic profiles (1256 samples), genetic variations (322 re-sequenced germplasms harbouring ∼175 million variants), and curated functional annotations. Core functionalities incorporate five interactive modules (Genome, Transcriptome, Variome, Feature Genes, and Varieties) alongside specialized toolkits for sequence analysis, gene function elucidation, and CRISPR off-target prediction. The platform facilitates unprecedented cross-species comparative genomics and batch analyses, thereby expediting the translation of genomic discoveries into breeding implementations. This resource holds considerable promise for advancing sustainable agriculture by bridging genomic research and precision breeding for the global scientific and breeding communities.
Variation
GVM
The GVM (https://ngdc.cncb.ac.cn/gvm) [63, 64] (reference in this issue) is a global repository that collects, integrates, and facilitates the submission of genomic variations—including single nucleotide polymorphisms and small insertions/deletions (InDels)—from a wide range of species worldwide. As of August 2025, GVM houses ∼2.09 billion variants from 73 species, deriving from 437 projects and 101 967 samples, all of which have been manually curated and processed through a standardized analysis pipeline. Meanwhile, it has archived 880 data submissions covering 754 879 samples across 78 species submitted from 182 organizations. In this version, a new model with deleterious variant information and population genetic selection signals has been launched. A new online tool (VersionMap) is introduced, which enables convenient mapping of variants across different genome assemblies. Collectively, these updates will strengthen GVM’s capacity and enhance its value as a global genomic variation resource.
TOAnnoPriDB
TOAnnoPriDB (http://bigdata.ibp.ac.cn/TOAnnoPriDB) is a comprehensive resource for annotation and prioritization of non-coding variants across human genome [65]. Based on the NyuWa genome resource [66], TOAnnoPriDB aims to help emphasize the potential functional effects of variants and the relationship between variants and human disease. TOAnnoPriDB covers ∼98% of the non-coding region of human genome in total and integrates trans-omic information from 147 public resources, including several databases we developed previously, namely NCVD [66–70], NONCODE [71], NPInter [72], piRBase [73], SmProt [74], and LncVar [75]. Based on the data integration, TOAnnoPriDB constructs a framework to prioritize variations according to evidence that supports the functional impact of variants and provides a user-friendly web interface to help users search and analyze variants. JBrowse 2 is incorporated to visualize annotation information. The allele frequency, gene expression, and molecular interaction related to the variations are also visualized. TOAnnoPriDB can serve as a powerful tool for variant annotation and prioritization, which can help users explore and understand the association between non-coding variants and human diseases.
Transcriptome
Gene Expression Nebulas
Gene Expression Nebulas (GEN; https://ngdc.cncb.ac.cn/gen) is a data portal integrating transcriptomic profiles from both bulk and single-cell levels in various conditions across multiple species [76]. As of August 2025, GEN has collectively accumulated 91 species, 629 datasets, 86 584 samples, 19 281 005 cells, and 584 publications, marking a significant increase compared to the previous version. This year, we have systematically incorporated 50 new species to expand the species diversity, including 40 plants, 5 animals, and 5 protists. Moreover, 26 816 samples related to aplastic anemia and Alzheimer’s disease have been integrated and analyzed to enrich the data volume. Additionally, we have updated the GENToolkit to encompass standardized analyses of global gene expression profiles, differential gene expression, and functional enrichment, offering more comprehensive and multi-tiered analytical and visualization capabilities, along with improved robustness and user-friendliness.
LncExpDB
LncExpDB (https://ngdc.cncb.ac.cn/lncexpdb) is a comprehensive resource that integrates and rigorously curates human long non-coding RNA (lncRNA) expression profiles across diverse biological contexts [77] (reference in this issue). Built upon the standardized gene reference from LncBook [78], it evaluates expression reliability, highlights featured genes, and identifies lncRNA–messenger RNA (mRNA) interactions by co-expression networks. This year’s update introduces three additional biological contexts: neurodegenerative disease, reproduction, and wound healing. It also incorporates two newly developed tools: LncNet for visualizing and exploring lncRNA–mRNA co-expression networks and LncImm for examining lncRNA correlations with immune checkpoint genes across four cancers with immunotherapy data. These enhancements are part of the major release of LncExpDB 2.0, launched in 2025. Compared with version 1.0, launched in 2020, LncExpDB 2.0 expands the data coverage from 9 to 15 biological contexts, identifying 39 253 featured genes (among 101 293) and >28 million co-expression pairs. It further enhances functionality with new analysis tools, a pipeline module, and in-depth functional analyses.
CROST
CROST (https://ngdc.cncb.ac.cn/crost) integrates data, knowledge, and tools for spatial transcriptomics [79]. Since August 2023, it has added three model organisms (Sus scrofa, Gallus gallus, Oryctolagus cuniculus) and expanded datasets covering disease, environmental, and developmental contexts. A new tool, SCAN, was introduced for cell type annotation. CROST now contains 1313 high-quality samples from 234 projects and 11 species, a 30% increase over the previous version. SCAN integrates six marker databases to provide predictions comparable to manual annotation, and users can upload data and export results. Together with SpatialAP, CROST enables one-stop single-cell and spatial transcriptomic analysis, offering a comprehensive view of tissue architecture and advancing understanding of biological mechanisms, particularly in tumors.
TWAS Atlas
TWAS Atlas (https://ngdc.cncb.ac.cn/twas/) is a comprehensive data resource that systematically consolidates published transcriptome-wide association study (TWAS) findings [80]. Recently, we curated high-quality gene-trait associations and expanded phenotypic coverage by conducting TWAS analysis based on public GWAS datasets. Overall, the number of associations increased by 274 972, derived from 224 publications and 171 datasets. Regarding feature updates, we improved the interactive visualization in TWAS Atlas 2.0. The Knowledge Graph now includes supporting evidence descriptions extracted from original publications. Furthermore, we added multidimensional analysis modules, including functional enrichment, Mendelian randomization, colocalization, and fine-mapping, providing researchers with a comprehensive toolkit to investigate causal gene-trait relationships. TWAS Atlas 2.0 offers a comprehensive platform for identifying risk loci for complex traits and exploring potential regulatory mechanisms underlying various diseases.
TEDD
TEDD (https://ngdc.cncb.ac.cn/tedd) is a publicly accessible database dedicated to the systematic collection, analysis, and visualization of human translation efficiency dynamics (TE, TR, EVI) with a strong emphasis on regulatory 5′/3′ UTR features (reference in this issue). The current release contains 279 datasets comprising 1518 samples from 143 projects, spanning 24 tissues/cell types, 74 cell lines, and 52 conditions. The translation efficiency dynamics metrics are calculated at both the gene and transcript levels, complemented by detailed UTR annotations to support fine-grained analysis of translational regulation. An integrated online analysis platform allows multidimensional comparisons of TE/TR/EVI across genes, transcripts, KEGG/GO-defined gene sets, and UTR elements. Overall, TEDD provides a comprehensive resource for advancing both basic research and translational applications in translatomics—particularly in areas such as mRNA vaccine development, synthetic biology, gene therapy, and enzyme engineering—by supporting rational design of gene expression systems for efficient protein production.
Epigenome
EWAS Open Platform
EWAS Open Platform (https://ngdc.cncb.ac.cn/ewas) serves as a continuously updated resource dedicated to epigenome-wide association studies (EWAS), integrating data, knowledge, and toolkit [81] (reference in this issue). A unified retrieval function and an AI-based Q&A Assistant have been introduced, substantially enhancing the efficiency of cross-module knowledge retrieval and interactive exploration. On the data side, GMQN, a batch effect correction tool optimized for the latest DNA methylation 935K arrays, has been upgraded [82] and 20 373 new high-quality samples have been incorporated into EWAS Data Hub [83]. On the knowledge side, EWAS Open Platform has added 54 667 epigenetic associations curated from publications [84], and an interactive multi-omics regulatory network integrating 17 494 causal relationships among DNA methylation, gene expression, and diverse traits. These updates aim to facilitate deeper exploration of the regulatory mechanisms underlying disease onset and progression.
MethBank
The Methylation Bank (MethBank; https://ngdc.cncb.ac.cn/methbank/) [85–87] is a comprehensive repository of whole-genome, single-base resolution DNA methylation across multiple species and biological contexts. To support metadata standardization, AI-assisted Metadata Curation System that integrates pretrained LLMs with ontology hierarchies has been implemented, enabling precise alignment of attributes such as tissue/cell line and disease. Plant-specific attributes, including geographic location, isolate, ecotype, and generation, have also been incorporated. In its latest release, MethBank documents a 49.08% increase in data volume, with the addition of 588 human samples (Homo sapiens) and 1156 plant samples (Arabidopsis thaliana). As of August 2025, MethBank integrates ~648 million gene-level methylation profiles from 26 species, which are derived from 410 high-quality projects and 5297 samples, and curated under a structured metadata framework with standardized analytical pipelines to ensure reproducibility and cross-study comparability.
Radiome
OBIA
The Open Biomedical Imaging Archive (OBIA; https://ngdc.cncb.ac.cn/obia) [52, 88] is a repository of biomedical images alongside its clinical data. Data are organized into five hierarchical objects—Collection, Individual, Study, Series, and Image—with Individuals linked to GSA-Human via accession numbers to enable multi-omics research. To ensure data privacy and integrity, OBIA implements standardized de-identification and quality control procedures and provides both open and controlled access. Beyond conventional web-based querying and browsing, OBIA supports advanced image retrieval. As of September 2025, OBIA has housed a total of 1094 individuals, 4349 studies, 24 918 series, and 2 012 007 images covering 9 modalities and 30 anatomical sites.
OPIA
The OPIA (https://ngdc.cncb.ac.cn/opia/) [89] is an open-access platform for archiving and sharing plant image datasets and associated image-based traits (i-traits) derived from high-throughput phenotyping technologies. As of August 2025, OPIA hosts 89 datasets from 43 plant species, comprising 573 731 images and 2 424 692 labeled instances. These datasets are AI-ready, supporting classification and identification tasks. OPIA also offers a suite of online tools for image pre-processing and intelligent phenotypic analysis. Among them, Img2Variety is a new CNN-based framework for crop accession identification using whole-plant images across all developmental stages. It achieves an accuracy of 88.66% for rice and ∼80% for maize intraspecific variety identification. In summary, OPIA is a useful resource for advancing crop research, enhancing breeding efficiency, and ultimately driving agricultural progress.
Multimodal
IDP
Immunity Deciphering Project (IDP; https://ngdc.cncb.ac.cn/idp/) is a comprehensive resource for the submission, management, integration, and sharing of immunity-related data, designed to facilitate the systematic digital decoding of human immunity. Supported by the National Natural Science Foundation of China’s major research program, the platform has established standardized terminology and data management systems and has facilitated the submission of 42 projects, encompassing 11 246 samples with 42.68 TB of data. Building on this foundation, IDP has standardized and integrated 617 multi-source and multi-modal datasets, covering 92 662 samples across 61 disease types. These datasets span diverse modalities, including scRNA-seq, scBCR-seq, scTCR-seq, ATAC-seq, Bisulfite-seq, and Hi-C, and are made accessible through cross-type search and metadata download functionalities. Overall, IDP provides high-quality data submission and sharing support for understanding the immune system and enabling immune interventions and personalized medicine.
PreDigs
PreDigs (https://www.biosino.org/predigs/) is a user-friendly database of predicted cell-type signatures in the digestive system, integrating single-cell RNA sequencing data from gastrointestinal tumors to explore cellular composition and heterogeneity [90]. It contains 124 curated datasets encompassing over 3.4 million cells, all of which are available for download. Subtype labels are unified, and a cell ontology tree was constructed with 142 cell types across eight hierarchical levels. Meanwhile, we calculated three different context-specific cell-type markers, including “Cell Markers,” “Subtype Markers,” and “TPN Markers,” based on various application requirements within or across tissues. Through the integrated analysis of PreDigs data, we identified distinct cell subpopulations exclusive to tumors, one of which corresponds to tumor-specific endothelial cells. Additionally, PreDigs offers online cell annotation tools, allowing users to classify single cells with greater flexibility. Overall, the database provides valuable insights into tumor heterogeneity and its impact on disease progression and treatment response.
CAVDdb
Calcific Aortic Valve Disease Database (CAVDdb; https://ngdc.cncb.ac.cn/cavd/) is an integrated multi-omics resource for calcific aortic valve disease (CAVD), the leading cause of aortic stenosis worldwide. The current release consolidates diverse datasets from human tissues and cell line models, including 24 projects and 515 samples. Specifically, it integrates transcriptomic data from 14 BioProjects and 214 BioSamples, covering 35 894 genes in tissues and 19 297 genes in cell lines, with 5453 differentially expressed genes across five groups. It also includes proteomic profiles of 5234 proteins (303 differentially expressed), metabolomic data of 480 metabolites (214 differentially expressed), single-cell transcriptomic data from 145 045 cells across six cell types, and epigenomic modules including DNA methylation levels of 78 311 genes from aortic valve samples, as well as one ATAC-seq and six ChIP-seq samples. CAVDdb offers three interactive tools, including a genome browser and functional enrichment modules. Overall, CAVDdb provides a comprehensive resource, facilitating mechanism studies and therapeutic target discovery.
MTB-KB
MTB-KB (https://ngdc.cncb.ac.cn/mtbkb/) is the first literature-curated knowledgebase dedicated to Mycobacterium tuberculosis (MTB) and tuberculosis (TB). It integrates 74 408 high-confidence associations manually extracted from 1187 publications, including 18 232 biological entities across eight major sections: Epidemiology, Diagnosis, Drug, Regimen, Vaccine, Drug Resistance, Virulence Factor, and Immune Mechanism. All entities are annotated and standardized using authoritative resources, thereby improving consistency, interoperability, and clinical relevance. Especially, it constructs an interactive knowledge graph that reveals cross-sectional relationships regarding MTB-host interactions, treatment strategies, and vaccine development opportunities, enabling multi-dimensional analysis, association, and inference. MTB-KB also features user-friendly modules for section-based browsing, quick and advanced search, statistical visualization, and data download. MTB-KB fills a critical gap in TB research by systematically consolidating literature-based knowledge and enhancing its interpretability through standardized annotation and graph analytics, providing a valuable and innovative platform that supports basic research, translational applications, and global TB control efforts.
ResMicroDb
ResMicroDb (https://resmicrodb.cncb.ac.cn/) is a comprehensive database and analysis platform dedicated to human respiratory microbiome research. It integrates 106 464 samples from 514 projects, spanning 10 respiratory tract sites and 146 phenotypes, along with 31 curated metadata fields (reference in this issue). The database also includes 11 908 microbe-disease associations identified from 132 case-control studies. ResMicroDb features a user-friendly web interface for querying, browsing, visualizing, and analyzing respiratory microbiome data. It provides three integrated analytical tools: Microbiome Composition, for visualizing taxonomic profiles; Sample Similarity Search, for inferring sample characteristics through similarity-based comparisons; and Cross-study Analysis, for identifying shared and specific microbial features across cohorts, sites, and diseases. In summary, ResMicroDb serves as a vital resource for advancing research on the respiratory microbiome and its clinical applications.
TE-SCALE
TE-SCALE (https://ngdc.cncb.ac.cn/te-scale/) is a comprehensive database for exploring transposable element (TE) expression across human cancers at single-cell resolution, addressing the underrepresentation of TEs in conventional gene-centric analyses (reference in this issue). It is constructed from publicly available single-cell transcriptomic datasets comprising over 1.3 million high-quality cells from 330 samples across 20 cancer types and 12 tissue origins. Powered by the in-house streamlined computational pipeline scTEfinder, TE-SCALE enables robust quantification of 1051 curated TE subfamilies, integrates gene and TE expression profiles, and provides precise cell-type annotation. It delivers a pan-cancer TE expression atlas with interactive, multi-scale exploration and three analytical modules: differential TE expression, TE-gene co-expression network, and functional enrichment. Notably, TE-SCALE is the first database, to our knowledge, that identifies tumor-specific TEs preferentially expressed in particular cancer types or pathological states, underscoring their potential as biomarkers for diagnosis, disease monitoring, and immunotherapeutic targeting.
scMultiModalMap
The single-cell multimodal data map (scMultiModalMap; https://ngdc.cncb.ac.cn/scmultimodalmap/) is a comprehensive resource for collecting, integrating, visualizing, and analyzing single-cell multimodal data. It currently includes 174 datasets covering two types of multimodality, 15 sequencing protocols, 3 antibody panels, and 79 conditions, encompassing over 3.2 million cells. Users can access detailed information about each dataset, including sample quality control, multimodal integration-based UMAPs, differential analyses (gene expression, protein abundance, and chromatin accessibility), and enrichment analysis results. Seven online analytical modules are available, consisting of two single-modality modules, three cross-modality modules, and two cell-level modules. These modules enable users to visualize modality-specific features across different cell types, explore relationships among features across various modalities, and analyze changes in cell composition or cell–cell communication events related to biological environments. Overall, scMultiModalMap serves as a crucial resource for exploring single-cell multimodal data, enabling users to understand cellular heterogeneity and advance their knowledge of cellular functions.
RDBSB
The Registry and Database of Bioparts for Synthetic Biology (RDBSB; https://www.biosino.org/rdbsb/) is a comprehensive, open-access resource addressing the dispersion of catalytic biopart functional data across databases and literature, a key barrier to designing metabolic pathways and optimizing cell factories in synthetic biology [91]. In its current release, RDBSB systematically curates 83 193 catalytic bioparts with experimental evidence, providing detailed qualitative and quantitative catalytic information, such as activities, substrates, optimal pH/temperature, and chassis specificity. The platform features an interactive search engine, visualization tools, and practical utilities like biopart finder, structure prediction, and pathway design tools, directly supporting synthetic biology workflow needs. Additionally, RDBSB supports community-driven biopart submission, which has facilitated over 1000 user contributions to date, enabling rapid data sharing. As a freely available resource, RDBSB significantly enriches pathway design resources and serves as an essential tool for advancing synthetic biology research, from basic biopart characterization to applied cell factory optimization.
Tool
Web server
BIT
Bioinformatics Toolkits (BIT; https://ngdc.cncb.ac.cn/bit/) is a cloud-based bioinformatics platform for online computing and analysis. It integrates a variety of frequently used bioinformatics tools, allowing users to select the tools they need, upload data, customize parameters, and quickly complete various analysis tasks. Users can also browse and download the analysis results. The current version of the software integrates 10 categories of tools, including visualization, sequence alignment, image processing, and omics data analysis, among others, and deploys 120 analysis tools in total. BIT offers a user manual that provides operational guidance and usage examples to facilitate quick initiation. It also features a tool navigation bar and search functionality to help users quickly locate the tools they need, thereby enhancing the user experience.
TonguExpert
TonguExpert (https://www.biosino.org/TonguExpert) is a free automated platform for archiving, analyzing, and extracting detailed phenotypes from tongue images to enhance the objectivity and accuracy of tongue diagnosis in clinical and research settings [92]. It integrates deep learning algorithms (such as YOLOv8 and ResNet50) with refined preprocessing, hosting the largest publicly available tongue image database to date, comprising 5992 high-quality images from a Chinese population, and has established a fine-grained phenotype library containing 773 phenotypes—355 global phenotypes from the whole tongue, tongue body, and tongue coating, and 408 local phenotypes from tongue fissures and tooth marks. TonguExpert demonstrates high predictive performance for four Traditional Chinese Medicine (TCM) phenotypes (ROC-AUC 0.89–0.99 for color) and exhibits strong generalization to novel phenotypes such as greasy coating. TonguExpert provides a user-friendly web interface for image upload and phenotype extraction, advancing automated, interpretable tongue diagnosis for TCM and precision health research.
miMatch
miMatch (https://www.biosino.org/iMAC/mimatch) is a web server designed to promote the integration of metagenomic cohorts and strengthen causal relationships in metagenomic research [93]. Low concordance across metagenomic studies remains a major challenge, largely due to host-related confounders such as genetics, environment, and lifestyle. To address this, miMatch leverages the microbial metabolic background as a comprehensive reference for host-related variables and applies propensity score matching to construct balanced case-control pairs. miMatch shows robust performance in both simulated and real-world data, effectively reducing false positives in microbial signature detection and enhancing result consistency and model generalizability across cohorts. The web server provides an easy-to-use platform for applying this framework, enabling users to easily incorporate miMatch into their studies. By constructing well-matched cohorts, miMatch helps researchers generate more reliable and generalizable insights into microbiome-disease associations.
xMarkerFinder
xMarkerFinder (https://www.biosino.org/xmarkerfinder/) is a computational platform for robust biomarker discovery across heterogeneous multi-cohort datasets [94]. Variability introduced by different cohorts, sequencing platforms, and study designs often limits reproducibility in microbiome and other omics research. xMarkerFinder addresses this challenge by integrating statistical modeling, cohort harmonization, and rigorous feature selection to minimize batch effects and ensure cross-study consistency. Primarily designed for microbiome analysis, it can also be extended to other omics, phenotypic, and environmental datasets. The web server provides both a streamlined one-click workflow for rapid biomarker identification and a customizable stepwise mode for interactive exploration. Curated metadata from diverse publicly available microbiome studies allows researchers to efficiently identify and select relevant datasets and facilitates data reuse. By combining reproducibility, scalability, and accessibility, xMarkerFinder offers a versatile platform for reliable biomarker discovery and translational applications.
VISTA
VISTA (https://ngdc.cncb.ac.cn/vista) is a computational tool that integrates a novel pairwise sequence-comparison system with an automatic threshold-identification framework for virus taxonomy [95]. By leveraging k-mer profiles, physico-chemical property sequences, and machine-learning techniques, VISTA constructs a robust distance-based model for taxonomic assignment. The tool has been applied to 22 036 complete viral genomes, covering six Baltimore classes, the class Caudoviricetes, and 39 additional families, demonstrating scalability across both prokaryotic and eukaryotic viruses. Importantly, VISTA also assigns previously unclassified viral genomes, providing objective species- or genus-level assignment. Application to 679 unclassified prokaryotic genomes from metagenomic datasets led to the recognition of 46 novel virus families [95]. VISTA-generated demarcation thresholds constitute data-driven “gold standards” for genus and species boundaries and are now being promoted for families such as Arteriviridae, Dicistroviridae, and Filoviridae. Overall, VISTA provides an efficient framework for genome-based virus classification and facilitates the integration of newly sequenced genomes into official taxonomy frameworks.
On-premise
Dingent
Dingent (https://ngdc.cncb.ac.cn/biocode/tool/BT008001) is a lightweight, user-friendly agent framework designed to simplify and accelerate the development of agent applications powered by LLMs. It integrated a backend service (LangGraph; https://www.langchain.com/langgraph), a ready-to-use chat interface, and a full-featured admin dashboard, significantly reducing the amount of repetitive “glue code” typically needed when building agent-based applications. A key feature of Dingent is its powerful web-based admin dashboard, which enables users to configure assistants, build workflows, and adjust settings through an intuitive graphical interface, eliminating the need for manual file editing. Additionally, Dingent supports instant project initialization with a single command and offers an extensible plugin system for incorporating custom tools. Overall, by streamlines the development process, Dingent allows developers to concentrate on core logic and facilitates the efficient creation of sophisticated, customizable AI agents.
OmniExtract
OmniExtract (https://ngdc.cncb.ac.cn/biocode/tool/BT007992) is an LLM-based automatic extraction tool specifically designed for information extraction from literature and documents. It utilizes prompt optimization engineering to enhance extraction performance based on curated data and provides various file format parsing tools. The tool supports batch extraction of multi-property entities from original documents (such as PDF or XML) as well as tabular files. OmniExtract offers a user-friendly approach: the entire extraction process, including model configuration, file parsing, prompt optimization, and information extraction, can be easily customized by modifying configuration files. Evaluation results show that when using open-source models, OmniExtract achieves a high accuracy with ranges from 82.63% to 89.00% across three public datasets: the WikiReading Recycled dataset [96], the CriticalCoolingRates dataset [97], and the Yidu-S4K dataset. Moreover, when applied to extract phenotypic information from dog breed standard documents, the accuracy exceeds 90% after optimization. Overall, OmniExtract provides convenient and reliable information extraction functionality with consistently stable performance.
scVar
scVar (https://ngdc.cncb.ac.cn/biocode/tool/BT008000) is a comprehensive workflow for the detection and functional characterization of single-nucleotide variants from 10× Genomic scRNA-seq data. This tool demonstrates high accuracy in identifying somatic mutations, exhibiting enhanced sensitivity particularly for low-frequency variants, and operates effectively without requiring paired normal samples or predefined cell-type annotations. Subsequent to variant detection, scVar integrates multiple databases to perform functional and clinical annotations of the identified variants. Additionally, the workflow encompasses a suite of downstream analytical modules tailored to distinct cellular subpopulations, including mutational signature analysis, quantification of tumor mutational burden and clonal diversity, functional enrichment analysis of mutated genes, and differential gene expression profiling within mutated cells. Overall, scVar improves the accuracy of identifying low-frequency variants and enables joint analysis of expression and variation at single-cell resolution, offering new insights into tumor heterogeneity.
Concluding remarks
Amid the exponential growth of multi-omics and multi-modal data, CNCB–NGDC is committed to providing an increasingly comprehensive and intelligent suite of database resources, facilitating efficient, high-quality data archiving, integration, and utilization, and offering broad support to the global research community to drive transformative advances in life, health, and medical sciences. Looking ahead, CNCB–NGDC will continue to deeply integrate AI technologies to enhance the management, integration, and analysis of multi-omics data, optimize big data storage and computing platforms, and develop efficient analytical tools and pipelines for addressing complex biological questions. These efforts are expected to accelerate knowledge discoveries in genomics, molecular biology, and biotechnology, foster cross-disciplinary research and innovation, and support a wide range of applications worldwide in personalized medicine, precision diagnostics, drug discovery, crop breeding, and biosafety.
Acknowledgements
We thank our users for submitting data, sending suggestions, reporting bugs, and getting involved in community curation. CNCB–NGDC is indebted to its funders, including the Ministry of Science & Technology and the Ministry of Finance of the People’s Republic of China, as well as Chinese Academy of Sciences.
Appendix
Corresponding author: Yiming Bao1,2,3,*
Co-corresponding authors: Zhang Zhang1,2,3,*, Wenming Zhao1,2,3,*, Jingfa Xiao1,2,3,*, Shuhui Song1,2,3,*, Shunmin He3,4,*, Guoqing Zhang5,6,*, Yixue Li5,7,*, Guoping Zhao5,8,*, Runsheng Chen3,4,*
CNCB–NGDC MEMBERS (Arranged by project role and then by contribution except for Team Leader (TL), as indicated)
Hiland Resource: Yibo Wang1,2,3,#, Weijie Zhang3,9,#, Xiaoning Chen1,2,3,#, Yanling Sun1,2,#, Bixia Tang1,2, Yu Zhang9, Kai Liu3,9, Wenming Zhao1,2,3,* (TL), Bing Su9,10 (TL), Yaoxi He9,10 (TL)
TOAnnoPriDB: Tingrui Song4,#, Yirong Shi3,4,#, Yanyan Li4, Di Hao4, Kaixin Zhan4, Tao Xu11,12, Runsheng Chen3,4,*, Shunmin He3,4,*
TEDD: Wenyan Lei2,3,13,#, Cuidan Li2,13,#, Hengyu Zhou2,13,14,#, Zhi Nie1,2,3,#, Anke Wang1,2,#, Pan Li2,3,13, Peihan Wang2,3,13, Zhuojing Fan1,2, Rongxi Zhu15, Haoyu Cheng16, Yuxian Guo15, Liya Yue2,13, Xiaoyuan Jiang2,13, Renjun Gao16, Yongjie Sheng16, Haitao Niu17, Tuohetaerbaike·Bahetibieke18, Wenbao Zhang18, Wenming Zhao1,2,3, Jingfa Xiao1,2,3,* (TL), Fei Chen2,3,13,17,18,# (TL)
PreDigs: Jiayue Meng5,#, Mengyao Han19,#, Yuwei Huang5,#, Liyun Yuan5,#, Guoqing Zhang5,6,*
scMultiModalMap: Zhi Nie1,2,3,#, RuiKun Xue1,2,3,#, Zhuojing Fan1,2, Jingfa Xiao1,2,3,*, Jingyao Zeng1,2,#
TE-SCALE: Xini Meng2,3,13,#, Zhi Nie1,2,3,#, Qifei Wang2,3,13, Yiwen Hu2,3,13, Yulan Deng20,21, Na Ai2,13,22,23, Zheng Huang2,3,13, Yun Li2,3,13, Yang Yuan3,24, Jingfa Xiao1,2,3, Jingyao Zeng1,2,#, Guochao Li2,3,13,#, Lan Jiang2,3,13,25,#
TonguExpert: Ting Li5,#, Ling Zuo5,26,#, Jianxin Chen26,27,#, Qianqian Peng5,#, Guoqing Zhang5,6,*, Sijia Wang5,28,#
CAVDdb: Sicheng Wu1,2,3,#, Hao Jiang29,30,#, Yibo Wang1,2,3,#, Hailong Kang1,2,3, Jiawei Shi31, Bixia Tang1,2, Nianguo Dong31 (TL), Wenming Zhao1,2,3,* (TL), Ximiao He29,30,32 (TL)
IDP: Fei Yang1,2,#, Shuai Jiang33,#, Zhenxian Han1,2,#, Xue Bai1,2,#, Dong Zou1,2, Sisi Zhang1,2, Yi Wang1,2,3, Zhijian Duan1,2,3, Lun Li1,2, Zhuojing Fan1,2, Shuhui Song1,2,3,*
MTB-KB: Pan Li2,3,13,#, Cuidan Li2,13,# (TL), Rongxi Zhu15, Wenjing Sun16, Hengyu Zhou2,13,14, Zhuojing Fan1,2, Liya Yue2,13, Sijia Zhang16, Xiaoyuan Jiang2,13, Quan Luo16, Jinying Han16, Hairong Huang34, Adong Shen35, Tuohetaerbaike·Bahetibieke18, Jing Wang18, Wenbao Zhang18, Hao Wen18, Haitao Niu17, Congfan Bu1,2, Zhang Zhang1,2,3, Jingfa Xiao1,2,3, Renjun Gao16,# (TL), Fei Chen2,3,13,17,18,# (TL)
ResMicroDb: Xiaotong Ji1,2,3,36,#, Qiheng Qian1,2,#, Hao Zhang1,2,3,#, Qingyun Cai1,2,3,#, Kaiwen Zhang1,2,3, Jingfa Xiao1,2,3, Xiaoqing Jiang1,2,#, Mingkun Li2,3,13,#
SugarcaneOmics: Hong Luo1,2,#, Xue Bai1,2,#, Zishan Wu1,2,3,#, Siwei Ren1,2,3, Haixia Xie1,2,3, Zhixiang Yuan1,2, Dongmei Tian1,2,#, Shuhui Song1,2,3,*
Dingent: Demian Kong1,2,3,#, Shaoqi Bei1,2,3, Yueyue Wu1,2,3, Bixia Tang1,2,# (TL), Wenming Zhao1,2,3,* (TL)
miMatch: Lei Liu37,#, Suqi Cao38,#, Weili Lin37,#, Guoqing Zhang5,6,*, Ruixin Zhu37,#, Dingfeng Wu38,#
OmniExtract: Yibo Wang1,2,3,#, Bixia Tang1,2,# (TL), Sicheng Wu1,2,3, Yuyan Meng1,2,3, Demian Kong1,2,3, Wenming Zhao1,2,3,* (TL)
RDBSB: Wan Liu5,#, Pingping Wang39,#, Xinhao Zhuang5,#, Xing Yan40,#, Zhihua Zhou39,#, Guoqing Zhang5,6,*
xMarkerFinder: Wenxing Gao37,#, Weili Lin37,#, Qiang Li5,#, Guoqing Zhang5,6,*, Ruixin Zhu37,#, Na Jiao38,41,#
VISTA: Yiyun Liu1,2,3,#, Lili Tian1,2,3,#, Yiming Bao1,2,3,*
scVar: Wei Zhao1,2,3,#, Fei Yang1,2, Wenting Zong1,2, Xinchang Zheng1,2,#, Yiming Bao1,2,3,*
BioProject & BioSample & GSA & GSA-Human: Xu Chen1,2,#, Tingting Chen1,2,#, Sisi Zhang1,2,#, Xiaolong Zhang1,2,#, Yubo Zhou1,2, Anke Wang1,2, Junwei Zhu1,2, Bing Xu1,2, Lili Dong1,2, Caixia Yu1,2, Wenjie Li1,2, Zhuojing Fan1,2, Shuang Zhai1,2, Yubin Sun1,2, Qiancheng Chen1,2, Yanqing Wang1,2,# (TL), Wenming Zhao1,2,3,* (TL)
OMIX: Anke Wang1,2,#, Caixia Yu1,2,#, Yanqing Wang1,2, Sisi Zhang1,2,# (TL)
GenBase: Congfan Bu1,2,#, Xuetong Zhao1,2,#, Xue Bai1,2,#, Jingfa Xiao1,2,3, Zhang Zhang1,2,3, Wenming Zhao1,2,3, Bixia Tang1,2,# (TL), Yiming Bao1,2,3,*
Database Commons: Shaosen Zhang1,2,3,#, Dong Zou1,2,#, Jinbiao Wang1,2,3,#, Yuhao Zeng1,2,3,#, Zheng Luo1,2,3,#, Yiran Zhan1,2,3, Zihan Wang1,2,3,36, Xi Zhao1,2,3, Yuxi Liu1,2,3, Lina Ma1,2,3,# (TL)
Genome Warehouse: Xuetong Zhao1,2,#, Zhenxian Han1,2,#, Yingke Ma1,2,#, Meili Chen1,2,3,# (TL)
GVM: Dongmei Tian1,2,#, Xue Bai1,2,#, Haixia Xie1,2,3,#, Siwei Ren1,2,3,#, Bixia Tang1,2,#, Shuhui Song1,2,3,* (TL)
Gene Expression Nebulas: Tongtong Zhu1,2,#, Wenzhuo Cheng1,2,3,#, Yuan Chu1,2,3,#, Dong Zou1,2,#, Zhenxian Han1,2,#, Ming Chen1,2,3, Zihan Wang1,2,3,36, Yiran Zhan1,2,3, Zheng Luo1,2,3, Tianyi Xu1,2,#(TL), Zhang Zhang1,2,3,*
TWAS Atlas: Hao Gao1,2,3,#, Si Zheng42,43,#, Congfan Bu1,2,#, Jialin Mai1,2,3,#, Jinbei Wang42, Rui Tang42, Jingyao Zeng1,2,#, Jiao Li42,#, Jingfa Xiao1,2,3,*
LncExpDB: Yue Qi1,2,3,#, Zhao Li1,2,3, Lina Ma1,2,3,# (TL)
EWAS Open Platform: Fei Yang1,2,#, Wenting Zong1,2,#, Zhuang Xiong44,#, Demian Kong1,2,3,#, Bixia Tang1,2, Xupeng Chen44, Yaoke Wei1,2,3, Xiangyu Yu1,2,3, Yiran Zhang1,2,3, Rujiao Li1,2,3,# (TL)
MethBank: Mochen Zhang1,2,#, Yaoke Wei1,2,3,#, Huiying Chen1,2,3,36,#, Rujiao Li1,2,3,# (TL)
CROST: Guoliang Wang2,3,13,#, Song Wu1,2,3,#, Hongzhu Qu2,3,13,# (TL), Yiming Bao1,2,3,* (TL), Xiangdong Fang2,3,13,# (TL)
OBIA: Enhui Jin1,2,3,#, Dongli Zhao45,#, Gangao Wu1,2,3,#, Junwei Zhu1,2, Zhonghuang Wang1,2,3, Zhiyao Wei45, Sisi Zhang1,2, Anke Wang1,2, Bixia Tang1,2, Xu Chen1,2, Yanling Sun1,2,#, Zhe Zhang46,#, Wenming Zhao1,2,3,*, Yuanguang Meng45,46,#
OPIA: Yongrong Cao1,2,3,#, Dongmei Tian1,2,#, Shuhui Song1,2,3,*
BIG Search: Dong Zou1,2,# (TL), Zhixiang Yuan1,2, Zhang Zhang1,2,3,*
BioCode: Dong Zou1,2,# (TL), Zhixiang Yuan1,2,#, Tianyi Xu1,2, Zhang Zhang1,2,3,*
BIT: Dong Zou1,2,# (TL), Zhenxian Han1,2,#, Zhixiang Yuan1,2, Zhang Zhang1,2,3,*
OpenLB: Dong Zou1,2,# (TL), Zhang Zhang1,2,3,*
Writing Group: Mochen Zhang1,2,#, Fei Yang1,2,#, Zhuojing Fan1,2, Shuhui Song1,2,3,*, Wenming Zhao1,2,3,*, Jingfa Xiao1,2,3,*, Zhang Zhang1,2,3,*, Yiming Bao1,2,3,*
CNCB–NGDC SubCenters (Listed in alphabetical order by database names)
NGDC-BDV: Xuemei Lu3,9,47,48, Yanan Wang9,47,48
NGDC-TCM: Yuan Yuan49,*, Wei Liu49
NGDC-TGD: Jinyan Huang50
NGDC-NSM: Yanfang Jiang51,52,53, Guoyue Lv51,52
NGDC-MET: Xinyu Liu3,54,55, Guowang Xu3,54,55
NGDC-MOG: Xumin Wang56, JiangYong Qu56
NGDC-PMO: Baisheng Li57,58,59, Chang Zhang57,58,59
NGDC-GMG: Xiaofeng Zou60, Guoxi Zhang60, You Guo61
CNCB–NGDC PARTNERS (Listed in alphabetical order by database names)
Animal-APA: Weiwei Jin62, Jing Gong62
Animal-eRNA: Weiwei Jin62, Jing Gong62
Animal-SNPAtlas: Xiaohui Niu62, Jing Gong62
AnimalTFDB: Wenkang Shen63, Anyuan Guo63
BBCancer: Zhixiang Zuo64, Jian Ren64
CancerSEA: Xinxin Zhang65, Yun Xiao65, Xia Li65
CellMarker: Xinxin Zhang65, Yun Xiao65, Xia Li65
CGDB: Dan Liu66, Yu Xue66
CGGA: Zheng Zhao67, Tao Jiang67
circAtlas: Wanying Wu68, Fangqing Zhao3,68,69, Jinyang Zhang68
CirFunBase: Xianwen Meng70, Ming Chen70
ConsRM: Bowen Song71, Jia Meng72
CPLM: Yujie Gou66, Miaomiao Chen66
dbPSP & THANATOS: Di Peng66, Yu Xue66
DEG & DoriC: Hao Luo73,74,75, Feng Gao73,74,75
DirectRMDB: Jie Jiang71,72, Kunqi Chen76,77
DrLLPS: Xinhe Huang66, Yu Xue66
eLMSG: Wan Liu5, Guoqing Zhang5,6
EPSD: Chi Zhang66, Yu Xue66
EVAtlas: Chunjie Liu63, Anyuan Guo63
EVmiRNA: Guiyan Xie63, Anyuan Guo63
GenTree: Hao Yuan3,78, Tianhan Su3,78, Yong E. Zhang3,78
GTDB: Chenfen Zhou5, Guoqing Zhang5,6
HCL: Yincong Zhou70, Ming Chen70, Guoji Guo79
hTFtarget: Qiong Zhang63, Anyuan Guo63
iEKPD: Shanshan Fu66, Miaoying Zhao66
IMP: Tong Chen80, Yuan Yuan49
iPCD: Dachao Tang66, Yu Xue66
iUUCD: Ming Lei66, Yu Xue66
LeukemiaDB: Mei Luo63, Anyuan Guo63
lnCAR: Yubin Xie64, Jian Ren64
lncRNASNP2: Yaru Miao63, Anyuan Guo63
lncRNASNP3: Anyuan Guo63, Jing Gong62
m5C-Atlas: Jiongming Ma76, Kunqi Chen76
m6A-Atlas: Haokai Ye71,72, Kunqi Chen76
m6A-TSHub: Bowen Song81, Daiyun Huang72
m7GHub: Yuxin Zhang72,81, Bowen Song81
MCA: Yincong Zhou70, Ming Chen70, Guoji Guo79
MiCroKiTS: Di Zhang66, Jianzhen Peng66
miRNASNP: Chunjie Liu63, Anyuan Guo63
msRepDB: Xingyu Liao82,83, Xin Gao82, Jianxin Wang83
ncRNA-eQTL: Jiang Li62, Jing Gong62
Pancan-mnvQTL: Xiaohui Niu62, Jing Gong62
PEA: Guiyan Xie63, Anyuan Guo63
PceRBase: Chunhui Yuan70, Ming Chen70
PlantRegMap: Dechang Yang84, Feng Tian84, Ge Gao84
Plant-ImputeDB: Xiaohui Niu62, Jing Gong62
PncStres: Wenyi Wu70, Ming Chen70
PTMD: Cheng Han66, Yu Xue66
RhesusBase: Juntian Qi85, Ni A. An85, Chuan-Yun Li85
RMDisease: Xuan Wang72, Zhen Wei72,86
RMVar: XiaoTong Luo64, Jian Ren64
ScRAPdb: Jiaxing Yue87, Zepu Miao87
SEECancer: Xinxin Zhang65, Yun Xiao65, Xia Li65
SEGreg: Qing Tang63, Anyuan Guo63
SNP2APA: Anyuan Guo63, Jing Gong62
THANATOS: Zihao Feng66, Yu Xue66
VFDB: Bo Liu88, Jian Yang88
WERAM: Chenyu Yang66, Leming Xiao66
ZCURVE_CoVdb: Hao Luo73,74,75, Feng Gao73,74,75
1. National Genomics Data Center, China National Center for Bioinformation, Beijing 100 101, China.
2. Beijing Institute of Genomics, Chinese Academy of Sciences, Beijing 100 101, China.
3. University of Chinese Academy of Sciences, Beijing 100 049, China.
4. Key Laboratory of Epigenetic Regulation and Intervention, Institute of Biophysics, Chinese Academy of Sciences, Beijing 100 101, China.
5. National Genomics Data Center & Bio-Med Big Data Center, Shanghai Institute of Nutrition and Health, University of the Chinese Academy of Sciences, Chinese Academy of Sciences, Shanghai 200 031, China.
6. Shanghai Sixth People’s Hospital, Shanghai 200 233, China.
7. Guangzhou Laboratory, Guangzhou 510 005, China.
8. Hangzhou Institute for Advanced Study, University of Chinese Academy of Sciences, Hangzhou 310 024, China.
9. National Key Laboratory of Genetic Evolution and Animal Model, Kunming Institute of Zoology, Chinese Academy of Sciences, Kunming 650 201, China.
10. Yunnan Key Laboratory of Integrative Anthropology, Kunming 650 201, China.
11. National Laboratory of Biomacromolecules, CAS Center for Excellence in Biomacromolecules, Institute of Biophysics, Chinese Academy of Sciences, Beijing 100 101, China.
12. Shandong First Medical University & Shandong Academy of Medical Sciences, Taian 271 016, China.
13. China National Center for Bioinformation, Beijing 100 101, China.
14. School of Advanced Interdisciplinary Science, University of Chinese Academy of Sciences, Beijing 101 408, China.
15. School of Life Science and Bio-Pharmaceutics, Shenyang Pharmaceutical University, Shenyang 110 016, China.
16. Key Laboratory for Molecular Enzymology and Engineering of Ministry of Education, School of Life Sciences, Jilin University, Changchun 130 021, China.
17. Key Laboratory of Viral Pathogenesis & Infection Prevention and Control (Jinan University), Ministry of Education, Jinan University, Guangzhou 510 632, China.
18. State Key Laboratory of Pathogenesis, Prevention and Treatment of High Incidence Diseases in Central Asia, Clinical Medicine Institute, The First Affiliated Hospital of Xinjiang Medical University, Urumqi 830 013, China.
19. Putuo People’s Hospital, Shanghai Key Laboratory of Signaling and Disease Research, School of Life Sciences and Technology, Tongji University, Shanghai 200 092, China.
20. Department of Thoracic Surgery and Institute of Thoracic Oncology, West China Hospital of Sichuan University, Chengdu 610 041, China.
21. Western China Collaborative Innovation Center for Early Diagnosis and Multidisciplinary Therapy of Lung Cancer, Sichuan University, Chengdu 610 041, China.
22. Beijing Life Science Academy, Lutuan East Road, Beijing 102 200, China.
23. Key Laboratory of Tobacco Biological Effects, Lutuan East Road, Beijing 102 200, China.
24. Institute for Stem Cell and Regeneration, Chinese Academy of Sciences, Beijing 100 101, China.
25. College of Future Technology College, University of Chinese Academy of Sciences, Beijing 100 049, China.
26. School of Traditional Chinese Medicine, Beijing University of Chinese Medicine, Beijing 100 029, China.
27. School of Life Sciences, Beijing University of Chinese Medicine, Beijing 100 029, China.
28. Center for Excellence in Animal Evolution and Genetics, Chinese Academy of Sciences, Kunming 650 201, China.
29. Department of Physiology, School of Basic Medicine, Tongji Medical College, Huazhong University of Science and Technology, Wuhan 430 030, Hubei, China.
30. Hubei Key Laboratory of Drug Target Research and Pharmacodynamic Evaluation, Huazhong University of Science and Technology, Wuhan 430 030, Hubei, China.
31. Department of Cardiovascular Surgery, Union Hospital, Tongji Medical College, Huazhong University of Science and Technology, Wuhan 430 022, China.
32. Key Laboratory of Vascular Aging, Ministry of Education, Tongji Hospital, Tongji Medical College, Huazhong University of Science and Technology, Wuhan 430 030, Hubei, China.
33. Institute for Data-Driven Tumor Immunology, Chongqing Medical University, Chongqing 400 016, China.
34. National Clinical Laboratory on Tuberculosis, Beijing Key Laboratory for Drug-Resistant Tuberculosis Research, Beijing Chest Hospital, Capital Medical University, Beijing Tuberculosis and Thoracic Tumor Institute, Beijing 101 149, China.
35. Key Laboratory of Major Diseases in Children, Ministry of Education, National Key Discipline of Pediatrics (Capital Medical University), Beijing Key Laboratory of Pediatric Respiratory Infection Diseases, Beijing Pediatric Research Institute, Beijing Children’s Hospital, Capital Medical University, Beijing 100 020, China.
36. Sino-Danish College, University of Chinese Academy of Sciences, Beijing 100 049, China.
37. The Shanghai Tenth People’s Hospital, School of Life Sciences and Technology, Tongji University, Shanghai 200 072, China.
38. National Center, Children’s Hospital, Zhejiang University School of Medicine, National Clinical Research Center for Child Health, Hangzhou 310 052, China.
39. CAS-Key Laboratory of Synthetic Biology, CAS Center for Excellence in Molecular Plant Sciences, Chinese Academy of Sciences, Shanghai 200 032, China.
40. Zhangke (Shanghai) Jianwu Biotech Services Co., Ltd, Shanghai 200 120, China.
41. State Key Laboratory of Genetic Engineering, Fudan Microbiome Center, School of Life Sciences, Fudan University, Shanghai 200 433, China.
42. Institute of Medical Information, Chinese Academy of Medical Sciences and Peking Union Medical College, Beijing 100 020, China.
43. Department of Computer Science and Technology, Tsinghua University, Beijing 100 084, China.
44. Interdisciplinary Institute for Medical Engineering, Fuzhou University, Fuzhou 350 002, China.
45. Chinese People’s Liberation Army (PLA) Medical School, Beijing 100 853, China.
46. Department of Obstetrics and Gynecology, Seventh Medical Center of Chinese PLA General Hospital, Beijing 100 700, China.
47. Biodiversity Data Center of Kunming Institute of Zoology, Chinese Academy of Sciences, Kunming 650 201, China.
48. Yunnan Key Laboratory of Biodiversity Information, Kunming Institute of Zoology, Chinese Academy of Sciences, Kunming 650 201, China.
49. Experimental Research Center, China Academy of Chinese Medical Sciences, Beijing 100 700, China.
50. Biomedical big data center, the First Affiliated Hospital, Zhejiang University School of Medicine, Zhejiang 310 003, China.
51. Northeast Subcenter of Medical Genomics, in The First Hospital of jilin University, Changchun, China.
52. Department of Hepatobiliary and Pancreatic Surgery, General Surgery Center, First Hospital of Jilin University, Changchun, China.
53. Genetic Diagnosis Center, First Hospital of Jilin University, Changchun, China.
54. State Key Laboratory of Medical Proteomics, Dalian Institute of Chemical Physics, Chinese Academy of Sciences, Dalian 116 023, China.
55. Liaoning Province Key Laboratory of Metabolomics, Dalian 116 023, China.
56. School of Life Sciences, Yantai University, Yantai 264 005, China.
57. Guangdong Provincial Center for Disease Control and Prevention, Guangzhou 511 400, China.
58. Guangdong Provincial Key Laboratory of Pathogen Detection for Emerging Infectious Disease Response, Guangzhou 511 400, China.
59. Guangdong Workstation for Emerging Infectious Disease Control and Prevention, Chinese Academy of Medical Sciences, Guangzhou 511 400, China.
60. Institute of Urology, First Affiliated Hospital of Gannan Medical University, Ganzhou 341 000, China.
61. Medical Big Data and Bioinformatics Research Centre, First Affiliated Hospital of Gannan Medical University, Ganzhou 341 000, China.
62. Hubei Key Laboratory of Agricultural Bioinformatics, College of Informatics, Huazhong Agricultural University, Wuhan 430 070, China.
63. Department of Thoracic Surgery, West China Biomedical Big Data Center, West China Hospital, Sichuan University, Chengdu 610 041, China.
64. State Key Laboratory of Oncology in South China, Cancer Center, Collaborative Innovation Center for Cancer Medicine, School of Life Sciences, Sun Yat-sen University, Guangzhou 510 060, China.
65. College of Bioinformatics Science and Technology, Harbin Medical University, Harbin 150 081, China.
66. MOE Key Laboratory of Molecular Biophysics, Hubei Bioinformatics and Molecular Imaging Key Laboratory, College of Life Science and Technology, Huazhong University of Science and Technology, Wuhan 430 074, China.
67. Beijing Neurosurgical Institute, Capital Medical University, Beijing 100 070, China.
68. Beijing Institutes of Life Science, Chinese Academy of Sciences, Beijing 100 101, China.
69. Key Laboratory of Systems Biology, Hangzhou Institute for Advanced Study, University of Chinese Academy of Sciences, Chinese Academy of Sciences, Hangzhou 310 024, China.
70. Department of Bioinformatics, College of Life Sciences, The First Affiliated Hospital, School of Medicine, Zhejiang University, Hangzhou 310 058, China.
71. Institute of Systems, Molecular and Integrative Biology, University of Liverpool, Liverpool L69 7ZB, Liverpool, UK.
72. Department of Biological Sciences, Xi’an Jiaotong-Liverpool University, Suzhou 215 123, China.
73. Department of Physics, School of Science, Tianjin University, Tianjin 300 072, China.
74. Frontiers Science Center for Synthetic Biology and Key Laboratory of Systems Bioengineering (Ministry of Education), Tianjin University, Tianjin 300 072, China.
75. SynBio Research Platform, Collaborative Innovation Center of Chemical Science and Engineering (Tianjin), Tianjin 300 072, China.
76. Key Laboratory of Gastrointestinal Cancer (Fujian Medical University), Ministry of Education, Fuzhou, China.
77. Fujian Key Laboratory of Tumor Microbiology, Department of Medical Microbiology, School of Basic Medical Sciences, Fujian Medical University, Fuzhou 350 004, China.
78. Key Laboratory of Zoological Systematics and Evolution, Institute of Zoology, Chinese Academy of Sciences, Beijing, China.
79. Center for Stem Cell and Regenerative Medicine, Zhejiang University School of Medicine, Hangzhou, China.
80. State Key Laboratory for Quality Ensurance and Sustainable Use of Dao-di Herbs, National Resource Center for Chinese Materia Medica, China Academy of Chinese Medical Sciences, Beijing 100 000, China.
81. Department of Public Health, School of Medicine, Nanjing University of Chinese Medicine, Nanjing 210 023, China.
82. Computational Bioscience Research Center (CBRC), Computer, Electrical and Mathematical Sciences and Engineering Division, King Abdullah University of Science and Technology (KAUST), Thuwal 23 955, Saudi Arabia.
83. Hunan Provincial Key Lab on Bioinformatics, School of Computer Science and Engineering, Central South University, Changsha 410 083, China.
84. State Key Laboratory of Protein and Plant Gene Research, School of Life Sciences, Biomedical Pioneering Innovative Center (BIOPIC) & Beijing Advanced Innovation Center for Genomics (ICG), Center for Bioinformatics (CBI), Peking University, Beijing 100 871, China.
85. State Key Laboratory of Protein and Plant Gene Research, Laboratory of Bioinformatics and Genomic Medicine, Institute of Molecular Medicine, College of Future Technology, Peking University, Beijing, China.
86. Institute of Life Course and Medical Sciences, University of Liverpool, Liverpool L7 8TX, UK.
87. Sun Yat-sen University Cancer Center, Guangzhou 510 060, China.
88. NHC Key Laboratory of Systems Biology of Pathogens, Institute of Pathogen Biology, Chinese Academy of Medical Sciences & Peking Union Medical College, Beijing, China.
Contributor Information
CNCB–NGDC Members and Partners:
Yiming Bao, Zhang Zhang, Wenming Zhao, Jingfa Xiao, Shuhui Song, Shunmin He, Guoqing Zhang, Yixue Li, Guoping Zhao, Runsheng Chen, Yibo Wang, Weijie Zhang, Xiaoning Chen, Yanling Sun, Bixia Tang, Yu Zhang, Kai Liu, Bing Su, Yaoxi He, Tingrui Song, Yirong Shi, Yanyan Li, Di Hao, Kaixin Zhan, Tao Xu, Wenyan Lei, Cuidan Li, Hengyu Zhou, Zhi Nie, Anke Wang, Pan Li, Peihan Wang, Zhuojing Fan, Rongxi Zhu, Haoyu Cheng, Yuxian Guo, Liya Yue, Xiaoyuan Jiang, Renjun Gao, Yongjie Sheng, Haitao Niu, Tuohetaerbaike Bahetibieke, Wenbao Zhang, Fei Chen, Jiayue Meng, Mengyao Han, Yuwei Huang, Liyun Yuan, RuiKun Xue, Jingyao Zeng, Xini Meng, Qifei Wang, Yiwen Hu, Yulan Deng, Na Ai, Zheng Huang, Yun Li, Yang Yuan, Guochao Li, Lan Jiang, Ting Li, Ling Zuo, Jianxin Chen, Qianqian Peng, Sijia Wang, Sicheng Wu, Hao Jiang, Hailong Kang, Jiawei Shi, Nianguo Dong, Ximiao He, Fei Yang, Shuai Jiang, Zhenxian Han, Xue Bai, Dong Zou, Sisi Zhang, Yi Wang, Zhijian Duan, Lun Li, Wenjing Sun, Sijia Zhang, Quan Luo, Jinying Han, Hairong Huang, Adong Shen, Jing Wang, Hao Wen, Congfan Bu, Xiaotong Ji, Qiheng Qian, Hao Zhang, Qingyun Cai, Kaiwen Zhang, Xiaoqing Jiang, Mingkun Li, Hong Luo, Zishan Wu, Siwei Ren, Haixia Xie, Zhixiang Yuan, Dongmei Tian, Demian Kong, Shaoqi Bei, Yueyue Wu, Lei Liu, Suqi Cao, Weili Lin, Ruixin Zhu, Dingfeng Wu, Yuyan Meng, Wan Liu, Pingping Wang, Xinhao Zhuang, Xing Yan, Zhihua Zhou, Wenxing Gao, Qiang Li, Na Jiao, Yiyun Liu, Lili Tian, Wei Zhao, Wenting Zong, Xinchang Zheng, Xu Chen, Tingting Chen, Xiaolong Zhang, Yubo Zhou, Junwei Zhu, Bing Xu, Lili Dong, Caixia Yu, Wenjie Li, Shuang Zhai, Yubin Sun, Qiancheng Chen, Yanqing Wang, Xuetong Zhao, Shaosen Zhang, Jinbiao Wang, Yuhao Zeng, Zheng Luo, Yiran Zhan, Zihan Wang, Xi Zhao, Yuxi Liu, Lina Ma, Yingke Ma, Meili Chen, Tongtong Zhu, Wenzhuo Cheng, Yuan Chu, Ming Chen, Tianyi Xu, Hao Gao, Si Zheng, Jialin Mai, Jinbei Wang, Rui Tang, Jiao Li, Yue Qi, Zhao Li, Zhuang Xiong, Xupeng Chen, Yaoke Wei, Xiangyu Yu, Rujiao Li, Mochen Zhang, Huiying Chen, Guoliang Wang, Song Wu, Hongzhu Qu, Xiangdong Fang, Enhui Jin, Dongli Zhao, Gangao Wu, Zhonghuang Wang, Zhiyao Wei, Zhe Zhang, Yuanguang Meng, Yongrong Cao, Xuemei Lu, Yanan Wang, Wei Liu, Jinyan Huang, Yanfang Jiang, Guoyue Lv, Xinyu Liu, Guowang Xu, Xumin Wang, JiangYong Qu, Baisheng Li, Chang Zhang, Xiaofeng Zou, Guoxi Zhang, You Guo, Weiwei Jin, Jing Gong, Xiaohui Niu, Wenkang Shen, Anyuan Guo, Zhixiang Zuo, Jian Ren, Xinxin Zhang, Yun Xiao, Xia Li, Dan Liu, Yu Xue, Zheng Zhao, Tao Jiang, Wanying Wu, Fangqing Zhao, Jinyang Zhang, Xianwen Meng, Ming Chen, Bowen Song, Jia Meng, Yujie Gou, Miaomiao Chen, Di Peng, Hao Luo, Feng Gao, Jie Jiang, Kunqi Chen, Xinhe Huang, Chi Zhang, Chunjie Liu, Guiyan Xie, Hao Yuan, Tianhan Su, Yong E Zhang, Chenfen Zhou, Yincong Zhou, Guoji Guo, Qiong Zhang, Shanshan Fu, Miaoying Zhao, Tong Chen, Yuan Yuan, Dachao Tang, Ming Lei, Mei Luo, Yubin Xie, Yaru Miao, Jiongming Ma, Haokai Ye, Bowen Song, Daiyun Huang, Yuxin Zhang, Di Zhang, Jianzhen Peng, Xingyu Liao, Xin Gao, Jianxin Wang, Jiang Li, Chunhui Yuan, Dechang Yang, Feng Tian, Ge Gao, Wenyi Wu, Cheng Han, Juntian Qi, Ni A An, Chuan-Yun Li, Xuan Wang, Zhen Wei, XiaoTong Luo, Jiaxing Yue, Zepu Miao, Qing Tang, Zihao Feng, Bo Liu, Jian Yang, Chenyu Yang, and Leming Xiao
Conflict of interest statement
All authors have confirmed that there are no conflicts of interest to disclose.
Funding
Strategic Priority Research Program of the Chinese Academy of Sciences [XDA0460200, XDA0460100, XDA0460400, XDA0450100]; Informatization Plan of Chinese Academy of Sciences [CAS-WX2024GC-06]; National Key Research & Development Program of China [2023YFC2605700, 2021YFF0703700, 2021YFF0703702, 2021YFF0703704, 2021YFF0704500, 2020YFA0907001, 2024YFA1307701, 2025YFF1207901, 2021YFF0703701, 2023YFC2604402, 2023YFC2605702, 2024YFC2311300]; National Natural Science Foundation of China [92374201, 32270019, T2425005, 31871328, 32100520, 32170669, 32100506, 32100511, 32270718, 32030021, 32300542, 32300468, 32470608, 32170678]; Biological Breeding-National Science and Technology Major Project [2022ZD0401701]; The Alliance of National and International Science Organizations for the Belt and Road Regions [ANSO-CR-KP-2022-09, ANSO-PA-2023-07]; Genomics Data Center Operation and Maintenance of Chinese Academy of Sciences [CAS-WX2022SDC-XK05]; Outstanding Member of the Youth Innovation Promotion Association, CAS [Y2023027]; The International Partnership Program of the Chinese Academy of Sciences [161GJHZ2022002MI]; the Key Technology Talent Program of the Chinese Academy of Sciences [2021000014, 2023000007]; The Open Biodiversity and Health Big Data Programme of IUBS; the Youth Innovation Promotion Association of the Chinese Academy of Sciences [2022098]. Funding to pay the Open Access publication charges for this article was provided by Strategic Priority Research Program of the Chinese Academy of Sciences.
Data availability
All resources and services are publicly available on the home page of CNCB–NGDC (https://ngdc.cncb.ac.cn).
References
- 1. Bao Y, Xue Y.. From BIG data center to China National Center for Bioinformation. Genomics Proteomics Bioinformatics. 2023;21:900–3. 10.1016/j.gpb.2023.10.0010. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 2. Piwecka M, Rajewsky N, Rybak-Wolf A.. Single-cell and spatial transcriptomics: deciphering brain complexity in health and disease. Nat Rev Neurol. 2023;19:346–62. 10.1038/s41582-023-00809-y. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 3. Vandereyken K, Sifrim A, Thienpont Bet al. Methods and applications for single-cell and spatial multi-omics. Nat Rev Genet. 2023;24:494–515. 10.1038/s41576-023-00580-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 4. Gulati GS, D'Silva JP, Liu Yet al. Profiling cell identity and tissue architecture with single-cell and spatial transcriptomics. Nat Rev Mol Cell Biol. 2025;26:11–31. 10.1038/s41580-024-00768-2. [DOI] [PubMed] [Google Scholar]
- 5. All of Us Research Program Genomics, I . Genomic data in the all of us research program. Nature. 2024;627:340–6. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 6. Cook MB, Sanderson SC, Deanfield JEet al. Our future health: a unique global resource for discovery and translational research. Nat Med. 2025;31:728–30. 10.1038/s41591-024-03438-0. [DOI] [PubMed] [Google Scholar]
- 7. Rozenblatt-Rosen O, Stubbington MJT, Regev Aet al. The Human Cell Atlas: from vision to reality. Nature. 2017;550:451–3. 10.1038/550451a. [DOI] [PubMed] [Google Scholar]
- 8. Lewin HA, Robinson GE, Kress WJet al. Earth BioGenome project: sequencing life for the future of life. Proc Natl Acad Sci USA. 2018;115:4325–33. 10.1073/pnas.1720115115. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 9. Papatheodorou I, Moreno P, Manning Jet al. Expression atlas update: from tissues to single cells. Nucleic Acids Res. 2020;48:D77–83. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 10. Bycroft C, Freeman C, Petkova Det al. The UK Biobank resource with deep phenotyping and genomic data. Nature. 2018;562:203–9. 10.1038/s41586-018-0579-z. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 11. Bhattacharya S, Andorf S, Gomes Let al. ImmPort: disseminating data to the public for the future of immunology. Immunol Res. 2014;58:234–9. 10.1007/s12026-014-8516-1. [DOI] [PubMed] [Google Scholar]
- 12. Ju LF, Xu HJ, Yang YGet al. Omics views of mechanisms for cell fate determination in early mammalian development. Genomics Proteomics Bioinformatics. 2023;21:950–61. 10.1016/j.gpb.2023.03.001. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 13. Geng F, Zhang X, Ma Jet al. Genome assembly and winged fruit gene regulation of Chinese wingnut: insights from genomic and transcriptomic analyses. Genomics Proteomics Bioinformatics. 2025;22:1–16. 10.1093/gpbjnl/qzae087. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 14. Chen M, Lv J, Zhang Qet al. Spatial transcriptomics unveils the blueprint of mammalian lung development. Genomics Proteomics Bioinformatics. 2025.; 10.1093/gpbjnl/qzaf053. [DOI] [PubMed] [Google Scholar]
- 15. Jardine L, Webb S, Goh Iet al. Blood and immune development in human fetal bone marrow and Down syndrome. Nature. 2021;598:327–31. 10.1038/s41586-021-03929-x. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 16. Du T, Gao C, Lu Set al. Differential transcriptomic landscapes of SARS-CoV-2 variants in multiple organs from infected rhesus macaques. Genomics Proteomics Bioinformatics. 2023;21:1014–29. 10.1016/j.gpb.2023.06.002. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 17. Zhou JT, Xu Y, Liu XHet al. Single-cell RNA-seq reveals the inhibitory effect of methamphetamine on liver immunity with the involvement of dopamine receptor D1. Genomics Proteomics Bioinformatics. 2024;22:1–17. 10.1093/gpbjnl/qzae060. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 18. Yang Y, Lu X, Liu Net al. Metformin decelerates aging clock in male monkeys. Cell. 2024;187:6358–78.e29. 10.1016/j.cell.2024.08.021. [DOI] [PubMed] [Google Scholar]
- 19. Zhao X, Yang A, Ding Jet al. Regulatory genomic circuitry of brain age by integrative functional genomic analyses. Genomics Proteomics Bioinformatics. 2025. 10.1093/gpbjnl/qzaf064. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 20. Zhang R, Hu M, Liu Yet al. Integrative omics uncovers low tumorous magnesium content as a driver factor of colorectal cancer. Genomics Proteomics Bioinformatics. 2024;22:qzae053. 10.1093/gpbjnl/qzae053. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 21. Li Y, van den Berg EH, Kurilshikov Aet al. Genome-wide studies reveal genetic risk factors for hepatic fat content. Genomics Proteomics Bioinformatics. 2024;22:1–12. 10.1093/gpbjnl/qzae031. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 22. Wang X, Zuo L, Yu Yet al. Clonal hematopoietic mutations in plasma cell disorders: clinical subgroups and shared pathogenesis. Genomics Proteomics Bioinformatics. 2025;23:qzaf027. 10.1093/gpbjnl/qzaf027. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 23. Sammut SJ, Crispin-Ortuzar M, Chin SFet al. Multi-omic machine learning predictor of breast cancer therapy response. Nature. 2022;601:623–9. 10.1038/s41586-021-04278-5. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 24. Huo Z, Duan Y, Zhan Det al. Proteomic stratification of prognosis and treatment options for small cell lung cancer. Genomics Proteomics Bioinformatics. 2024;22:1–12. 10.1093/gpbjnl/qzae033. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 25. Li W, Gao L, Yi Xet al. Patient assessment and therapy planning based on homologous recombination repair deficiency. Genomics Proteomics Bioinformatics. 2023;21:962–75. 10.1016/j.gpb.2023.02.004. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 26. Zhang Z. Expanding bioinformatics: toward a paradigm shift from data to theory. Fundamental Research. 2024. 10.1016/j.fmre.2024.11.019. [DOI] [Google Scholar]
- 27. Jumper J, Evans R, Pritzel Aet al. Highly accurate protein structure prediction with AlphaFold. Nature. 2021;596:583–9. 10.1038/s41586-021-03819-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 28. Theodoris CV, Xiao L, Chopra Aet al. Transfer learning enables predictions in network biology. Nature. 2023;618:616–24. 10.1038/s41586-023-06139-9. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 29. Cui H, Wang C, Maan Het al. scGPT: toward building a foundation model for single-cell multi-omics using generative AI. Nat Methods. 2024;21:1470–80. 10.1038/s41592-024-02201-0. [DOI] [PubMed] [Google Scholar]
- 30. Wilkinson MD, Dumontier M, Aalbersberg IJet al. The FAIR guiding principles for scientific data management and stewardship. Sci Data. 2016;3:160018. 10.1038/sdata.2016.18. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 31. BIG Data Center Members . The BIG data center: from deposition to integration to translation. Nucleic Acids Res. 2017;45:D18–24. 10.1093/nar/gkw1060. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 32. BIG Data Center Members . Database resources of the BIG data center in 2018. Nucleic Acids Res. 2018;46:D14–20. 10.1093/nar/gkx897. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 33. BIG Data Center Members . Database resources of the BIG data center in 2019. Nucleic Acids Res. 2019;47:D8–D14. 10.1093/nar/gky993. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 34. National Genomics Data Center Members and Partners . Database resources of the national genomics data center in 2020. Nucleic Acids Res. 2020;48:D24–33. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 35. CNCB–NGDC Members and Partners . Database resources of the National Genomics Data Center, China National Center for Bioinformation in 2021. Nucleic Acids Res. 2021;49:D18–28. 10.1093/nar/gkaa1022. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 36. CNCB–NGDC Members and Partners . Database resources of the National Genomics Data Center, China National Center for Bioinformation in 2022. Nucleic Acids Res. 2022;50:D27–38. 10.1093/nar/gkab951. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 37. CNCB–NGDC Members and Partners . Database resources of the National Genomics Data Center, China National Center for Bioinformation in 2023. Nucleic Acids Res. 2023;51:D18–28. 10.1093/nar/gkac1073. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 38. CNCB–NGDC Members and Partners . Database resources of the National Genomics Data Center, China National Center for Bioinformation in 2024. Nucleic Acids Res. 2024;52:D18–D32. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 39. CNCB–NGDC Members and Partners . Database resources of the National Genomics Data Center, China National Center for Bioinformation in 2025. Nucleic Acids Res. 2025;53:D30–44. 10.1093/nar/gkae978. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 40. Ma L, Zou D, Liu Let al. Database commons: a catalog of worldwide biological databases. Genomics Proteomics Bioinformatics. 2023;21:1054–8. 10.1016/j.gpb.2022.12.004. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 41. Sherman BT, Hao M, Qiu Jet al. DAVID: a web server for functional enrichment analysis and functional annotation of gene lists (2021 update). Nucleic Acids Res. 2022;50:W216–21. 10.1093/nar/gkac194. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 42. Kanehisa M, Furumichi M, Sato Yet al. KEGG: biological systems database as a model of the real world. Nucleic Acids Res. 2025;53:D672–7. 10.1093/nar/gkae909. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 43. Cerami E, Gao J, Dogrusoz Uet al. The cBio cancer genomics portal: an open platform for exploring multidimensional cancer genomics data. Cancer Discov. 2012;2:401–4. 10.1158/2159-8290.CD-12-0095. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 44. Szklarczyk D, Nastou K, Koutrouli Met al. The STRING database in 2025: protein networks with directionality of regulation. Nucleic Acids Res. 2025;53:D730–7. 10.1093/nar/gkae1113. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 45. Varadi M, Bertoni D, Magana Pet al. AlphaFold protein structure database in 2024: providing structure coverage for over 214 million protein sequences. Nucleic Acids Res. 2024;52:D368–75. 10.1093/nar/gkad1011. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 46. UniProt C . UniProt: the universal protein knowledgebase in 2025. Nucleic Acids Res. 2025;53:D609–17. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 47. Glockner FO, Yilmaz P, Quast Cet al. 25 years of serving the community with ribosomal RNA gene reference databases and tools. J Biotechnol. 2017;261:169–76. 10.1016/j.jbiotec.2017.06.1198. [DOI] [PubMed] [Google Scholar]
- 48. Consortium EP, Moore JE, Purcaro MJet al. Expanded encyclopaedias of DNA elements in the human and mouse genomes. Nature. 2020;583:699–710. 10.1038/s41586-020-2493-4. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 49. Karczewski KJ, Francioli LC, Tiao Get al. The mutational constraint spectrum quantified from variation in 141,456 humans. Nature. 2020;581:434–43. 10.1038/s41586-020-2308-7. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 50. Fairley S, Lowy-Gallego E, Perry Eet al. The international genome sample resource (IGSR) collection of open human genomic variation resources. Nucleic Acids Res. 2020;48:D941–7. 10.1093/nar/gkz836. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 51. Fiorini N, Lipman DJ, Lu Z.. Towards PubMed 2.0. eLife. 2017;6:e28801. 10.7554/eLife.28801. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 52. Zhang S, Chen X, Jin Eet al. The GSA family in 2025: a broadened sharing platform for multi-omics and multimodal data. Genomics Proteomics Bioinformatics. 2025;23:qzaf072. 10.1093/gpbjnl/qzaf072. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 53. Chen T, Chen X, Zhang Set al. The genome sequence archive family: toward explosive data growth and diverse data types. Genomics Proteomics Bioinformatics. 2021;19:578–83. 10.1016/j.gpb.2021.08.001. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 54. Wang Y, Song F, Zhu Jet al. GSA: genome sequence archive. Genomics Proteomics Bioinformatics. 2017;15:14–8. 10.1016/j.gpb.2017.01.001. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 55. Zhang SS, Chen X, Chen TTet al. GSA-Human: genome sequence archive for human. Yi Chuan. 2021;43:988–93. [DOI] [PubMed] [Google Scholar]
- 56. Bu C, Zheng X, Zhao Xet al. GenBase: a nucleotide sequence database. Genomics Proteomics Bioinformatics. 2024;22:qzae047. 10.1093/gpbjnl/qzae047. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 57. Arita M, Karsch-Mizrachi I, Cochrane G.. The international nucleotide sequence database collaboration. Nucleic Acids Res. 2021;49:D121–4. 10.1093/nar/gkaa967. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 58. Goldfarb T, Kodali VK, Pujar Set al. NCBI RefSeq: reference sequence standards through 25 years of curation and annotation. Nucleic Acids Res. 2025;53:D243–57. 10.1093/nar/gkae1038. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 59. Ma Y, Zhao X, Jia Yet al. The updated Genome Warehouse: enhancing data value, security, and usability to address data expansion. Genomics Proteomics Bioinformatics. 2025;23:1–6. 10.1093/gpbjnl/qzaf010. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 60. Feldgarden M, Brover V, Gonzalez-Escalona Net al. AMRFinderPlus and the reference gene catalog facilitate examination of the genomic links among antimicrobial resistance, stress response, and virulence. Sci Rep. 2021;11:12728. 10.1038/s41598-021-91456-0. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 61. Zhang W, Chen X, Sun Yet al. HiLand Resource: a comprehensive database of highland human populations. Genomics Proteomics Bioinformatics. 2025. 10.1093/gpbinl/azaf083. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 62. Luo H, Bai X, Wu Zet al. SugarcaneOmics: an integrative multi-omics platform for sugarcane research. Plant Commun. 2025;6:101489. 10.1016/j.xplc.2025.101489. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 63. Li C, Tian D, Tang Bet al. Genome Variation Map: a worldwide collection of genome variations across multiple species. Nucleic Acids Res. 2021;49:D1186–91. 10.1093/nar/gkaa1005. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 64. Song S, Tian D, Li Cet al. Genome Variation Map: a data repository of genome variations in BIG data center. Nucleic Acids Res. 2018;46:D944–9. 10.1093/nar/gkx986. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 65. Song T, Shi Y, Li Yet al. TOAnnoPriDB: an integrative database for trans-omic annotation and prioritization of non-coding variants across human genome. Sci Bull (Beijing). 2025;70:1757–60. 10.1016/j.scib.2024.12.030. [DOI] [PubMed] [Google Scholar]
- 66. Zhang P, Luo H, Li Yet al. NyuWa genome resource: a deep whole-genome sequencing-based variation profile and reference panel for the Chinese population. Cell Reports. 2021;37:110017. 10.1016/j.celrep.2021.110017. [DOI] [PubMed] [Google Scholar]
- 67. Niu Y, Teng X, Zhou Het al. Characterizing mobile element insertions in 5675 genomes. Nucleic Acids Res. 2022;50:2493–508. 10.1093/nar/gkac128. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 68. Shi Y, Niu Y, Zhang Pet al. Characterization of genome-wide STR variation in 6487 human genomes. Nat Commun. 2023;14:2092. 10.1038/s41467-023-37690-8. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 69. Liu S, Luo H, Zhang Pet al. Adaptive selection of cis-regulatory elements in the Han Chinese. Mol Biol Evol. 2024;41:msae034. 10.1093/molbev/msae034. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 70. Luo H, Zhang P, Zhang Wet al. Recent positive selection signatures reveal phenotypic evolution in the Han Chinese population. Science Bulletin. 2023;68:2391–404. 10.1016/j.scib.2023.08.027. [DOI] [PubMed] [Google Scholar]
- 71. Zhao L, Wang J, Li Yet al. NONCODEV6: an updated database dedicated to long non-coding RNA annotation in both animals and plants. Nucleic Acids Res. 2021;49:D165–71. 10.1093/nar/gkaa1046. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 72. Zheng Y, Luo H, Teng Xet al. NPInter v5.0: ncRNA interaction database in a new era. Nucleic Acids Res. 2023;51:D232–9. 10.1093/nar/gkac1002. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 73. Wang J, Shi Y, Zhou Het al. piRBase: integrating piRNA annotation in all aspects. Nucleic Acids Res. 2022;50:D265–72. 10.1093/nar/gkab1012. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 74. Li Y, Zhou H, Chen Xet al. SmProt: a reliable repository with comprehensive annotation of small proteins identified from ribosome profiling. Genomics Proteomics Bioinformatics. 2021;19:602–10. 10.1016/j.gpb.2021.09.002. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 75. Chen X, Hao Y, Cui Yet al. LncVar: a database of genetic variation associated with long non-coding genes. Bioinformatics. 2017;33:112–8. 10.1093/bioinformatics/btw581. [DOI] [PubMed] [Google Scholar]
- 76. Zhang Y, Zou D, Zhu Tet al. Gene Expression Nebulas (GEN): a comprehensive data portal integrating transcriptomic profiles across multiple species at both bulk and single-cell levels. Nucleic Acids Res. 2022;50:D1016–24. 10.1093/nar/gkab878. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 77. Li Z, Liu L, Jiang Set al. LncExpDB: an expression database of human long non-coding RNAs. Nucleic Acids Res. 2021;49:D962–8. 10.1093/nar/gkaa850. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 78. Li Z, Liu L, Feng Cet al. LncBook 2.0: integrating human long non-coding RNAs with multi-omics annotations. Nucleic Acids Res. 2023;51:D186–91. 10.1093/nar/gkac999. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 79. Wang G, Wu S, Xiong Zet al. CROST: a comprehensive repository of spatial transcriptomics. Nucleic Acids Res. 2024;52:D882–90. 10.1093/nar/gkad782. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 80. Lu M, Zhang Y, Yang Fet al. TWAS atlas: a curated knowledgebase of transcriptome-wide association studies. Nucleic Acids Res. 2023;51:D1179–87. 10.1093/nar/gkac821. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 81. Xiong Z, Yang F, Li Met al. EWAS open platform: integrated data, knowledge and toolkit for epigenome-wide association study. Nucleic Acids Res. 2022;50:D1004–9. 10.1093/nar/gkab972. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 82. Xiong Z, Li M, Ma Yet al. GMQN: a reference-based method for correcting batch effects and probe bias in HumanMethylation BeadChip. Front. Genet. 2021;12:810985. 10.3389/fgene.2021.810985. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 83. Xiong Z, Li M, Yang Fet al. EWAS Data Hub: a resource of DNA methylation array data and metadata. Nucleic Acids Res. 2020;48:D890–5. 10.1093/nar/gkz840. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 84. Li M, Zou D, Li Zet al. EWAS Atlas: a curated knowledgebase of epigenome-wide association studies. Nucleic Acids Res. 2019;47:D983–8. 10.1093/nar/gky1027. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 85. Zhang M, Zong W, Zou Det al. MethBank 4.0: an updated database of DNA methylation across a variety of species. Nucleic Acids Res. 2023;51:D208–16. 10.1093/nar/gkac969. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 86. Li R, Liang F, Li Met al. MethBank 3.0: a database of DNA methylomes across a variety of species. Nucleic Acids Res. 2018;46:D288–95. 10.1093/nar/gkx1139. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 87. Zou D, Sun S, Li Ret al. MethBank: a database integrating next-generation sequencing single-base-resolution DNA methylation programming data. Nucleic Acids Res. 2015;43:D54–8. 10.1093/nar/gku920. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 88. Jin E, Zhao D, Wu Get al. OBIA: an open biomedical imaging archive. Genomics Proteomics Bioinformatics. 2023;21:1059–65. 10.1016/j.gpb.2023.09.003. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 89. Cao Y, Tian D, Tang Zet al. OPIA: an open archive of plant images and related phenotypic traits. Nucleic Acids Res. 2024;52:D1530–7. 10.1093/nar/gkad975. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 90. Meng J, Han M, Huang Yet al. PreDigs: a database of context-specific cell-type markers and precise cell subtypes for digestive cell annotation. Genomics Proteomics Bioinformatics. 2025; 23:qzaf066. 10.1093/gpbjnl/qzaf066. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 91. Liu W, Wang P, Zhuang Xet al. RDBSB: a database for catalytic bioparts with experimental evidence. Nucleic Acids Res. 2025;53:D709–16. 10.1093/nar/gkae844. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 92. Li T, Zuo L, Wang Pet al. TonguExpert: a deep learning-based algorithm platform for fine-grained extraction and classification of tongue phenotypes. Phenomics. 2025;5:109–22. 10.1007/s43657-024-00210-9. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 93. Liu L, Cao S, Lin Wet al. miMatch: a microbial metabolic background matching tool for mitigating host confounding in metagenomics research. Gut Microbes. 2024;16:2434029. 10.1080/19490976.2024.2434029. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 94. Gao W, Lin W, Li Qet al. Identification and validation of microbial biomarkers from cross-cohort datasets using xMarkerFinder. Nat Protoc. 2024;19:2803–30. 10.1038/s41596-024-00999-9. [DOI] [PubMed] [Google Scholar]
- 95. Zhang T, Liu Y, Guo Xet al. VISTA: a tool for fast taxonomic assignment of viral genome sequences. Genomics Proteomics Bioinformatics. 2025;23:qzae082. 10.1093/gpbjnl/qzae082. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 96. Dwojak T, Pietruszka M, Borchmann Łet al. From Dataset Recycling to Multi-Property Extraction and Beyond. Proceedings of the 24th Conference on Computational Natural Language Learning. 2020, 641–51. 10.18653/v1/P17. [DOI] [Google Scholar]
- 97. Polak MP, Morgan D.. Extracting accurate materials data from research papers with conversational language models and prompt engineering. Nat Commun. 2024;15:1569. 10.1038/s41467-024-45914-8. [DOI] [PMC free article] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Data Availability Statement
All resources and services are publicly available on the home page of CNCB–NGDC (https://ngdc.cncb.ac.cn).




