Skip to main content
Nucleic Acids Research logoLink to Nucleic Acids Research
. 2025 Nov 20;54(D1):D1143–D1151. doi: 10.1093/nar/gkaf1248

InsectBase 3.0: a comprehensive multi-omics resource for insects

Zuoqi Wang 1, Hao Chen 2, Shuo Jin 3, Zonghuan Li 4, Yang Mei 5,6, Yiqi Xiao 7, Chenfan Zhao 8, Tianyu Zhou 9, Fei Li 10, Ying Liu 11,, Kang He 12,
PMCID: PMC12807663  PMID: 41263103

Abstract

Insects represent the most diverse animal group and play essential roles in ecosystems, agriculture, and human health. The rapid accumulation of high-quality genomes and diverse omics datasets provides unprecedented opportunities to advance research in insect biology and evolution, yet also poses challenges in data integration, accessibility, and reuse. To address these demands, we developed InsectBase 3.0 (www.insect-genome.com), an upgraded platform that integrates extensive insect multi-omics data and provides user-friendly tools for furthor analysis. InsectBase 3.0 covers 3020 species across 24 insect orders, including 1651 chromosome-level assemblies and 61 353 curated transcriptomes. In addition, the new version incorporates a large-scale protein structure dataset (474 300 predicted models and 381 experimental entries), genome resequencing and variant resources, curated transposable element libraries, and 3D morphological reconstructions. To organize these heterogeneous resources, InsectBase 3.0 introduces knowledge graphs centered on species and genes, enabling users to navigate cross-dataset relationships more intuitively. In addition, the web framework has been modernized to improve efficiency, interactivity, and visualization. Together, these updates establish InsectBase 3.0 as the most comprehensive open-access insect omics platform to date, providing a valuable resource for evolutionary, functional, and applied research in insects.

Graphical Abstract

Graphical Abstract.

Graphical Abstract

Introduction

Insects represent a cornerstone of global biodiversity and, through their diverse interactions with plants, animals, and humans, they shape ecosystem services [1, 2], influence agricultural productivity [3, 4], and contribute to biomedical discoveries [5], reflecting their central importance across ecological, evolutionary, agricultural, and applied sciences [6]. Recent years have witnessed rapid advances in sequencing technologies and the sharp decline in costs [7, 8], which have fueled an unprecedented accumulation of insect omics datasets. Beyond foundational genomic datasets, emerging datasets such as protein structural datasets, single-cell transcriptomes, higher-order chromatin organization, and 3D morphological reconstructions offer complementary perspectives on insect biology. However, the explosive growth and increasing diversity of these datasets have posed significant challenges, as much of the information remains fragmented, inconsistently annotated, or difficult to access, underscoring the urgent need for a comprehensive and well-curated insect omics database.

To cope with these challenges, a wide range of insect genomic databases have progressively expanded, from early community-driven resources (e.g. FlyBase, VectorBase, and i5k Workspace@NAL) to more recent platforms (e.g. SilkMeta, InsectTFDB, and GPIBase) [914]. However, most of these resources are limited to particular insect taxa or specific omics themes, and thus fall short of providing a unified and comprehensive platform for the insect research community. To bridge this gap, we first launched InsectBase in 2016 [15], which integrated nearly all available insect genome data with advanced analysis tools, and subsequently released a major update as InsectBase 2.0 in 2022 [16]. Compared to other insect-omics databases [1014, 20], InsectBase has from the outset aimed to cover a broad taxonomic spectrum and to integrate multiple omics data types within a single one-stop platform. This combination of wide taxonomic coverage and cross-omics integration allows InsectBase to complement existing resources and facilitate comparative and functional studies across diverse insect groups. Reflecting these advantages, InsectBase 2.0 has since gained broad recognition, attracting >365 000 visits from 77 countries and regions and being cited 137 times, highlighting growing influence in the field.

Recent advances in sequencing technologies and the expanding taxonomic coverage of insect research have greatly increased the scale and quality of insect multi-omics datasets. Coupled with growing feedback from the community on InsectBase 2.0, these developments highlight the need for broader data integration and improved functionality. Therefore, we now present InsectBase 3.0 with four major improvements: (i) The quantity and quality of insect genomic data have been substantially expanded. InsectBase 3.0 now includes genomes from 3020 species across 24 insect orders, with 1651 assemblies at the chromosome level, and offers 61 353 transcriptomes with manually curated metadata on tissue types and developmental stages; (ii) a large-scale dataset of insect protein structures, comprising 474 300 predicted models and 381 experimentally determined entries, accompanied by multiple analysis utilities; (iii) the integration of newly available omics resources, including 3D morphological dataset, genome resequencing and variant dataset, and transposable element (TE) profiles; and (iv) to accommodate the increasing data volume and community usage, the web framework has been modernized at both server and interface levels, ensuring greater efficiency in data processing and a smoother user experience.

Materials and methods

Genome

Genomic data were retrieved from a range of public repositories such as the National Center for Biotechnology Information (NCBI) [17], National Genomics Data Center (NGDC) [18], BIPAA (https://bipaa.genouest.org/is/), GigaDB, i5k Workspace@NAL [9], LepBase (http://lepbase.org/), VectorBase [11], DNA Data Bank of Japan (DDBJ) [19], and SilkDB 3.0 [20], supplemented by genomes generated from our own sequencing projects (Supplementary Table S1). All assemblies were subjected to quality assessment using BUSCO (v5.7.1) [21] and Compleasm (v0.2.6) [22] to ensure consistency and reliability with the insecta_odb10 lineage dataset. For genome annotation, we applied a standardized local pipeline to all assemblies. First, repeat elements were identified de novo using RepeatModeler2 (v2.0.7) [23], combined with previously constructed insect repeat library, and subsequently masked with RepeatMasker (v4.2.1) [24]. Next, three lines of evidence were generated for gene annotation. For ab initio prediction, BRAKER3 (v3.0.8) [25] was employed to produce de novo gene models. For RNAseq-based evidence, RNA-seq reads were aligned to the genome using HISAT2 (v2.2.1) [26], and assembled into transcripts with StringTie (v2.1.5) [27]. For homology-based evidence, protein sequences were aligned to the genome using miniprot (v0.12) [28]. Finally, we integrated three types of evidences by EVidenceModeler (v2.1) [29] to obtain the final gene annotation. OMArk (v0.3.1) [30] was further used for annotation quality evaluation with the LUCA reference database. All scripts and parameter settings used in the above genome workflow have been archived in Zenodo (DOI: [10.5281/zenodo.17399264]) and are also available from our GitHub repository (https://github.com/Wzuoqi/InsectBase).

Gene

Functional annotation of the predicted gene sets was performed by BLASTP alignment against the Swiss-Prot protein database [31, 32], implemented with DIAMOND (v2.1.13) [33]. Reference KEGG Orthology (KOs) for each gene were assigned using EggNOG-mapper (v2.1.13) [34]. Potential HGT genes were identified from BLAST searches against the NCBI NR/NT databases [17], considering genes as HGT candidates when at least 15 of the top 20 hits were from non-insect species.

Protein structure

Protein structure prediction for insects was performed using AlphaFold3 (v3.0.0) [35] with default model parameters and reference databases. In total, 474 300 protein structures from 27 key gene families across 231 insect species were predicted (Supplementary Table S2). Prediction confidence was evaluated using metrics such as pLDDT and pTM. For visualization of the predicted three-dimensional structures and their associated pLDDT scores, Mol* Viewer [36] and NGL Viewer [37] were employed. In addition, 381 experimentally determined insect protein structures were collected from the Protein Data Bank (PDB) [38] through keyword-based retrieval followed by manual curation. These structures encompass three major experimental techniques: electron microscopy, X-ray diffraction, and solution NMR. Large-scale multiple structural alignments were performed for each gene family using FoldMason (v9), and protein phylogenetic trees were constructed based on structural similarities. For online utilities, the predicted protein structures support an online all-versus-all structural comparison implemented with Foldseek (v10.941cd33) [39] on the server side. Structural similarity was assessed using TM-align as the alignment method, with normalized TM-scores (relative to query length) applied to evaluate the degree of similarity.

Genome resequencing and variant data

We retrieved insect genome resequencing data from the NCBI SRA database and used Python-based web spider scripts to obtain associated sample metadata, including strain and geographic information etc. After quality control and removal of incomplete records, 3223 genome resequencing samples from 99 species across 6 orders were retained. Variant identification was then performed for each resequencing sample using Bcftools mpileup (v1.6) [40], and functional annotation of SNPs and InDels was conducted with SnpEff (v5.2a) [41]. The variant data summary for each sample dataset is displayed on the website, and the corresponding VCF files are also available for download.

Transposable element

TEs were annotated across 426 representative insect genomes, selected to cover 18 orders and subjected to quality filtering prior to analysis. Annotation was performed using HiTE (v3.1.2) [42], which combines structural and homology-based strategies to identify diverse TE classes, including LTR, non-LTR, TIR, and Helitron elements, and produces non-redundant, high-quality TE libraries for each genome.

3D images

A total of 101 insect samples were collected, representing 7 orders. Specimens were obtained from laboratory field collections or intercepted by customs, immediately preserved in 75% ethanol, and stored at 4°C until processing. Prior to scanning, all samples underwent a staining procedure. Briefly, specimens were fixed in Bouin’s solution, stained with standard Lugol’s iodine, rinsed with 1× phosphate-buffered saline (PBS), and subsequently stored in 1× PBS. Prepared samples were scanned within one week of staining. Micro-CT scanning was performed using a SkyScan 1272 desktop high-resolution 3D X-ray microscope (Bruker). Raw projection data were reconstructed with NRecon software (v1.7.4, Bruker), and the reconstructed datasets were further processed in Dragonfly (v2022.2, Object Research Systems, Canada) to generate three-dimensional structural models, which were exported in STL format.

Implementation of database

The updated InsectBase 3.0 was implemented using a modular architecture. The frontend was developed with React (v18.3.1) and Next.js (v14.2.7) frameworks, complemented by TailwindCSS (v3.4.1) for modern and consistent styling. Interactive web visualization was supported by ECharts (v5.6.0) and Three.js (v0.161.0). The backend was built upon the Django framework (v4.2) connected to a PostgreSQL (v16.1) database for efficient data management. Deployment was achieved through Docker (v26.1.4) containerization with Nginx (v1.16.1) serving as the reverse proxy and load balancer.

Updates in insectbase 3.0

From data to insight: improved coverage, annotation, and curation

The quantity and quality of insect genomic resources in InsectBase 3.0 have been substantially expanded [7, 8]. The database now spans 3020 species representing 24 insect orders, with 1651 assemblies reaching the chromosome level, providing a much broader phylogenetic framework for comparative analyses (Fig. 1B). To ensure consistency, all genomes were uniformly annotated using our standardized in-house pipeline, and, in response to community feedback, we implemented a quality-ranking system based on OMArk and BUSCO assessments [21, 30]. This system allows users to intuitively evaluate both the absolute and relative quality of genome annotations through rankings and visualizations, thereby facilitating informed decisions on their suitability for downstream research.

Figure 1.

Figure 1.

Data expansion and curation of InsectBase 3.0. (A) Data summary of main modules of InsectBase 3.0. (B) Taxonomic distribution of insect genomes in InsectBase 3.0. (C) The transcriptome samples with curated metadata in InsectBase 3.0.

As for transcriptomes, benefiting from advances in third-generation sequencing technologies [43, 44], transcriptomes in InsectBase 3.0 has surged to 61 353. However, the rapid expansion has led to inconsistencies in sample metadata, limiting the efficient use of these transcriptome resources. To address this issue, we manually curated the transcriptomes based on their original descriptions and associated metadata, and established a unified classification system according to sex, developmental stage, and tissue type, thereby facilitating their effective use in downstream research and comparative analyses (Fig. 1C). In addition, other data modules such as miRNA, lncRNA, and chromosome have also been updated. Overall, InsectBase 3.0 not only achieves a substantial expansion in data quantity but also places greater emphasis on systematic curation and quality assessment, thereby providing a more efficient platform for transforming genomic data into biological insight (Table 1).

Table 1.

Data summary of InsectBase 3.0

Feature Units v1.0 v2.0 v3.0 Fold increase
Genomes Species 138 815 3,093 3.79
Transcriptomes Runs 116 25 805 61 421 2.38
Coding Sequences Transcripts 160 905 15 045 111 40 987 806 2.72
Chromosomes Species 207 1570 7.58
miRNAs 7544 71 524 202 880 2.83
lncRNAs 2439 1 011 568 1 403 399 1.38
Insect Viruses 1524 1742 1.14
Predicted Protein Structure 474 300 New
Experimental Protein Structure 381 New
3D Images Species 101 New
TEs Species 919 New
Genome Resequencing Samples 3223 New

From sequence to structure: insect protein structure dataset with utilities

InsectBase 3.0 introduces a major update in protein structural resources, providing both predicted and experimentally insect protein structures. Using AlphaFold3, we generated 367 538 structures across 27 key gene families from 231 insect species, complemented by 381 curated experimental structures from the PDB. This large-scale expansion substantially enhances the functional dimension of InsectBase, enabling researchers to explore not only sequences but also structural and evolutionary relationships. By integrating FoldMason-based multiple structural alignments and protein phylogenies, InsectBase 3.0 offers a novel framework for structural comparative genomics. Furthermore, the implementation of Foldseek enables fast, server-side all-versus-all structural comparisons, allowing users to efficiently assess protein similarity beyond sequence conservation. These structural resources provide a new dimension to insect omics, enabling functional annotation and evolutionary inference that go beyond sequence-based analysis.

From genotype to phenotype: 3D morphological reconstructions of insects

The extraordinary diversity of insect species presents major challenges for extracting biological insights from existing genomic data alone. Phenotypic datasets, such as external and internal morphological traits, can guide more effective mining and comparison of omics resources, thereby generating new perspectives on insect biology [45, 46]. In InsectBase 3.0, we introduce a newly established micro-CT morphological dataset comprising 101 insect samples representing multiple orders. Each species was processed through standardized fixation, staining, and high-resolution scanning, and the resulting three-dimensional reconstructions in STL format for direct reuse. This dataset provides researchers with access to detailed external and internal insect morphological features, offering a phenotypic dimension that complements genomic datasets and strengthens integrative analyses.

Introduction of other omics datasets

Genome resequencing data represent a key resource in insect research, enabling investigations of population diversity, insecticide resistance, strain breading, and the evolutionary dynamics of invasive insects [4750]. Therefore, InsectBase 3.0 incorporates 3223 resequencing datasets from 99 species across 6 insect orders, each processed to provide high-confidence SNP and InDel annotations. By centralizing these datasets, InsectBase 3.0 establishes a unified resource that makes population-level variation data more accessible and reusable for the insect research community.

As for TEs, they constitute a substantial fraction of insect genomes and are increasingly recognized as critical players in insect research. They were proved to shape chromosome organization and drive long-term evolutionary diversification. Also, TEs participate in epigenetic regulation, host adaptation, reproductive processes, and telomere maintenance, underscoring their growing importance for understanding insect biology [5155]. In InsectBase 3.0, TEs were systematically annotated in 426 high-quality genomes spanning 18 insect orders, generating comprehensive libraries of LTR, non-LTR, TIR, and Helitron elements. In addition, curated TE libraries at Insecta and order-specific levels are provided for download to support repeat annotation and other downstream analysis.

Enhanced user interface and integrated knowledge graph

InsectBase 3.0 features an updated, streamlined web architecture that reorganizes and extends the platform into 17 data modules, namely “organism,” “genome,” “chromosome,” “genome reseq,” “protein-coding gene,” “gene family,” “KEGG pathway,” “protein structure,” “microRNA,” “lncRNA,” “UTR,” “transposable element,” “potential HGT,” “3D image,” and “classification” (Fig. 2). In addition, the database offers a suite of online functional modules, such as BLAST search, genome synteny visualization, genome browser, structure alignment, and a dedicated download center, which facilitate efficient data exploration and retrieval. The framework was redesigned to accommodate rapidly growing and increasingly heterogeneous datasets, to enable cross-module queries at interactive speed, and to provide richer, more consistent visualizations and metadata.

Figure 2.

Figure 2.

Enhanced user interface features and data query workflow of InsectBase 3.0.

Beyond these improvements, InsectBase 3.0 introduces knowledge graphs as a new approach to connect heterogeneous datasets and reveal relationships among entities, supporting integrative exploration of insect research. Two complementary knowledge graphs were constructed, centered on insect species and genes, and can be accessed through the “organism” and “gene” modules. These graphs were built by integrating curated metadata, cross-module associations in the database, and external references, and are intended to help researchers navigate biological networks more effectively. For example, when accessing a specific species through the “organism” knowledge graph, users can obtain links to all information available in the database for that species, including a species overview, classification tags, available genomes, transcriptomes, coding genes, noncoding RNAs, resequencing data, and 3D images. The graph also connects to external resources such as related virus, symbiont, publications, and recommendations of taxonomically or functionally similar species, extending the scope of exploration beyond the database. In addition, both knowledge graphs support export in JSON format, allowing users to readily incorporate the data into downstream analyses or language model training.

Case study: linking sine oculis homeobox (six) gene family to flight capability in nilaparvata lugens

The brown planthopper (BPH), Nilaparvata lugens, is a major migratory pest of rice, whose devastating impact is directly linked to its long-distance flight capability [5658]. This flight capacity is crucially dependent on the development and performance of the thoracic flight muscles [59, 60]. The sine oculis homeobox (SIX) gene family, a master regulator of myogenesis across animal species, is therefore a prime candidate for investigating the genetic basis of this critical trait [6163]. The following case study will demonstrate how InsectBase can be utilized for a preliminary investigation into the relationship between a target gene and a complex phenotypic trait like flight capability.

First, we identified the SIX gene family in N. lugens by searching the InsectBase Gene module with the keywords “Homeobox protein SIX.” This query returned three candidate genes (SIX1, SIX4, and SIX6), together their sequences and EggNog annotation (Supplementary Table S3). To complement and validate this initial identification, we also performed an online BLAST search against the N. lugens genome using verified SIX protein sequences from the Swiss-Prot database as queries. ​ Subsequently, to characterize the evolutionary landscape of the SIX gene family across insects, we used the Organism module to find closely related hemipteran species and several model insects with available genomes, from which potential SIX homologs were identified (Fig. 3A). The Genome Browser and Genome Synteny modules were used to examine the genomic context and syntenic conservation of each SIX gene across related species (Fig. 3B and C). Furthermore, to extend the comparison from sequence to structure, we retrieved the predicted protein structures for each SIX gene from the Protein Structure module (Fig. 3D). Using the integrated online FoldSeek tool, we performed comparisons of these structures to assess explore their structural conservation and potential functional similarities. To investigate expression patterns, we utilized the Transcriptome module, which provides 460 N. lugens transcriptomes (89 with curated metadata) with their aligned GTF files for downstream analysis. ​​ Expression values extracted for SIX1 (Nlug008426.1) from GTF files revealed marked upregulation in thoracic tissues, including wing buds and nota, suggesting its involvement in wing morphogenesis and flight muscle development (Fig. 3E). Complementing the molecular data, 3D micro-CT images provide quantitative baseline morphological data of flight muscles for N. lugens (Fig. 3F). These image resources enable comparative studies of flight muscle architecture across related planthopper species or furnish an essential phenotypic reference for assessing the outcomes of future CRISPR or RNA inteference functional experiments.

Figure 3.

Figure 3.

Case study: linking sine oculis homeobox (SIX) gene family to flight capability in Nilaparvata lugens. (A) Distribution of the SIX gene family across insect species. (B) Visualization of the gene SIX1 (Nlug008426.1) location in the N. lugens genome with the Genome Browser module. (C) Chromosomal synteny between N. lugens and Laodelphax striatellus in the Genome Synteny module. (D) Predicted protein structure of the N. lugens gene SIX1 (Nlug008426.1) previewed in the Protein Structure module. (E) Differential expression levels of SIX1 across N. lugens tissues. (F) 3D micro-CT image of N. lugens.

Discussion and future development

Advances in sequencing technologies have led to a rapid accumulation of high-quality insect genomes and multi-omics datasets, with greatly expanded taxonomic coverage [7, 8]. As a result, the main challenge for insect omics databases is shifting from simply providing access to data toward enabling their efficient integration and use to generate biological insights. Therefore, in addition to the dramatic expansion of data volume, InsectBase 3.0 prioritizes features and tools that facilitate the effective use of omics datasets. For genomes, InsectBase 3.0 applies a standardized annotation pipeline and implements quality-ranking metrics, enabling users to assess data reliability and select suitable assemblies for downstream studies. For transcriptomes, the database provides 61 353 entries with manually curated metadata on sex, developmental stage, and tissue type, ensuring consistent organization and facilitating comparative analysis. Furthermore, InsectBase 3.0 incorporates 3D phenotypic data, which serve as a complementary layer to inform and enhance the mining of genomic datasets. Beyond data expansion and curation, the database now features knowledge graphs organized around species and genes, enabling cross-linking among internal datasets and external references to generate new biological insights.

To ensure the long-term sustainability of InsectBase, we are implementing a clear update and maintenance strategy. Data will be synchronized with external databases such as NCBI and PDB on an annual schedule each September, with version control systems to track updates and maintain data provenance for reproducibility. Long-term maintenance will be supported by institutional commitments and funding, and contingency plans will be in place to guarantee service continuity during technical issues or funding gaps. These measures will also be documented in the Help section of the website for transparency.

To better support the insect research community and enable deeper biological insights, InsectBase will continue to update in the following directions: (i) Integration of emerging omics data. New modules will be added for novel data types relevant to insect research, as exemplified in this release by protein structures and resequencing datasets. Rapidly developing fields such as insect single-cell omics and telomere-to-telomere genomes are expected to be incorporated in future updates. (ii) Enabling phenotype-informed analyses. Future updates will enrich InsectBase with curated phenotypic datasets—for example, the micro-CT 3D images in this update—allowing users to connect morphological variation with underlying genetic and evolutionary mechanisms. (iii) Implementation of a public API: To facilitate programmatic access and automated data integration, InsectBase will provide an open API for querying and retrieving data, thereby improving interoperability with external bioinformatics tools and workflows. (iv) Knowledge graph-enabled language models: Future updates will explore integrating species- and gene-centered knowledge graphs with large language models, supported by agent-based approaches and retrieval-augmented generation (RAG). This direction is expected to enable intelligent query, automated knowledge discovery, and more context-aware interpretation of insect omics data.

Supplementary Material

gkaf1248_Supplemental_File

Acknowledgements

Author contributions: Zuoqi Wang (Conceptualization [equal], Data curation [equal], Formal analysis [equal], Supervision [equal], Validation [equal], Visualization [equal], Writing—original draft [equal]), Hao Chen (Data curation [equal], Formal analysis [equal]), Shuo Jin (Data curation [equal], Formal analysis [equal]), Zonghuan Li (Data curation [equal], Methodology [equal]), Yang Mei (Data curation [equal], Formal analysis [equal]), Chenfan Zhao (Data curation [equal]), Tianyu Zhou (Data curation [equal]), Liu Ying (Conceptualization [equal], Funding acquisition [equal], Writing—review & editing [equal]), and Kang He (Conceptualization [equal], Funding acquisition [equal], Writing—review & editing [equal]).

Contributor Information

Zuoqi Wang, State Key Laboratory of Rice Biology and Breeding and Ministry of Agricultural and Rural Affairs Key Laboratory of Molecular Biology of Crop Pathogens and Insect Pests, Institute of Insect Sciences, Zhejiang University, Hangzhou 310058, China.

Hao Chen, State Key Laboratory of Rice Biology and Breeding and Ministry of Agricultural and Rural Affairs Key Laboratory of Molecular Biology of Crop Pathogens and Insect Pests, Institute of Insect Sciences, Zhejiang University, Hangzhou 310058, China.

Shuo Jin, State Key Laboratory of Rice Biology and Breeding and Ministry of Agricultural and Rural Affairs Key Laboratory of Molecular Biology of Crop Pathogens and Insect Pests, Institute of Insect Sciences, Zhejiang University, Hangzhou 310058, China.

Zonghuan Li, State Key Laboratory of Rice Biology and Breeding and Ministry of Agricultural and Rural Affairs Key Laboratory of Molecular Biology of Crop Pathogens and Insect Pests, Institute of Insect Sciences, Zhejiang University, Hangzhou 310058, China.

Yang Mei, State Key Laboratory of Rice Biology and Breeding and Ministry of Agricultural and Rural Affairs Key Laboratory of Molecular Biology of Crop Pathogens and Insect Pests, Institute of Insect Sciences, Zhejiang University, Hangzhou 310058, China; College of Plant Protection, Jilin Agricultural University, Changchun 130118, China.

Yiqi Xiao, State Key Laboratory of Rice Biology and Breeding and Ministry of Agricultural and Rural Affairs Key Laboratory of Molecular Biology of Crop Pathogens and Insect Pests, Institute of Insect Sciences, Zhejiang University, Hangzhou 310058, China.

Chenfan Zhao, State Key Laboratory of Rice Biology and Breeding and Ministry of Agricultural and Rural Affairs Key Laboratory of Molecular Biology of Crop Pathogens and Insect Pests, Institute of Insect Sciences, Zhejiang University, Hangzhou 310058, China.

Tianyu Zhou, State Key Laboratory of Rice Biology and Breeding and Ministry of Agricultural and Rural Affairs Key Laboratory of Molecular Biology of Crop Pathogens and Insect Pests, Institute of Insect Sciences, Zhejiang University, Hangzhou 310058, China.

Fei Li, State Key Laboratory of Rice Biology and Breeding and Ministry of Agricultural and Rural Affairs Key Laboratory of Molecular Biology of Crop Pathogens and Insect Pests, Institute of Insect Sciences, Zhejiang University, Hangzhou 310058, China.

Ying Liu, Key Laboratory of Green Prevention and Control of Agricultural Transboundary Pests of Yunnan Province/Agricultural Environment and Resource Research Institute, Yunnan Academy of Agricultural Sciences, Kunming 650205, China.

Kang He, State Key Laboratory of Rice Biology and Breeding and Ministry of Agricultural and Rural Affairs Key Laboratory of Molecular Biology of Crop Pathogens and Insect Pests, Institute of Insect Sciences, Zhejiang University, Hangzhou 310058, China.

Supplementary data

Supplementary data is available at NAR online.

Conflict of interest

The authors declare no competing interest.

Funding

This work was supported by 2024 Yangtze River Delta Science and Technology Innovation Community Joint Research (Basic Research) Project [2024CSJZN0900], National Key Research and Development Program of China [2022YFD1401600, 2021YFD1400100], National Natural Science Foundation of China [32202366] and Natural Science Foundation of Zhejiang Province [LZ23C140002]. Funding to pay the Open Access publication charges for this article was provided by Yangtze River Delta Science and Technology Innovation Community Joint Research (Basic Research) Project.

Data availability

All data in InsectBase 3.0 are freely available at http://www.insect-genome.com/. The database workflows have been archived in Zenodo (DOI: [10.5281/zenodo.17399264]).

References

  • 1. Medina-Serrano N, Hossaert-McKey M, Diallo Aet al. Insect-flower interactions, ecosystem functions, and restoration ecology in the northern Sahel: current knowledge and perspectives. Biol Rev Camb Philos Soc. 2025;100:969–95. 10.1111/brv.13170. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 2. Losey JE, Vaughan M. The economic value of ecological services provided by insects. Bioscience. 2006;56:311–23. 10.1641/0006-3568(2006)56[311:TEVOES]2.0.CO;2. [DOI] [Google Scholar]
  • 3. Meier R, Lim GS. Conflict, convergent evolution, and the relative importance of immature and adult characters in endopterygote phylogenetics. Annu Rev Entomol. 2009;54:85–104. 10.1146/annurev.ento.54.110807.090459. [DOI] [PubMed] [Google Scholar]
  • 4. Poelchau MF, Coates BS, Childers CPet al. Agricultural applications of insect ecological genomics. Curr Opin Insect Sci. 2016;13:61–9. 10.1016/j.cois.2015.12.002. [DOI] [PubMed] [Google Scholar]
  • 5. Ratcliffe NA, Mello CB, Garcia ESet al. Insect natural products and processes: new treatments for human disease. Insect Biochem Molec. 2011;41:747–69. 10.1016/j.ibmb.2011.05.007. [DOI] [PubMed] [Google Scholar]
  • 6. van Huis A. Potential of insects as food and feed in assuring food security. Annu Rev Entomol. 2013;58:563–83. 10.1146/annurev-ento-120811-153704. [DOI] [PubMed] [Google Scholar]
  • 7. Li F, Wang X, Zhou X. The genomics revolution drives a new era in entomology. Annu Rev Entomol. 2025;70:379–400. 10.1146/annurev-ento-013024-013420. [DOI] [PubMed] [Google Scholar]
  • 8. Poelchau MF, Coates BS, Childers CPet al. Agricultural applications of insect ecological genomics. Curr Opin Insect Sci. 2016;13:61–9. 10.1016/j.cois.2015.12.002. [DOI] [PubMed] [Google Scholar]
  • 9. Poelchau M, Childers C, Moore Get al. The i5k Workspace@NAL–enabling genomic data access, visualization and curation of arthropod genomes. Nucleic Acids Res. 2015;43:D714–9. 10.1093/nar/gku983. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 10. Thurmond J, Goodman JL, Strelets VBet al. FlyBase 2.0: the next generation. Nucleic Acids Res. 2019;47:D759–65. 10.1093/nar/gky1003. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 11. Giraldo-Calderón GI, Harb OS, Kelly SAet al. VectorBase.org updates: bioinformatic resources for invertebrate vectors of human pathogens and related organisms. Curr Opin Insect Sci. 2022;50:100860. 10.1016/j.cois.2021.11.008. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12. Lu K, Pan Y, Shen Jet al. SilkMeta: a comprehensive platform for sharing and exploiting pan-genomic and multi-omic silkworm data. Nucleic Acids Res. 2024;52:D1024–32. 10.1093/nar/gkad956. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 13. Zhou H, Tang J, Cheng Zet al. InsectTFDB: a comprehensive database and analysis platform for insect transcription factors. Mol Ecol Resour. 2025;25:e70006. 10.1111/1755-0998.70006. [DOI] [PubMed] [Google Scholar]
  • 14. Wang Y, Mei Y, Su Cet al. GPIBase: a comprehensive resource for geminivirus-plant-insect research. Mol Plant. 2023;16:647–9. 10.1016/j.molp.2023.02.007. [DOI] [PubMed] [Google Scholar]
  • 15. Yin C, Shen G, Guo Det al. InsectBase: a resource for insect genomes and transcriptomes. Nucleic Acids Res. 2016;44:D801–7. 10.1093/nar/gkv1204. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16. Mei Y, Jing D, Tang Set al. InsectBase 2.0: a comprehensive gene resource for insects. Nucleic Acids Res. 2022;50:D1040–5. 10.1093/nar/gkab1090. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 17. Sayers EW, Bolton EE, Brister JRet al. Database resources of the national center for biotechnology information. Nucleic Acids Res. 2022;50:D20–6. 10.1093/nar/gkab1112. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 18. CNCB-NGDC Members and Partners . Database Resources of the National Genomics Data Center, China National Center for Bioinformation in 2024. Nucleic Acids Res. 2024;52:D18–32. 10.1093/nar/gkad1078. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19. Okido T, Kodama Y, Mashima Jet al. DNA Data Bank of Japan (DDBJ) update report 2021. Nucleic Acids Res. 2022;50:D102–5. 10.1093/nar/gkab995. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 20. Lu F, Wei Z, Luo Yet al. SilkDB 3.0: visualizing and exploring multiple levels of data for silkworm. Nucleic Acids Res. 2020;48:D749–55. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 21. Seppey M, Manni M, Zdobnov EM. BUSCO: assessing genome assembly and annotation completeness. Methods Mol Biol. 2019;1962:227–45. [DOI] [PubMed] [Google Scholar]
  • 22. Huang N, Li H. compleasm: a faster and more accurate reimplementation of BUSCO. Bioinformatics. 2023;39:btad595. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 23. Flynn JM, Hubley R, Goubert Cet al. RepeatModeler2 for automated genomic discovery of transposable element families. P Natl Acad Sci Usa. 2020;117:9451–7. 10.1073/pnas.1921046117. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 24. Tarailo-Graovac M, Chen N. Using RepeatMasker to identify repetitive elements in genomic sequences. Curr Protoc Bioinform. 2009;Chapter 4:4–10. [DOI] [PubMed] [Google Scholar]
  • 25. Gabriel L, Brůna T, Hoff KJet al. BRAKER3: fully automated genome annotation using RNA-seq and protein evidence with GeneMark-ETP, AUGUSTUS, and TSEBRA. Genome Res. 2024;34:769–77. 10.1101/gr.278090.123. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 26. Kim D, Paggi JM, Park Cet al. Graph-based genome alignment and genotyping with HISAT2 and HISAT-genotype. Nat Biotechnol. 2019;37:907–15. 10.1038/s41587-019-0201-4. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 27. Pertea M, Pertea GM, Antonescu CMet al. StringTie enables improved reconstruction of a transcriptome from RNA-seq reads. Nat Biotechnol. 2015;33:290–5. 10.1038/nbt.3122. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 28. Li H. Protein-to-genome alignment with miniprot. Bioinformatics. 2023;39:btad014. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 29. Haas BJ, Salzberg SL, Zhu Wet al. Automated eukaryotic gene structure annotation using EVidenceModeler and the Program to Assemble Spliced Alignments. Genome Biol. 2008;9:R7. 10.1186/gb-2008-9-1-r7. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 30. Nevers Y, Warwick Vesztrocy A, Rossier Vet al. Quality assessment of gene repertoire annotations with OMArk. Nat Biotechnol. 2025;43:124–33. 10.1038/s41587-024-02147-w. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 31. Boutet E, Lieberherr D, Tognolli Met al. UniProtKB/Swiss-Prot, the manually annotated section of the UniProt knowledgebase: how to use the entry view. Methods Mol Biol. 2016;1374:23–54. [DOI] [PubMed] [Google Scholar]
  • 32. Zaru R, Orchard S. UniProt Tools: BLAST, Align, Peptide Search, and ID Mapping. Curr Protoc. 2023;3:e697. 10.1002/cpz1.697. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 33. Buchfink B, Xie C, Huson DH. Fast and sensitive protein alignment using DIAMOND. Nat Methods. 2015;12:59–60. 10.1038/nmeth.3176. [DOI] [PubMed] [Google Scholar]
  • 34. Cantalapiedra CP, Hernández-Plaza A, Letunic Iet al. eggNOG-mapper v2: functional annotation, orthology assignments, and domain prediction at the metagenomic scale. Mol Biol Evol. 2021;38:5825–9. 10.1093/molbev/msab293. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 35. Abramson J, Adler J, Dunger Jet al. Accurate structure prediction of biomolecular interactions with AlphaFold 3. Nature. 2024;630:493–500. 10.1038/s41586-024-07487-w. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 36. Sehnal D, Bittrich S, Deshpande Met al. Mol* Viewer: modern web app for 3D visualization and analysis of large biomolecular structures. Nucleic Acids Res. 2021;49:W431–7. 10.1093/nar/gkab314. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 37. Rose AS, Bradley AR, Valasatava Yet al. NGL viewer: web-based molecular graphics for large complexes. Bioinformatics. 2018;34:3755–8. 10.1093/bioinformatics/bty419. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 38. Burley SK, Berman HM, Duarte JMet al. Protein Data Bank: a comprehensive review of 3D structure holdings and worldwide utilization by researchers, educators, and students. Biomolecules. 2022;12:1425. 10.3390/biom12101425. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 39. van Kempen M, Kim SS, Tumescheit Cet al. Fast and accurate protein structure search with Foldseek. Nat Biotechnol. 2024;42:243–6. 10.1038/s41587-023-01773-0. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 40. Danecek P, McCarthy SA. BCFtools/csq: haplotype-aware variant consequences. Bioinformatics. 2017;33:2037–9. 10.1093/bioinformatics/btx100. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 41. Cingolani P, Platts A, Wang LLet al. A program for annotating and predicting the effects of single nucleotide polymorphisms, SnpEff: sNPs in the genome of Drosophila melanogaster strain w1118; iso-2; iso-3. Fly. 2012;6:80–92. 10.4161/fly.19695. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 42. Hu K, Ni P, Xu Met al. HiTE: a fast and accurate dynamic boundary adjustment approach for full-length transposable element detection and annotation. Nat Commun. 2024;15:5573. 10.1038/s41467-024-49912-8. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 43. Wang Y, Zhao Y, Bollas Aet al. Nanopore sequencing technology, bioinformatics and applications. Nat Biotechnol. 2021;39:1348–65. 10.1038/s41587-021-01108-x. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 44. Espinosa E, Bautista R, Larrosa Ret al. Advancements in long-read genome sequencing technologies and algorithms. Genomics. 2024;116:110842. 10.1016/j.ygeno.2024.110842. [DOI] [PubMed] [Google Scholar]
  • 45. Schmidt J, Scholz S, Wiesner Jet al. MicroCT data provide evidence correcting the previous misidentification of an Eocene amber beetle (Coleoptera, Cicindelidae) as an extant species. Sci Rep. 2023;13:14743. 10.1038/s41598-023-39158-7. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 46. Ren C, Wen Y, Zheng Set al. Two transcriptional cascades orchestrate cockroach leg regeneration. Cell Rep. 2024;43:113889. 10.1016/j.celrep.2024.113889. [DOI] [PubMed] [Google Scholar]
  • 47. Peng Y, Mao K, Li Het al. Extreme genetic signatures of local adaptation in a notorious rice pest, Chilo suppressalis. Natl Sci Rev. 2025;12:nwae221. 10.1093/nsr/nwae221. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 48. Du Z, Wang X, Duan Yet al. Global invasion history and genomic signatures of adaptation of the highly invasive sycamore lace bug. Genomics Proteomics Bioinf. 2025;22:qzae074. 10.1093/gpbjnl/qzae074. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 49. Peng Y, Jin M, Li Zet al. Population genomics provide insights into the evolution and adaptation of the Asia Corn Borer. Mol Biol Evol. 2023;40:msad112. 10.1093/molbev/msad112. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 50. Ma L, Cao L, Chen Jet al. Rapid and repeated climate adaptation involving chromosome inversions following invasion of an Insect. Mol Biol Evol. 2024;41:msae044. 10.1093/molbev/msae044. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 51. Cong Y, Ye X, Mei Yet al. Transposons and non-coding regions drive the intrafamily differences of genome size in insects. iScience. 2022;25:104873. 10.1016/j.isci.2022.104873. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 52. Liu Q, Jiang F, Li Ret al. Chromatin dynamics of a large-sized genome provides insights into polyphenism and X0 dosage compensation of locusts. Nat Genet. 2025;57:2264–75. 10.1038/s41588-025-02330-y. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 53. Zhou B, Hu P, Liu Get al. Evolutionary patterns and functional effects of 3D chromatin structures in butterflies with extensive genome rearrangements. Nat Commun. 2024;15:6303. 10.1038/s41467-024-50529-0. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 54. Coronado-Zamora M, González J. The epigenetics effects of transposable elements are genomic context dependent and not restricted to gene silencing in Drosophila. Genome Biol. 2025;26:251. 10.1186/s13059-025-03705-4. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 55. Jiang Y, Hu J, Li Yet al. Comprehensive genomic analysis reveals novel transposable element-derived microRNA regulating caste differentiation in honeybees. Mol Biol Evol. 2025;42:msaf 074. 10.1093/molbev/msaf074. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 56. Wu J, Ge L, Liu Fet al. Pesticide-induced planthopper population resurgence in rice cropping systems. Annu Rev Entomol. 2020;65:409–29. 10.1146/annurev-ento-011019-025215. [DOI] [PubMed] [Google Scholar]
  • 57. Shi S, Wang H, Zha Wet al. Recent advances in the genetic and biochemical mechanisms of rice resistance to brown planthoppers (Nilaparvata lugens Stål). Int J Mol Sci. 2023;24:16959. 10.3390/ijms242316959. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 58. Zhang C, Brisson JA, Xu H. Molecular mechanisms of wing polymorphism in insects. Annu Rev Entomol. 2019;64:297–314. 10.1146/annurev-ento-011118-112448. [DOI] [PubMed] [Google Scholar]
  • 59. Iwamoto H. Structure, function and evolution of insect flight muscle. Biophysics. 2011;7:21–8. 10.2142/biophysics.7.21. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 60. Hou L, Guo S, Ding Det al. Neuroendocrinal and molecular basis of flight performance in locusts. Cell Mol Life Sci. 2022;79:325. 10.1007/s00018-022-04344-9. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 61. Kumar JP. The sine oculis homeobox (SIX) family of transcription factors as regulators of development and disease. Cell Mol Life Sci. 2008;66:565. 10.1007/s00018-008-8335-4. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 62. Yajima H, Motohashi N, Ono Yet al. Six family genes control the proliferation and differentiation of muscle satellite cells. Exp Cell Res. 2010;316:2932–44. 10.1016/j.yexcr.2010.08.001. [DOI] [PubMed] [Google Scholar]
  • 63. Wu W, Huang R, Wu Qet al. The role of Six1 in the genesis of muscle cell and skeletal muscle development. Int J Biol Sci. 2014;10:983–9. 10.7150/ijbs.9442. [DOI] [PMC free article] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

gkaf1248_Supplemental_File

Data Availability Statement

All data in InsectBase 3.0 are freely available at http://www.insect-genome.com/. The database workflows have been archived in Zenodo (DOI: [10.5281/zenodo.17399264]).


Articles from Nucleic Acids Research are provided here courtesy of Oxford University Press

RESOURCES