Abstract
In recent years, many high-quality reference genome sequences for arthropod species have been generated. Although most genome papers describe their protocols and metrics, no consensus exists on the data that should be included in genome reports. Here, we review current standards across seven key stages of an arthropod genome project (budgeting, sourcing and vouchering, sample preparation and sequencing, genome assembly, analysis reproducibility, databasing, and genome annotation) and identify persistent gaps in standards as well as their implementation. To assess current standards reporting in the community, we surveyed 100 arthropod genome papers published in 2024. The use of long reads to assemble highly contiguous arthropod genomes is now standard practice when adequate input DNA is available, and basic assembly contiguity and conserved gene content statistics are consistently reported. However, there is less standardization in pre- and post-assembly procedures and metrics. When comparing Darwin Tree of Life (DToL) genome notes to other journals, publications from the latter group were less likely to describe compliance with ethical collection practices, sample vouchering, post-assembly curation steps, and assembly quality metrics beyond basic contiguity and completeness values. Genome annotation practices are highly variable: some genome note formats do not explicitly require annotation, and while the reporting rate of protein-coding gene annotations is higher in non-DToL publications, the submission rate of annotations to centralized sequence databases is much lower. Our findings highlight critical opportunities to harmonize reporting standards and promote their dissemination, ensuring that future arthropod genomes are both comparable and maximally reusable for large-scale comparative and applied research.
Keywords: i5k initiative, Earth BioGenome Project, genome assembly, genome metrics
Introduction
The genome sequence of an organism underpins modern biological research and discovery. Improvements in sequencing technologies and assembly algorithms over the past two decades have enabled the assembly of structurally and biologically complete genomes. The recent completion of a telomere-to-telomere (T2T) human genome (Nurk et al. 2022; Rhie et al. 2023) represents a gold standard for producing similarly high-quality assemblies for diverse organisms.
Genomic characterization of arthropod biodiversity is rapidly advancing via projects affiliated to the Earth BioGenome Project (EBP) (Lewin et al. 2018, 2022; Blaxter et al. 2025). New arthropod genomes are being produced by regional EBP affiliates such as the Darwin Tree of Life (DToL) project (DToL Project Consortium 2022) and taxonomy-based affiliates such as Ag100Pest (Childers et al. 2021), Project Psyche (Wright et al. 2025), and Beenome100 (https://www.beenome100.org/). These assemblies are often viewed as “reference genomes,” which we define here as genome assemblies that (i) are as complete and accurate as possible for the given species at the time they were generated, and (ii) are accepted by the scientific community as the “reference” for downstream analyses such as producing an authoritative annotation for the genome, understanding evolutionary relationships, or assessing genetic variation. The increased availability of high-quality genomes in public databases will continue to accelerate biological discoveries (Fig. 1).
Fig. 1.

Improvements in arthropod assembly production rate and assembly quality submitted to INSDC from 2015 to 2025. a) Total counts of submitted arthropod genome assemblies, excluding alternate haplotypes of diploid assemblies. Counts are subdivided into chromosome-level vs non-chromosome-level submissions. *Total count from 2021 was adjusted after removing a high number of submissions (n = 806) from a single submitter. b) Distribution of assembly contig N50 values. Y-axis uses log10 scaling. Data points further than 1.5 * IQR are plotted as individual outliers. *Median contig N50 value from 2021 was adjusted after removing a high number of submissions (n = 806) from a single submitter. Total arthropod assemblies and contig N50 stats were retrieved using NCBI Datasets (O’Leary et al. 2024).
Large-scale genomic data production enables comparisons across species to yield novel biological insights, but these are valid only if the underlying data are comparable and accessible. In practice, genome producers’ research interests shape the data generated, the analytical tools used, and the results reported, leading to disparities in data quality, reporting standards, and accessibility that can hinder scientific progress via data re-use. The EBP has published quality standards and recommendations for sample collection, genome assembly, genome annotation, genome analysis, and databasing (Lawniczak et al. 2022, 2025) and maintains updated guidance at https://www.earthbiogenome.org. It remains unclear how widely these recommendations will be adopted in arthropod genomics, particularly outside major consortia, and whether they address arthropod-specific challenges.
The 5,000 arthropod genomes initiative (i5k) is an international, volunteer-driven EBP affiliate that aims to ensure that scientists can access high-quality genomes from insects and other arthropods, as well as the information needed to use them effectively (i5K Consortium 2013). We formed the i5k standards working group, prompted by community, requests to identify (i) current standards (Table 1) and (ii) persistent gaps (Table 2) at seven key stages of an arthropod genome project (budgeting, specimen sourcing and vouchering, sample preparation and sequencing, genome assembly, analysis reproducibility, databasing, and genome annotation). Standards can be quantitative (eg contig N50, BUSCO score) or procedural, akin to best practices (eg creating a specimen voucher, submission of data to public databases); thus, gaps in standardization can be the lack of appropriate quantitative measurements or the lack of completing established procedures as part of community genome assembly protocols. We do not provide comprehensive recommendations on best practices here, as they are already fully documented where they are originally described (see “Current standards” subsections below). Rather, by surveying recent publications and consulting with experts in the arthropod genomics community, we highlight areas where harmonized standards and improved reporting will enhance the usability of arthropod genomes for large-scale comparative and applied research.
Table 1.
Current known standards and associated resources in arthropod genomics.
| Genome project key stage | Known standards | Published standards | Known resources |
|---|---|---|---|
| Budgeting | Quantitative:
|
Lou et al. (2021) |
https://www.earthbiogenome.org/report-on-assembly-standards (EBP Subcommittee for Sequencing and Assembly 2026) https://www.genome.gov/about-genomics/fact-sheets/DNA-Sequencing-Costs-Data (Wetterstrand 2025) https://goat.genomehubs.org/ (Challis et al. 2023) |
| Sourcing and vouchering | Quantitative:
|
Lawniczak et al. (2022, 2023, 2025), CBD COP10 (2010), Sherkow et al. (2022), Smee (2020), McCartney et al. (2021), and Corrales and Astrin (2023) |
https://www.earthbiogenome.org/sample-collection-processing-standards (Lawniczak et al 2025.) https://www.protocols.io (Teytelman et al. 2016) |
| Sample preparation and sequencing | Quantitative:
|
Lawniczak et al. (2022, 2025), Howard et al. (2025), and Mueller et al. (2016) |
https://www.earthbiogenome.org/report-on-assembly-recommendations (EBP Subcommittee for Sequencing and Assembly 2026b) https://www.earthbiogenome.org/report-on-annotation-standards (Martin et al. 2023) https://www.protocols.io (Teytelman et al. 2016) https://tinyurl.com/EBP-Anecdotes |
| Genome assembly | Quantitative:
|
Lawniczak et al. (2022), Howe et al. (2021), Guiglielmoni (2025), Wang and Wang (2023), Rhie et al. (2021), and Cannon et al. (2025) |
https://www.earthbiogenome.org/report-on-assembly-standards (EBP Subcommittee for Sequencing and Assembly 2026) https://www.earthbiogenome.org/report-on-assembly-recommendations (EBP Subcommittee for Sequencing and Assembly 2026b) https://pipelines.tol.sanger.ac.uk/treeval https://gitlab.com/wtsi-grit/rapid-curation |
| Analysis reproducibility | Quantitative:
|
Lawniczak et al. (2022), Ziemann et al. (2023), Goble et al. (2021), and Wilkinson et al. (2016, 2025) |
https://www.earthbiogenome.org/it-and-informatics-standards (Poelchau et al. 2026) https://snakemake.readthedocs.io/en/stable/ (Mölder et al. 2021) https://www.nextflow.io/ (Di Tommaso et al. 2017) |
| Databasing | Quantitative:
|
Lawniczak et al. (2022) and Wilkinson et al. (2016) |
https://www.earthbiogenome.org/data-sharing-management-best-practices (McCartney et al. 2022) https://www.earthbiogenome.org/it-and-informatics-standards (Poelchau et al. 2026) https://www.insdc.org/ (Karsch-Mizrachi et al. 2025) https://i5k.nal.usda.gov/ (Poelchau et al. 2015) |
| Genome annotation | Quantitative:
|
Lawniczak et al. (2022) |
https://www.earthbiogenome.org/report-on-annotation-standards (Martin et al. 2023) https://www.earthbiogenome.org/report-on-annotation-recommended-tools (EBP Subcommitee for Annotation 2026) |
Table 2.
Current gaps in arthropod genomics standards and proposed action items to close gaps.
| Genome Project key stage | Gap summary | Action item |
|---|---|---|
| Budgeting | Unpredictability of genome project costs | Resources/training for estimating costs of a genome project |
| Budgeting | Redundancies in sequencing activities | Awareness of current arthropod sequencing projects to reduce redundancy (eg resources such as GoaT; https://goat.genomehubs.org/) |
| Budgeting | Limited knowledge of budgeting for data management activities | Resources/training for Data Management Plans (DMPs) |
| Sourcing and vouchering | Challenges with sample preservation | Documentation of preservation and extraction methods on protocols.io |
| Sourcing and vouchering | Inconsistent vouchering | Review best practices, development of a voucher preservation field kit |
| Sample preparation and sequencing | Lack of extraction quality metrics | Awareness of and contributions to sequencing protocols on protocols.io and anecdotes on https://tinyurl.com/EBP-Anecdotes |
| Sample preparation and sequencing | Tissue selection (sample pooling, taxon-Specific sequencing targets, RNAseq tissue diversity) | Resources/training for genome sequencing best practices, research efforts for annotation best practices |
| Genome assembly | T2T assembly: sequencing of telomeres and centromeres is still challenging | Improvements in long-read sequencing technologies and assembly algorithms |
| Genome assembly | Challenges with locating assembly quality metrics in publications | Develop and use a consistent, concise notation for the most important assembly contiguity/accuracy metrics; consult GoaT |
| Genome assembly | Lack of consistent assembly evaluation metrics used in genome notes publications | Task force to coordinate with journals to refine reporting requirements |
| Genome assembly | Lack of standardized curation pipelines | Resources/training, automation of curation pipelines |
| Analysis reproducibility | Limited reproducibility of researcher analytical pipelines | Increased adoption of workflow systems (eg Galaxy, Nextflow, Snakemake) |
| Databasing | Siloed data in community databases, lack of standardized connections between databases | Task force to understand researcher behavior to maximize database discoverability |
| Databasing | Need for interoperability of data across databases | Submission to INSDC in standard formats; continued development of tools with standard APIs |
| Genome annotation | Inconsistent prediction of noncoding regions and use of functional annotation | Continued development of automated pipelines to combine complex annotation tasks |
| Genome annotation | Challenges with annotation for certain gene families (eg chemosensory genes, selenoproteins) | Resources/training for manual curation, improvements to automated prediction tools |
| Genome annotation | Lack of community standards for annotation quality | Task force to identify and update measurable quality metrics |
| Genome annotation | Lack of standardization in GFF/GTF format | Disseminate minimum requirements for annotation submission to public databases |
| Genome annotation | Lack of consistent submission to INSDC | Resources/training for best practices, coordinate with journals to consider requiring annotation submission |
Materials and methods
We downloaded summary information for arthropod genomes submitted to INSDC between 2015 and 2025 using NCBI Datasets v18.2.2 (O’Leary et al. 2024). We used a custom R script to plot submission rate and quality statistics (Supplementary File 1).
In October 2024, we virtually convened a meeting of 21 individuals with expertise in arthropod genomics to identify gaps in existing genome project standards and reporting practices (Supplementary Table 1). We divided the individuals into groups based on their expertise: a “budgeting and sample source” group, a “sequencing, assembly, and analyses” group, and an “annotation and databasing” group. We tasked each group with answering a standard set of questions designed to identify existing standards, any gaps in these standards, their relevance to arthropods, their relevance to the arthropod research community, and their importance (see Supplementary File 2). Their answers form the basis of the standards and gaps presented in this paper.
To assess the standards reported for newly sequenced genomes in the arthropod community, we conducted a meta-analysis of arthropod genome papers published in 2024. We identified arthropod genome publications from 2024 by searching PubMed using the query: ((((arthropod genome assembly[Title/Abstract]) OR (arthropod genome sequence[Title/Abstract])) OR (insect genome assembly[Title/Abstract])) OR (insect genome sequence[Title/Abstract])) AND ((“2024/01/01”[Date - Publication]: “2025/01/01”[Date - Publication])) (search conducted 2.10.2025). In 399 publications (Supplementary Table 2), 191 were published by the DToL project as “genome notes” in Wellcome Open Research. Since the DToL project uses a specific template for genome notes explicitly designed to meet reporting standards (Threlfall and Blaxter 2021), we assessed these papers separately. We randomly selected 50 DToL genome notes (Supplementary Table 3) and 50 arthropod genome papers published in other journals for review (Supplementary Table 4). We omitted publications that did not have an obvious connection to arthropod genomes or publications that reported assembly updates without full assembly pipelines (eg using long-range data to scaffold a previously assembled genome).
For each genome paper, we evaluated the completion (YES/NO) of standards in multiple categories corresponding to phases of a genome assembly project (Fig. 2; Supplementary Tables 3–5). We evaluated the reporting of standards rather than the achievement of a certain standard value (Supplementary Table 6). For example, we counted the reporting of any BUSCO/OMArk score performed on genome sequences as “YES” for the metric “BUSCO (genome mode),” even though some consortia target a certain minimum value. As a result, we measured whether readers can ascertain the quality of genome assemblies through the written content of publications. We used a custom R script to plot reporting rates (Supplementary File 1).
Fig. 2.

Frequency of genome assembly standards reporting in arthropod publications from 2024. For each standard, the reporting frequency is shown for 50 Darwin Tree of Life genome notes (DToL) and 50 publications from all other journals (Other). The extended description of the evaluation criteria used to determine reporting of standards is available in Supplementary Table 5.
Results and discussion
Budgeting
A robust project plan underpins every successful genome project, and a clear, detailed budget is central to that plan. Although the importance of standardizing budgeting practices may not seem obvious at first, inconsistent or incomplete cost estimates can compromise assembly quality and limit future use of genomic resources. Establishing and sharing budgeting practices can streamline resource allocation and help researchers anticipate challenges.
Current standards
Sequencing a new genome typically costs anywhere from a few hundred to several thousand dollars (Lou et al. 2021; Wetterstrand 2025; EBP Subcommittee for Sequencing and Assembly 2026), largely depending on the size of the genome and the choice of sequencing platform. In addition to the costs of sequencing, projects must also include estimates for the cost of collection, shipping and storage of specimens, the cost of computation (including data quality assessment, assembly, annotation, and downstream analyses), and the cost of staff effort through the project (Table 1).
Current gaps
The arthropod community would benefit from guidance and training to clarify the major sources of genome project costs, for example:
Field and vouchering expenses: the costs of permits, travel, specimen shipment (with or without cold chain), and specimen curation, including taxonomic work.
Extraction and library preparation: the costs and benefits of optimizing protocols for tiny or challenging taxa. Examples of optimization strategies with successful outcomes in arthropods include isopods (Isopoda) and jumping spiders (Salticidae) (Howard et al. 2025).
Genome-specific factors associated with data generation: genome size, ploidy, heterozygosity, repeat content, presence of sex chromosomes, and expected proportion of nontarget (cobiont) DNA. In arthropods, large repetitive genomes in Orthoptera (grasshoppers, locusts, and crickets) and Euphausiacea (marine crustaceans) (Alfsnes et al. 2017) and whole genome duplication events in horseshoe crabs (Thomas et al. 2024) are examples of genome factors that can complicate assembly and increase costs.
Choices of sequencing technology: selection of ONT or PacBio or both for long reads, which version of chromatin conformation capture or other long-range technology to deploy.
RNA-seq requirements: multiple sexes, tissues, and/or life stages for genome annotation, selection of sequencing technologies (short reads, long reads, combination). Specialized tissues in arthropods are desirable (eg antennae, venom glands), but access to limited samples may preclude the ability to produce sufficient volumes of RNA for multiple tissue types.
Compute infrastructure: secure data storage and CPU/GPU time for assembly, annotation, and other analyses.
Staff costs: estimating the cost of expert staff required for each stage of the process, including considerations of skill levels required.
Although data generation costs will evolve with technology, budgets should include contingency funds to absorb unexpected failures.
Developing common project cost documentation and training, and encouraging authors to share funding breakdowns, would improve understanding in the community of project planning and coordination (Table 2). In parallel, the compilation and up-to-date maintenance of a comprehensive database of funding sources, including grant opportunities, project funding, and the missions of funding organizations, would better direct dynamic planning of efforts.
Sharing project data can also help to address another project planning gap, namely the lack of awareness of ongoing arthropod genome sequencing projects (Table 2). Increased discoverability of project statuses and resources would help to prioritize budgeting for underrepresented organisms and sequencing data types and minimize redundant research efforts. The Genomes on a Tree (GoaT) database (Challis et al. 2023) was established to provide this kind of platform for coordination, and has been adopted by the EBP to coordinate genome sequencing efforts globally (see https://goat.genomehubs.org/help/use_cases). GoaT collates information on genome quality metrics, genome sizes, and karyotypes across Eukaryota. These data are used to estimate values for species where direct information is lacking. Researchers can also search GoaT for a species’ genome project status, based on reports provided by a wide range of initiatives worldwide, including the projects affiliated with the EBP. The i5k has its own GoaT page (https://goat.genomehubs.org/projects/i5K), and the i5k is currently developing a mechanism to assist arthropod researchers to announce their project progress on GoaT.
Successful applications for the arthropod genome project funding may be determined in part by the potential of data reuse in future studies. Many global funding agencies require DMPs in grant proposals to describe how data will be used during and after a project. Yet many researchers lack templates or training to craft and implement effective DMPs (Table 2). Developing community-endorsed DMP templates and workshops would ensure that data sharing and archiving costs are properly budgeted and executed. Funders should continue to support line items explicitly for data management and long-term accessibility (Byrd et al. 2020).
Sourcing and vouchering
Once funding is secured, adherence to sample acquisition and processing best practices is essential to generating reference-grade arthropod genomes. Sourcing specimens for sequencing requires consideration of availability, ethical and legal restrictions, and cost. Of these factors, availability most strongly determines project difficulty: species with small ranges, brief lifespans, or remote habitats demand more expertise, time, cost, and paperwork than those readily obtained from cultures or suppliers. Choices of strain, sex, life stage, and tissue also affect extraction success and ultimately may influence ploidy, sex-chromosome representation, and heterozygosity in the final assembly.
Current standards
Permitting and documentation requirements that govern the collection and shipment of specimens, as well as ethical sampling issues, are complex and highly variable by location. Minimal consideration must include Nagoya protocols (CDB COP10 2010) if there is to be transport across national borders, and the conservation and pest status of target species. These issues are well covered in best practice guidelines made available through large sequencing consortia, including the EBP, DToL, and ERGA (Table 1) (Smee 2020; McCartney et al. 2021; Sherkow et al. 2022; Lawniczak et al. 2025). Best practices for confirming species identity, recording metadata, vouchering, and preserving specimens are documented in EBP reports and allied resources (Table 1) (Lawniczak et al. 2022, 2023, 2025; Corrales and Astrin 2023; Reichel et al. 2026).
Current gaps
Although detailed best practices exist, their implementation is uneven. We highlight two areas for improved guidance: expanded testing of alternatives to cold-chain dependent sample preservation methods and taxa-specific vouchering recommendations (Table 2).
Methods that yield high-molecular-weight DNA without the need for preservation and shipping on dry ice or in liquid nitrogen (eg ethanol perfusion) have not been systematically tested across diverse arthropods. For example, the small bodies of mosquitoes make them easy to squish and perfuse with ethanol in a small tube without greatly diluting the DNA (Teltscher and Lawniczak 2023). To yield a similar success rate for species with drastically different body compositions, such as large chitinous beetles, existing methods can be adapted or alternative methods developed. Sharing both protocols and success-rate metadata (including taxon and preservation details) on platforms like protocols.io (https://www.protocols.io) (Teytelman et al. 2016) will build a community knowledge base and refine methods for different arthropod groups.
Vouchering, the preservation of a portion or representation of the specimen used to generate a genome assembly, is a best practice that is often neglected. Vouchers submitted to a registered repository (eg Global Genome Biodiversity Network data portal; https://www.ggbn.org/ggbn_portal/) (Droege et al. 2014) that is well-equipped to physically preserve them, maintain associated metadata records, and coordinate their availability are important to future research, especially efforts to link physical characteristics to genomic traits. Proper vouchering also provides evidence of compliance with legal and ethical collection protocols.
In a best-case scenario, voucher specimens are large enough to be preserved mostly intact, and only a subsample of tissue is needed for sequencing. When most or all of the specimens may be consumed in the sequence generation process, photographs of the specimen may serve as the voucher, along with remaining tissue or extracted DNA if available. The DToL project has a robust system to collect photographs of specimens as vouchers (Lawniczak et al. 2023). Because specimens are often damaged during collection and shipping, it is critical the photographs be taken at the time of sample collection before preservation. This is a challenge for collectors who are likely to have limited resources in the field and may not be aware of what the defining characteristics of the species are that should be included in the photograph.
Alongside photography, when multiple, presumably conspecific, specimens can be sourced in the same collection event, preservation of an exemplar specimen can serve as a voucher by proxy to ensure an intact physical example of the species is available for future study. Exemplar specimens are particularly valuable when most or all of a voucher specimen is used for sequence generation. Barcode or skim sequencing of exemplar specimens through nondestructive sampling will allow exemplars to be confidently linked to the voucher by sequence comparison to the genome assembly.
Unfortunately, the disconnect between collections experts and sequencing centers (including the complications of field collection and timely tissue preservation vs ideal voucher preservation in an equipped lab setting) means that vouchers in any form (specimen or portion thereof, photos, and/or exemplars) are not always appropriately prioritized, preserved, or submitted to a registered repository. We identified two recommendations to increase compliance with voucher guidelines.
First, consistent use of a voucher photography station would help increase compliance. By including 1-mm-sized scale paper with a standard color grid (see https://zenodo.org/records/15235194 for an example), field specimen photos can be collected with reference information before the specimen is preserved for shipping. Compilation of sedation methods for arthropod photography, for example, chilling and methods to generate CO2 with ingredients that can be easily obtained and safely transported, would aid efficient imaging efforts by collectors.
Second, efforts are needed to identify and bring together taxonomic and museum experts to review and improve the published best practices for collection and sequencing. Ideally, any tissue, body parts, or DNA from the voucher that remain after extraction for sequencing should be preserved. At a minimum, photographs of the voucher should be deposited in secure repositories, either within recognized museum catalogs or in open databases such as Bioimage Archive (https://www.ebi.ac.uk/bioimage-archive/), along with detailed, Darwin Core-compliant metadata. Any available exemplar specimens collected in the same location should also be deposited in a recognized museum collection along with metadata connecting the exemplar to voucher remnants, photographs, and/or sequencing data and assemblies. Taxon-specific vouchering guidance and considerations for field-collection limitations are also needed, for example, whether there are specific body parts that are generally used to distinguish different species and therefore should be prioritized for high-resolution photography. This taxon-specific guidance should then be maintained in a central, free and easily searchable location like protocols.io. Additionally, journals and sequence data repositories should encourage researchers to include voucher information with their submission to increase compliance.
Sample preparation and sequencing
Whole-genome shotgun sequencing technology has evolved rapidly over the past 20 years. The transitions from chain-termination (Sanger) reads to high-throughput short reads to high-throughput long reads have generated substantial improvements in the contiguity, accuracy, and cost of genome sequencing (Li and Durbin 2024).
Current standards
The EBP has released guidelines for specimen processing for sequencing, with many examples from insects, and specific considerations for very small specimens (Table 1) (Lawniczak et al. 2025). The DToL project has released wet lab protocols for sample preparation and extraction, with some tailored toward insects or arthropods (Howard et al. 2025). For example, the Picogramme input Multi-modal Sequencing (PiMmS) protocol is specifically designed for extracting and PCR-amplifying low-bias, long-read genomic shotgun libraries from small specimens (Laumer 2023). The EBP sequencing and assembly subcommittee has also recently released genome and transcriptome sequencing strategies that should work for many organisms, including arthropods (Table 1) (Martin et al. 2023; EBP Subcommittee for Sequencing and Assembly 2026). Updated recommendations for sequencing strategies are common due to improvements in the accuracy and cost-effectiveness of technologies. For example, the incorporation of long reads for sequencing full-length cDNA has proved to be powerful for gene prediction, especially in discerning alternatively spliced isoforms deriving from a single gene (Pardo-Palacios et al. 2024).
We note that reference genome sequencing initiatives typically focus first on species where ample amounts of high-quality DNA are available, whereas species with a single physical specimen as the definitive “name-bearer” (ie holotype) require preservation of the morphological integrity of the specimen by noninvasive sequencing methods that may yield insufficient DNA for long-read sequencing (Heckenhauer et al. 2023). Fragmented, short-read-only assemblies still fit our definition of a reference genome in these circumstances (see Introduction), and many of the standards discussed here can be reported for genomes regardless of assembly contiguity.
Current gaps
Major gaps in sample preparation standards include the lack of established nucleic acid extraction quality metrics and the need for improved guidance on specimen and tissue selection (Table 2). Neither of these issues is necessarily restricted to arthropods.
Nucleic acid extraction is a key step in the genome sequencing workflow, yet metrics that reliably inform sequencing success are lacking. Purity of the DNA or RNA sample can be assessed through UV spectrophotometry, and molecule length through electrophoresis. Automated analytical platforms report DNA (or RNA) Integrity Numbers (DIN or RIN), a metric that integrates length distributions (and for RNA, the intactness of the ribosomal RNAs) (Mueller et al. 2016). Although DNA molecules of sufficient length are necessary for long-read sequencing, a simple length measurement does not identify “damaged” DNA containing nicks and covalent adducts that result in very poor long-read sequencing performance, and spectrophotometric estimates of DNA purity can similarly miss critical contamination. The Sanger Tree of Life programme (ToL) has published a compendium of extraction protocols based on experience with thousands of species across Eukaryota, including many arthropods, and they offer advice on interpretation of QC measures (Denton et al. 2024).
A second gap is sample and tissue selection. Selecting the number of specimens to sequence, which specimens, and which tissues to include or exclude is not always straightforward, given the life history diversity inherent to the phylum. For arthropods, the following factors complicate sample and tissue selection for nucleotide extraction:
Specimen size. Sampling from the same individual to generate all libraries for long-read genome sequencing is recommended. However, a single arthropod specimen may not provide sufficient DNA for all libraries due to its small size.
Diversity of sex determination systems. Sex determination mechanisms are highly diverse across arthropods (Blackmon et al. 2017) and are not always known for a given species. However, the sex of the sequenced individual impacts the sequencing result. For example, in species with genetic heterogametic sex determination (XY or ZW), sequencing the heterogametic sex permits assembly of both sex chromosomes. In Hymenoptera, on the other hand, sequencing single haploid male hymenopterans avoids the need to identify and eliminate haplotypic duplication in the assembly.
Heterozygosity. Most wild arthropods are likely to be highly heterozygous given their large census population sizes. Genome assembly may be problematic if libraries are derived from multiple specimens from highly heterozygous species.
Microbiota. Arthropod digestive systems carry rich microbiomes, and specimens taken from the wild are often infected with pathogens and parasites. Any extraction of material that also samples from this unknown community of cobionts will include their DNA and RNA, and sequencing volumes will have to be adjusted to assure sufficient coverage of the target arthropod. While this “contamination” may be significant, the sequence data also allows the assembly of the genomes of the microbes associated with arthropods (Vancaester and Blaxter 2023, 2024). Thus, researchers must decide whether to remove guts prior to sequencing.
The EBP proposes two current mechanisms to help close these gaps. The first is to use protocols.io (https://www.protocols.io, which we additionally recommend in the “Sourcing and vouchering” section) to document protocols that lead to successful high-molecular-weight DNA extraction and genome sequencing, as well as specimen and tissue selection. Protocols.io allows commenting and “forking” of protocols, so that they can be adapted for other organisms. The second mechanism is to collect “anecdotes” of protocols that did or did not work for specific taxa and steps from specimen preservation to sequencing (https://tinyurl.com/EBP-Anecdotes). While the information in this sheet is less structured than in protocols.io, it provides a low barrier to entry for researchers to share information, including protocols that do not work for particular taxa or situations. We recommend that arthropod researchers contribute taxon-specific extraction methods and metrics, as well as tissue and specimen selection, on one or both platforms. Importantly, researchers should include not just successful protocols, but also identify those that have failed when using high-quality tissue. Detailed commentary on these methods by the arthropod genomics community could provide sorely needed information on methods to use or avoid for specific taxa.
Genome assembly
The ideal standard for genome assembly is a T2T representation containing the entire set of chromosomes without gaps or structural assembly errors and a quality value (QV) score of 60 or greater. However, this is currently costly, and successful attempts are nearly exclusively derived from cell lines. For standard biodiversity samples, a combination of both long-read and long-range data is now routinely used to generate chromosome-level assemblies. There are multiple types of graph-based assembly algorithms (Li and Durbin 2024) in addition to dedicated organelle assemblers (Dierckxsens et al. 2017; Uliano-Silva et al. 2023; Zhou et al. 2025); recommendations are out-of-scope for this analysis, but current examples are available at (EBP Subcommittee for Sequencing and Assembly 2026b). Following genome assembly, manual curation should be performed to improve assembly quality (Howe et al. 2021). A recently published insect assembly protocol is available (Guiglielmoni 2025).
Current standards
Unlike other key stages of assembly projects, quantitative standards for assembly are more well-defined (Table 1), and the EBP has proposed values defining high-quality assemblies (Blaxter et al. 2025; EBP Subcommittee for Sequencing and Assembly 2026). Measurable quality metrics for assembled genomes fall into three broad categories: contiguity, completeness, and correctness (Howe et al. 2021; Rhie et al. 2021; Wang and Wang 2023). The following should be measured for all new genomes (current EBP recommended values shown in parentheses):
contig N50/NG50 (1 Mbp)
scaffold NG50 (chromosome-level scaffolding)
percentage of contigs placed as chromosomes (90%)
k-mer completeness of the assembly (90%)
conserved gene-space completeness (eg BUSCO, OMArk) (90%)
QV score (Q40)
false haplotypic duplications (<5%)
transcript mappability (90%)
Additional procedural standards recommended by the EBP include resolution of the assembly with the species karyotype when possible, comparing total size against a size estimated using orthogonal k-mer-based (Ranallo-Benavidez et al. 2020) and/or flow cytometry (Johnston et al. 2019) methods, resolution of large-scale structural errors using corroborating evidence, the assembly and annotation of organelle genomes with specialized pipelines, and the separation of sequences corresponding to contaminants and cobionts. Arthropods often harbor biologically important cobionts (eg endosymbiotic bacteria) that often end up as contaminating sequences in assemblies submitted to public databases (Astashyn et al. 2024). Contaminants should be distinguished from legitimate horizontal gene transfer (HGT) events, which are known in insects harboring bacterial endosymbionts (Xing et al. 2023). Putative HGT events can be supported using some contamination detection tools (Astashyn et al. 2024) and/or read-mapping patterns across host-integrant junctions.
Curation steps should be completed prior to genome annotation and assembly submission. Researchers should identify and name chromosome units in the assembly, distinguishing them from unlocalized and unplaced sequences. Naming should follow the conventions used for new assemblies of the same species, otherwise they should be ordered numerically by size (Howe et al. 2021). Orientation should follow the p-to-q arm direction, made possible by identifying centromeric regions with tools such as ModDotPlot (Sweeten et al. 2024) and AniAnns (Sweeten et al. 2026). Researchers should take care with naming or orientation conventions that conflict with well-established reference genomes for the species, as such changes can complicate comparisons with the substantial body of published work and databasing built on those references. For long repeats that are challenging to assemble, having 1 to 2 representative copies of the repeat at the gap boundary and an annotation of the gap defining the repeat type can be useful for downstream analysis.
When submitting diploid assemblies, researchers should define the primary assembly and separate from the secondary assembly containing the alternate haplotype. Researchers should clearly define whether trio data is used to phase the whole genome (maternal + paternal haplotypes, phase concordance across chromosomes) or long-read/long-range data used to phase each chromosome independently (haplotype 1 + haplotype 2, no phase concordance across chromosomes). The assembly names themselves should contain unique identifiers (Cannon et al. 2025; Lawniczak et al. 2025). In a recent study reporting multiple new T2T genomes (Yoo et al. 2025), the haplotype assembly that was more complete (fewer gaps with both telomeres present) and higher quality (higher QV) was chosen as the primary assembly, unless only one haplotype contained the rDNA array (then it was chosen as primary regardless of the completeness criteria).
When possible, the heterogametic sex should be sequenced for representation of all sex chromosomes. Sometimes assembly algorithms will separate sex chromosomes across the two haplotypes of a genome, but for downstream analyses, it is often desired that the reference genome contains all sex chromosomes in haploid form. Therefore, both sex chromosomes should be added to one representation (primary or hap1) of the genome. For trio-phased assemblies either the maternal or paternal haplotype can be chosen. However, when sex chromosomes are derived from assemblies of different individuals, a synthetic genome should be constructed that is separate from the assemblies of the constituent individuals and contains both sex chromosomes.
Current gaps
True T2T status in eukaryote genomes requires suitable starting material and considerable sequencing and curation efforts to span repetitive regions (Table 2). There is a need to identify pipelines to consistently achieve T2T status for diverse species. In arthropods these efforts will be complicated by extensive variation in genome size, repeat content, and sex chromosome number and identity. Gigabase-scale genomes with more than 50% repeat content are common in many arthropod lineages, driven by expansions of often uncharacterized transposable elements (Petersen et al. 2019; Wu and Lu 2019); scaffolding across these large heterochromatic blocks requires long range data at depths and read lengths that may not be possible to generate from field specimens using currently available technology, and annotation of repetitive elements often lags far behind gene models in euchromatic sequence. Sex chromosome copy number variation (eg in crustaceans [Lécher et al. 1995], or spiders [Sember et al. 2020]) and rapid turnover of neo-sex chromosomes (Blackmon et al. 2017) create sequences with variable and unpredictable patterns of divergence from other genomic regions, and will require careful analysis and innovative methods to identify and assemble unambiguously (eg in ticks, Nuss et al. 2023; Tidwell et al. 2024). Bridging these assembly challenges requires adoption of new technologies and computational methods and elevates the importance of clear documentation and consistent reporting of maximally informative metrics for new assembly outputs.
Although most assembly standards are well-developed, it is critical to formulate a method to consistently communicate assembly quality with as few metrics as possible (Table 2). For example, the EBP uses the notation x.y to quantify assembly contiguity, where: x = log10[contig NG50], y = log10[scaffold NG50] (EBP Subcommittee for Sequencing and Assembly 2026). Support for this notation is available in GoaT (Challis et al. 2023), and increased use of this type of notation in publications will facilitate rapid assessment of the overall genome quality including comparative analyses of multiple assemblies within the same species or between closely related species.
As manual curation can affect contiguity metrics, researchers should record details on any post-assembly curation steps taken to provide justification for the structure of the final assembly (Table 2). There is a need for increased standardization in assembly curation. Future strategies should focus on automation and transparent documentation of curation activities. Examples include the ToL TreeVal pipeline (https://pipelines.tol.sanger.ac.uk/treeval) that automates the production of several data types used in curation which in turn can be assessed following best practices documentation to correct large-scale assembly errors (https://gitlab.com/wtsi-grit/rapid-curation), as well as Verkko-Fillet (Kim et al. 2025) which provides curation tools to improve graph-based assemblies. Promoting these pipelines and subsequent training opportunities for pipeline use and interpretation will increase the consistency of curation practices.
Analysis reproducibility
Critical to the scientific method is the ability to reproduce previous results given the same input data and analytical techniques (Goodman et al. 2016). Surveys of researchers across scientific disciplines reported widespread reproducibility issues in published studies (Baker 2016; Cobey et al. 2024). Missing data, code and/or documentation, broken software dependencies, and buggy code can contribute to the inability to reproduce complete computational biology workflows (Reiter et al. 2021; Ziemann et al. 2023).
Current standards
Automated workflows, code versioning, and containerization are key components of improving analysis reproducibility (Table 1) (Ziemann et al. 2023; Poelchau et al. 2026). Automated workflows for sequencing analysis, genome assembly, and post-assembly processing are available on Galaxy (Larivière et al. 2024) and Nextflow (Di Tommaso et al. 2017) (eg https://pipelines.tol.sanger.ac.uk/pipelines; Blaxter et al. 2025). Sharing workflows in publications and linking to repositories at platforms such as WorkflowHub (Goble et al. 2021), would greatly enhance reproducibility. FAIR (Findable, Accessible, Interoperable, Reusable) principles for research workflows have recently been developed by the FAIR Workflows Working Group (WCI-FW) (Wilkinson et al. 2025) which could both promote analysis reproducibility as well as provide concrete metrics to evaluate whether an analysis is reproducible. Finally, peer review at scientific journals provides the most straightforward mechanism to enforce reproducibility. Adoption of guidelines such as those developed by the WCI-FW should facilitate peer review.
Current gaps
Our expert panel noted a procedural gap in the insect genomics community—specifically, that the use and publication of workflows needs to become more routine (Table 2). While awareness of analysis reproducibility has increased, scientists may still struggle to adopt best practices in making their analyses reproducible given the rapidly changing nature of bioinformatics software and workflows. Although the adoption of workflow management frameworks such as Snakemake (Mölder et al. 2021) and Nextflow (Di Tommaso et al. 2017) are greatly improving analysis reproducibility, their effective use requires an initial learning investment, and it is unclear whether access to training and workflow software is globally universal.
One possible way to close this gap is a short-term task force, eg, via the i5k initiative's working groups, that assesses the arthropod genomics community’s ability to create reproducible workflows. Assessment outcomes would reinforce standards for reproducibility as well as provide recommendations on where investments are needed to ensure analysis reproducibility in insect genomics globally.
Databasing
To ensure the optimal use of the exponentially growing body of genomics data by the larger research community, data should be deposited into public databases. Funders may mandate data sharing standards.
Current standards
Stored data should support the FAIR principles for scientific data (Wilkinson et al. 2016). New reference genome assemblies, the raw data that was used to generate them, and sample and assembly metadata should be submitted to the International Nucleotide Sequence Database Collaboration (INSDC, comprising GenBank, the European Nucleotide Archive and the DNA Databank of Japan). INSDC members support the databasing of supporting data and metadata in standardized formats and share data nightly. The EBP has provided guidance on data management and sharing best practices (McCartney et al. 2022) as well as recommended databases for depositing basic genomics data types (Poelchau et al. 2026). There are a few reasons why this recommendation may not be met by arthropod genomics researchers. First, individual nations might have policies that restrict depositing on INSDC. Second, the requirement from funders or journals to deposit certain data types on INSDC databases may not be enforced, despite being part of an internationally agreed journal standard. Third, researchers may be challenged by the metadata deposition requirements of INSDC. Despite the necessary rigor imposed by correct and complete metadata and data submission to INSDC databases, arthropod genomics researchers should aspire to these standards if data are truly to be FAIR compliant.
RefSeq at NCBI (Pruitt et al. 2000; Goldfarb et al. 2025) and Ensembl at EBI (Dyer et al. 2025) offer additional layers of data representation, in particular of gene and repeat annotations of analyzed genomes. These representations are rich analytic platforms for the analysis of genes in their genomic context. Ensembl also performs clustering of gene products into COMPARA gene families and derived proteins are decorated with domain and motif functional annotations. These genome databases, and the INSDC databases in general, provide centralized and standardized data access, integration, and downloads across diverse taxa. Community-organized databases can provide a complementary way to share data and to collate collective understanding of the biology of the represented genomes. There are databases for well-studied model arthropods, including FlyBase (Drysdale 2008; Thurmond et al. 2018), VectorBase (Lawson et al. 2009; Giraldo-Calderón et al. 2022), iBeetleBase (Dönitz et al. 2015; Dönitz et al. 2018), and LepBase (Challi et al. 2016). The USDA i5K Workspace@NAL (Poelchau et al. 2015) hosts arthropod genomes and community annotations for the Ag100Pest Initiative (Childers et al. 2021), the Beenome100 Project (https://www.beenome100.org/), and several additional genome projects. Community-provided curated databases can also be a focused source of a single data type, such as repeats databases such as Repbase (Jurka et al. 2005; Bao et al. 2015) and Dfam (Storer et al. 2021), single cell transcriptomes at Fly Cell Atlas (Li et al. 2022) or regulatory DNA element catalogs at REDfly (Keränen et al. 2022), or contain specialized modalities such as InsectBase 2.0 (Mei et al. 2022) which includes HGT and RNA-RNA interactions. If INSDC or community-based databases cannot host specific data types, then generalist repositories (eg Dryad, FigShare, Zenodo) can provide secure storage of flat files. Stable repositories where data can be accessed using a persistent digital object identifier (DOI) are preferable over the use of institutional or private websites.
Current gaps
With the growing amount of data in both INSDC and community databases, the arthropod community would benefit from implementation of best practices for creating and maintaining connections between databases. This goal is complicated by the fact that there is no standard design or platform for community databases such that they are programmed with different functionalities in mind. Standardized interfaces are necessary to facilitate cross-database analyses. The NIH Comparative Genomics Resource (CGR) (Bornstein et al. 2023) provides organism-agnostic tools and resources to enhance INSDC submitted data and aims to provide standardized connections between NCBI data and community content. Command-line interfaces and APIs should support automated data download, such as gene, genome, and taxonomy data using NCBI Datasets (O’Leary et al. 2024). Interoperable identifiers such as ontology terms supported by the Open Biological and Biomedical Ontologies (Smith et al. 2007) are essential to provide connection points across databases.
Another current standards gap is the lack of synchronization of data between community-based resources and INSDC representations. For instance, whenever similar kinds of data such as genome sequences are hosted on multiple databases, each update in one database will lead to synchronization issues. Researchers should pay close attention to the frequency of database updates, and at the very minimum clear reporting of INSDC standard accessions (eg GCA_assembly accessions) including versions needs to be implemented such that users are aware of updates. Access to previous versions is essential for citing a version used for publications or replicating analyses. Synchronization issues also apply to metadata: although there are guidelines for requesting taxonomy identifiers when submitting genomes (Blaxter et al. 2024), taxonomic revisions may prevent updates to legacy datasets as submitters have the ownership of sample metadata, increasing metadata heterogeneity. Efforts are needed to generate standard operating procedures to reconcile discrepancies between taxonomy databases and INSDC.
Genome annotation
High-quality genome annotation includes the accurate prediction of structural boundaries of features and the subsequent assignment of functional information to features. Although there are a growing number of annotation pipelines available to the arthropod community, annotation may not be included in a new genome report, and there is a lack of standardization in producing, evaluating, and databasing annotations.
Current standards
The EBP has released a report on annotation standards discussing the key challenges and best practices for generating annotations (Martin et al. 2023), as well as a report on recommended tools for completing annotation tasks (Table 1) (EBP Subcommittee for Annotation 2026). The minimal EBP standard for genome annotation is the annotation of coding DNA sequence features of protein-coding genes (Martin et al. 2023). The prediction of additional feature types can enhance the quality of the annotation, including:
additional components of protein-coding genes (eg transcription start sites and untranslated regions)
noncoding RNAs
pseudogenes
regulatory sequences
repetitive elements
Quality assessments should be performed on completed genome annotations. A basic annotation statistics report should be included for comparisons to annotations from closely related species. The report should include the number of features for each type (protein-coding, noncoding RNA, pseudogenes) as well as noting potential quality errors, such the proportion of single-exon genes, the presence of frameshifted models, or ultrashort introns. BUSCO (Simão et al. 2015) should be run in protein mode on translated CDSs as a completeness metric in the annotation set.
Current gaps
Most current genome annotation tools focus on the annotation of protein-coding regions. Future development is required to optimize automated pipelines that handle multiple complex tasks in a comprehensive genome annotation, including the prediction of noncoding features and functional annotation. For example, EGAPx (https://github.com/ncbi/egapx) is a standalone version of the Eukaryotic Genome Annotation Pipeline (EGAP) used at NCBI. Unlike genome assembly, where there has been a concerted effort to generate best practices pipelines from the sequencing to final assembly, there has not been an equivalent effort in the annotation space, particularly for steps downstream of structural annotation. This process is underway through the Long-Read RNA-Seq Genome Annotation Assessment Project (LRGASP) (Pardo-Palacios et al. 2024). A “complete” genome annotation would account for the categorization of all functional elements in the genome sequence, including regions involved in transcriptional regulation. However, there is currently no standard for the proper identification of transcriptional regulatory sequences and their representation in annotation files.
Automated annotation pipelines have limitations for challenging gene families. In arthropods, a well-known example is the difficulty in annotating chemosensory gene families due to rapid diversification (Vizueta et al. 2018). Other examples of difficult-to-annotate gene families from arthropods include silk proteins (Frandsen et al. 2023) and venom (von Reumont et al. 2022). New tools may be valuable in addressing these cases (Vizueta et al. 2020), but resources and training to maintain manual curation expertise in the arthropod community are also needed (Poelchau et al. 2015). Current annotation challenges also include short, intronless, antisense, and chimeric transcript types as well as pseudogenes.
Unlike assembly, there is a general lack of quantitative annotation quality metrics. Assessing the presence of conserved genes in predicted gene sets using BUSCO (Simão et al. 2015), compleasm (Huang and Li 2023), or OMArk (Nevers et al. 2024) is the most commonly used method, with the latter able to detect taxonomically inconsistent proteins. However, methods to measure gene annotation completeness beyond conserved genes are lacking. PSAURON (Sommer et al. 2025) is a promising reference-free machine learning method that measures the likelihood of CDS or amino acid sequences to be genuinely protein-coding. PSAURON assigns a composite likelihood score to provide a general assessment of annotation quality as well as scores for individual sequences to potentially identify incorrect models. If a reference annotation from the same or similar species is available, measurable indicators such as the ratio of mono- to multi-exon genes and levels of sequence similarity can be reflective of annotation quality.
There is also a current gap in the annotation file format standard. Annotation tools typically produce output in General Feature Format (GFF) or Gene Transfer Format (GTF). Notably, there are inconsistencies in the feature types and attribute tags used to store information, making it difficult to write software that processes GTF or GFF3 files (Saha et al. 2022). For example, the phase field is often used incorrectly; the score field is ambiguous; and the role of the ID and Dbxref attributes is often misinterpreted (Saha et al. 2022). Documenting best practices for GFF3 formatting and usage is therefore essential.
Although genome annotation is often generated as part of a genome sequencing effort, it is also essential to include that annotation in the INSDC database submission to maximize its reuse. GenBank provides support for submitting annotation using GFF3 files through its table 2asn tool (https://www.ncbi.nlm.nih.gov/genbank/genomes_gff/), including logic to automatically handle many INSDC requirements such as locus_tag assignment for gene features. More frequent submission of arthropod annotation data to INSDC can help to further streamline the process. Some feature types might be better suited for specialized databases, such as the submission of repetitive elements to RepBase or Dfam. Although resources are available to support the curation (Goubert et al. 2022) and submission (Kohany et al. 2006) of repetitive elements, ultimately community-level efforts are needed to promote repetitive element data archiving activities.
Current procedural gaps in arthropod genomics
Our expert panel primarily identified persistent gaps in cases where a defined standard or metric does not yet exist. However, procedural gaps occur when existing standards are not consistently applied and reported in research papers. To assess whether the arthropod community is actively reporting standard metrics for newly sequenced genomes, we conducted a meta-analysis of arthropod genome papers published in 2024. We searched PubMed for open access papers (see “Materials and methods”) and randomly sampled 50 genome notes published in the Wellcome Open Research Tree of Life Gateway as part of the DToL project (hereafter, “DToL genome notes”) (Supplementary Table 3) and 50 arthropod genome papers from other journals (Supplementary Table 4). For each of these two categories we calculated the completion rate for standards in the seven key stages discussed above (Fig. 2).
DToL genome notes
DToL genome notes provide a detailed description of a newly assembled genome sequence including multiple quality assessments. Genome notes promote genome data discovery and reuse as well as give authorship credit to all researchers involved in the work (Threlfall and Blaxter 2021). DToL genome notes are produced using a Nextflow pipeline (https://github.com/sanger-tol/genomenote) to aggregate data and metadata to populate a draft template for genome notes. As a result, the reporting of standards in these genome notes is extremely consistent, with 100% completion rate in 26/36 measured standards. The genome notes had lower reporting of genome annotation activities: the completion rate of seven annotation standards was 50% or less in DToL genome notes, and 22/50 (44%) reported the public availability of annotations via Ensembl Rapid Release (Fig. 2; Supplementary Table 3). However, genome notes generated by DToL are intended as resource announcements and do not require the analysis and discussion of sequences relevant to organismal biology (eg repeat analyses or gene content). For DToL genomes, genome annotation is decoupled from genome submission, and annotations generated by Ensembl can be made public after the genome note is published.
Arthropod genome papers in other journals
In the 50 randomly sampled arthropod genome papers from other journals, reporting rates of standards were much less uniform: only 7/36 standards had 100% reporting rates and 10/36 had less than 50% reporting rates (Fig. 2; Supplementary Table 4). In particular, some post-assembly metrics proposed by the EBP (percentage of assembly in chromosomes, base accuracy, k-mer completeness) as well as standard curation steps recommended by the EBP (haplotype assembly/phasing, sex determination identification/assembly, organelle assembly, contamination screening) are not consistently reported by the broader arthropod genomics community. Other standards that show major differences between the DToL genome notes and other journals are the provision of ethics statements, voucher availability statements, and the use of workflow systems to conduct analyses. Genome annotation practices show the opposite pattern: protein-coding annotation is more commonly reported alongside the genome sequence (94%) relative to DToL genome notes (44%). Again, this is due to the focus of DToL genome notes on assembly. Other annotation activities have variable completion rates in non-DToL papers (40% to 86%). However, considering only papers with annotation reported, 77% (36/47) of annotations are publicly available and only 11% (5/47) of annotations are accessioned on centralized databases, whereas 100% of DToL genome notes (22/22) with annotation have accessioned data on Ensembl Rapid Release.
A subset of 38 (76%) non-DToL papers were published in journals with genome report formats, including 24 (48%) from the journal Scientific Data. There are differences in required or recommended analyses across these journals (Halfon et al. 2026) which may affect reporting patterns measured here. The sample sizes were too low to confidently assess standards reporting by journal, but a cursory evaluation indicated that individual journals (G3, GBE, Scientific Data) had similar reporting patterns as the set of 50 papers grouped together (data not shown). We note that G3, GBE, and Scientific Data all had 100% reporting rates of protein-coding annotation (Supplementary Table 4), although only G3 and GBE list annotation explicitly as a publication requirement.
Multiple factors contribute to procedural gaps
We hypothesize various reasons behind the observed patterns in procedural gaps. First, some standards are more commonly included as part of journal reporting requirements. As observed from a survey of “genome note” article types (Halfon et al. 2026), the submission of sequencing and assembly data is universally required, and there was near 100% compliance in the meta-analysis presented here (Fig. 2; Supplementary Table 5). Second, standards with more consistent reporting often reflect historical use and involve user-friendly metrics and tools. When researchers commonly produced draft-level assemblies using short reads, contig N50/NG50 was an appropriate metric to understand contiguity and compare between assemblies of the same or similar species. BUSCO software is easy to install and use and the output is relatively straightforward to interpret. These two metrics were only explicitly required in half of the genome note article types (Halfon et al. 2026), yet they both had 100% reporting rates in our meta-analysis (Fig. 2; Supplementary Table 5). Third, it is possible that some standards were performed but not reported in publications. This could be particularly true for manual curation activities. In the spirit of transparency and reproducibility, authors should strive to describe any changes in the underlying genome sequence data following assembly, including joins/splits, contamination removal, sex chromosome identification, and organelle genome assembly. Finally, related to the above, it is likely that some standards are intended to be reported as part of future studies. Unless genome annotation is explicitly required, authors may elect to limit the discussion of the genome report to the underlying assembly and its measured quality. The comparison and evolution of repeat elements and gene sets may involve a distinct set of researchers, and the timely release of genome reports for the sequence itself becomes necessary to encourage proper attribution as genomes are used for future studies. Currently the generation of a repeat or protein-coding gene annotation for a genome by a group distinct from the submitters of the genome sequence itself is designated a “third party annotation” and is not usually independently submittable to INSDC. This disconnect has meant that, for example, the protein sequences derived from Ensembl annotations of DToL arthropod genomes are not available in the NCBI protein sequence database for sequence similarity search, though they are made available in UniProt (UniProt Consortium 2018).
Bridging gaps in arthropod genome assembly standards
We have identified several standards gaps throughout the arthropod genome project cycle (Table 2). We highlight three high-priority gaps that can be addressed in the near future and would benefit the arthropod genomics community. First, we should address the current inconsistency of genome assembly metrics reported in publications (Fig. 2). This is largely a consequence of the lack of widely accepted community standards as well as inconsistent (or lack of) required reporting metrics in journals. We recommend a task force to liaison with journal editors to potentially modify reporting requirements to increase consistency. Second, we should promote the deposition of genome sequences, annotations, and rich metadata into INSDC databases whenever possible to increase discoverability and enhance comparative genomics efforts. Longer term goals should include increased adoption of automated pipelines to increase reproducibility and creating more robust connections between centralized databases and community databases with specialized data types. Third, we should provide more training opportunities for scientists to close gaps. In addition to enhancing mastery with tools and resources essential for high-quality genome assembly, training with experts in the field can help to create a better understanding of how to interpret standards measurements and access relevant information on public databases.
The continued development of new standards and the increased consistency in standards reporting in publications will improve the quality of arthropod genomes resources for the scientific community. The i5k standards working group will continue to seek opportunities for increasing community awareness and adoption of standards for newly sequenced arthropod genomes.
Supplementary Material
Acknowledgments
The authors are members of the i5k Working Group on Genome Project Standards. We thank experts in the arthropod genomics community (Supplementary Table 1) for contributions to the discussion of gaps in genome project standards. We thank Caroline Howard, Olga Vinnere Pettersson, Sheina Sim, and Mara Lawniczak for providing comments to improve the quality of the paper.
Contributor Information
Eric S Tvedte, National Center for Biotechnology Information, National Library of Medicine, National Institutes of Health, Bethesda, MD 20894, United States.
Gregor Bucher, Department of Evolutionary Developmental Genetics, University of Göttingen, Johann-Friedrich-Blumenbach Institute, GZMB, 37077 Göttingen, Germany.
David M Luecke, Veterinary Pest Genetics Research Unit, USDA, Agricultural Research Service, Kerrville, TX 78028, United States.
David C Molik, Center for Scholarly Publishing, Kansas State Libraries, Kansas State University, Manhattan, KS 66506, United States.
Terrence Sylvester, Department of Biological Sciences and Center for Biodiversity Research, University of Memphis, Memphis, TN 38152, United States.
Mark Blaxter, Tree of Life, Wellcome Sanger Institute, Cambridge CB10 1SA, United Kingdom.
Christine G Elsik, Divisions of Animal Sciences, University of Missouri, Columbia, MO 65211, United States.
Kerstin Howe, Tree of Life, Wellcome Sanger Institute, Cambridge CB10 1SA, United Kingdom.
Duane D McKenna, Department of Biological Sciences and Center for Biodiversity Research, University of Memphis, Memphis, TN 38152, United States.
Terence D Murphy, National Center for Biotechnology Information, National Library of Medicine, National Institutes of Health, Bethesda, MD 20894, United States.
Lukas Schrader, Institute for Evolution & Biodiversity, University of Münster, DE-48149 Muenster, Germany.
Cibele G Sotero-Caio, Tree of Life, Wellcome Sanger Institute, Cambridge CB10 1SA, United Kingdom.
Robert M Waterhouse, Environmental Bioinformatics Group, SIB Swiss Institute of Bioinformatics, 1015 Lausanne, Switzerland.
Anna K Childers, Bee Research Laboratory, Beltsville Agricultural Research Center, USDA, Agricultural Research Service, Beltsville, MD 20705, United States.
Marc S Halfon, Department of Biochemistry, University at Buffalo-State University of New York, Buffalo, NY 14203, United States; Department of Biomedical Informatics, University at Buffalo-State University of New York, Buffalo, NY 14203, United States; Department of Biological Sciences, University at Buffalo-State University of New York, Buffalo, NY 14260, United States.
Monica F Poelchau, National Agricultural Library, USDA, Agricultural Research Service, Beltsville, MD 20705, United States.
Data availability
The authors affirm that all data necessary to confirm the article’s conclusions are included in the article, the figures, and the supplementary files.
Supplemental material available at GENETICS online.
Funding
M.S.H. is supported by National Institutes of Health grant U24 GM142435. D.D.M. is supported by National Science Foundation grant DEB2110053. This work was supported in part by the U.S. Department of Agriculture, Agricultural Research Service. Mention of trade names or commercial products in this publication is solely for the purpose of providing specific information and does not imply recommendation or endorsement by the USDA. USDA is an equal opportunity provider and employer. This work was supported in part by the National Center for Biotechnology Information of the National Library of Medicine (NLM), National Institutes of Health (NIH). The contributions of the NIH author(s) are considered Works of the United States Government. The findings and conclusions presented in this paper are those of the author(s) and do not necessarily reflect the views of the NIH or the U.S. Department of Health and Human Services.
Literature cited
- Alfsnes K, Leinaas HP, Hessen DO. 2017. Genome size in arthropods; different roles of phylogeny, habitat and life history in insects and crustaceans. Ecol Evol. 7:5939–5947. 10.1002/ece3.3163. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Astashyn A et al. 2024. Rapid and sensitive detection of genome contamination at scale with FCS-GX. Genome Biol. 25:60. 10.1186/s13059-024-03198-7. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Baker M. 2016. 1,500 scientists lift the lid on reproducibility. Nature. 533:452–454. 10.1038/533452a. [DOI] [PubMed] [Google Scholar]
- Bao W, Kojima KK, Kohany O. 2015. Repbase update, a database of repetitive elements in eukaryotic genomes. Mob DNA. 6:11. 10.1186/s13100-015-0041-9. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Blackmon H, Ross L, Bachtrog D. 2017. Sex determination, sex chromosomes, and karyotype evolution in insects. J Hered. 108:78–93. 10.1093/jhered/esw047. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Blaxter M et al. 2025. The Earth BioGenome Project Phase II: illuminating the eukaryotic tree of life. Front Sci. 3:. 10.3389/fsci.2025.1514835. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Blaxter M, Pauperio J, Schoch C, Howe K. 2024. Taxonomy identifiers (TaxId) for biodiversity genomics: a guide to getting taxid for submission of data to public databases. Wellcome Open Res. 9:591. 10.12688/wellcomeopenres.22949.1. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Bornstein K, Gryan G, Chang ES, Marchler-Bauer A, Schneider VA. 2023. The NIH comparative genomics resource: addressing the promises and challenges of comparative genomics on human health. BMC Genomics. 24:575. 10.1186/s12864-023-09643-4. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Byrd JB, Greene AC, Prasad DV, Jiang X, Greene CS. 2020. Responsible, practical genomic data sharing that accelerates research. Nat Rev Genet. 21:615–629. 10.1038/s41576-020-0257-5. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Cannon EKS et al. 2025. Guidelines for gene and genome assembly nomenclature. Genetics. 229:iyaf006. 10.1093/genetics/iyaf006. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Challi RJ, Kumar S, Dasmahapatra KK, Jiggins CD, Blaxter M. 2016. Lepbase: the lepidopteran genome database [preprint]. bioRxiv 056994. 10.1101/056994. [DOI]
- Challis R, Kumar S, Sotero-Caio C, Brown M, Blaxter M. 2023. Genomes on a Tree (GoaT): a versatile, scalable search engine for genomic and sequencing project metadata across the eukaryotic tree of life. Wellcome Open Res. 8:24. 10.12688/wellcomeopenres.18658.1. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Childers AK et al. 2021. The USDA-ARS Ag100Pest initiative: high-quality genome assemblies for agricultural pest arthropod research. Insects. 12:626. 10.3390/insects12070626. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Cobey KD et al. 2024. Biomedical researchers’ perspectives on the reproducibility of research. PLoS Biol. 22:e3002870. 10.1371/journal.pbio.3002870. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Conference of the Parties to the Convention on Biological Diversity . 2010. Nagoya protocol on access to genetic resources and the fair and equitable sharing of benefits arising from their utilization [accessed 07.07.2026]. https://www.cbd.int/abs/doc/protocol/nagoya-protocol-en.pdf
- Corrales C, Astrin JJ. 2023. Biodiversity biobanking—a handbook on protocols and practices. Pensoft Publishers. Advanced Books. Vol. 1. [Google Scholar]
- Darwin Tree of Life Project Consortium . 2022. Sequence locally, think globally: the Darwin Tree of Life Project. Proc Natl Acad Sci U S A. 119:e2115642118. 10.1073/pnas.2115642118. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Denton A, Yatsenko H, Jay J, Houliston KE, Howard C. 2024. Sanger Tree of Life wet laboratory protocol collection v.2 [accessed 06.26.2026]. https://www.protocols.io/view/sanger-tree-of-life-wet-laboratory-protocol-collec-dtyf6ptn.pdf.
- Dierckxsens N, Mardulyn P, Smits G. 2017. NOVOPlasty: de novo assembly of organelle genomes from whole genome data. Nucleic Acids Res. 45:e18. 10.1093/nar/gkw955. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Di Tommaso P et al. 2017. Nextflow enables reproducible computational workflows. Nat Biotechnol. 35:316–319. 10.1038/nbt.3820. [DOI] [PubMed] [Google Scholar]
- Dönitz J et al. 2015. iBeetle-Base: a database for RNAi phenotypes in the red flour beetle Tribolium castaneum. Nucleic Acids Res. 43:D720–D725. 10.1093/nar/gku1054. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Dönitz J, Gerischer L, Hahnke S, Pfeiffer S, Bucher G. 2018. Expanded and updated data and a query pipeline for iBeetle-Base. Nucleic Acids Res. 46:D831–D835. 10.1093/nar/gkx984. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Droege G et al. 2014. The Global Genome Biodiversity Network (GGBN) data portal. Nucleic Acids Res. 42:D607–D612. 10.1093/nar/gkt928. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Drysdale R. 2008. Flybase. In: Dahmann C, editor. Drosophila: methods and protocols. Humana Press. p. 45–59. [Google Scholar]
- Dyer SC et al. 2025. Ensembl 2025. Nucleic Acids Res. 53:D948–D957. 10.1093/nar/gkae1071. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Earth BioGenome Project Subcommittee for Annotation . 2026. Report on Annotation: Recommended Tools v.4.0 [accessed 06.23.2026]. https://www.earthbiogenome.org/report-on-annotation-recommended-tools.
- Earth BioGenome Project Subcommittee for Sequencing and Assembly . 2026a. Report on Assembly Standards v.7.0 [accessed 06.23.2026]. https://www.earthbiogenome.org/report-on-assembly-standards.
- Earth BioGenome Project Subcommittee for Sequencing and Assembly . 2026b. Report on Assembly Recommendations v.4 [accessed 06.23.2026]. https://www.earthbiogenome.org/report-on-assembly-recommendations.
- Frandsen PB et al. 2023. Allelic resolution of insect and spider silk genes reveals hidden genetic diversity. Proc Natl Acad Sci U S A. 120:e2221528120. 10.1073/pnas.2221528120. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Giraldo-Calderón GI et al. 2022. VectorBase.Org updates: bioinformatic resources for invertebrate vectors of human pathogens and related organisms. Curr Opin Insect Sci. 50:100860. 10.1016/j.cois.2021.11.008. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Goldfarb T et al. 2025. NCBI RefSeq: reference sequence standards through 25 years of curation and annotation. Nucleic Acids Res. 53:D243–D257. 10.1093/nar/gkae1038. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Goodman SN, Fanelli D, Ioannidis JPA. 2016. What does research reproducibility mean? Sci Transl Med. 8:341ps312. 10.1126/scitranslmed.aaf5027. [DOI] [PubMed] [Google Scholar]
- Goubert C et al. 2022. A beginner's guide to manual curation of transposable elements. Mob DNA. 13:7. 10.1186/s13100-021-00259-7. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Guiglielmoni N. 2025. De novo genome assembly using long reads and chromosome conformation capture. In: Bonizzoni M, Ometto L, editors. Insect genomics: methods and protocols. Springer US. p. 1–27. [DOI] [PubMed] [Google Scholar]
- Halfon MS et al. 2026. How to write (and review) a genome report. In review.
- Heckenhauer J, Razuri-Gonzales E, Mwangi FN, Schneider J, Pauls SU. 2023. Holotype sequencing of Silvataresholzenthali Rázuri-Gonzales, Ngera & Pauls, 2022 (Trichoptera, Pisuliidae). Zookeys. 1159:1–15. 10.3897/zookeys.1159.98439. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Howard C et al. 2025. On the path to reference genomes for all biodiversity: laboratory protocols and lessons learned from processing over 2000 species in the Sanger Tree of Life. Gigascience. 14:giaf119. 10.1093/gigascience/giaf119. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Howe K et al. 2021. Significantly improving the quality of genome assemblies through curation. GigaScience. 10:giaa153. 10.1093/gigascience/giaa153. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Goble C et al. 2021. Implementing FAIR digital objects in the EOSC-Life workflow collaboratory. Zenodo. 10.5281/zenodo.4605654. [DOI]
- Huang N, Li H. 2023. Compleasm: a faster and more accurate reimplementation of BUSCO. Bioinformatics. 39:btad595. 10.1093/bioinformatics/btad595. [DOI] [PMC free article] [PubMed] [Google Scholar]
- i5K Consortium . 2013. The i5k initiative: advancing arthropod genomics for knowledge, human health, agriculture, and the environment. J Hered. 104:595–600. 10.1093/jhered/est050. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Johnston JS, Bernardini A, Hjelmen CE. 2019. Genome size estimation and quantitative cytogenetics in insects. In: Brown SJ, Pfrender ME, editors. Insect genomics: methods and protocols. Springer New York. p. 15–26. [DOI] [PubMed] [Google Scholar]
- Jurka J et al. 2005. Repbase update, a database of eukaryotic repetitive elements. Cytogenet Genome Res. 110:462–467. 10.1159/000084979. [DOI] [PubMed] [Google Scholar]
- Karsch-Mizrachi I et al. 2025. The international nucleotide sequence database collaboration (INSDC): enhancing global participation. Nucleic Acids Res. 53:D62–D66. 10.1093/nar/gkae1058. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Keränen SVE, Villahoz-Baleta A, Bruno AE, Halfon MS. 2022. Redfly: an integrated knowledgebase for insect regulatory genomics. Insects. 13:618. 10.3390/insects13070618. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Kim J et al. 2025. Finishing a complete giraffe genome from telomere to telomere with Verkko-Fillet [preprint]. bioRxiv 679366. 10.1101/2025.10.01.679366. [DOI] [PMC free article] [PubMed]
- Kohany O, Gentles AJ, Hankus L, Jurka J. 2006. Annotation, submission and screening of repetitive elements in Repbase: RepbaseSubmitter and Censor. BMC Bioinformatics. 7:474. 10.1186/1471-2105-7-474. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Larivière D et al. 2024. Scalable, accessible and reproducible reference genome assembly and evaluation in Galaxy. Nat Biotechnol. 42:367–370. 10.1038/s41587-023-02100-3. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Laumer C. 2023. Picogram input multimodal sequencing (pimms) v.1 [accessed 06.23.2026]. 10.17504/protocols.io.rm7vzywy5lx1/v1. [DOI]
- Lawniczak MK et al. 2023. Recording sample metadata for the Darwin Tree of Life Project (v2.5 12.01.2023) [accessed 03-21-2025]. https://github.com/darwintreeoflife/metadata/blob/main/DToL%20METADATA%20SOP%20v.2.5.pdf.
- Lawniczak MK et al. 2022. Standards recommendations for the Earth BioGenome Project. Proc Natl Acad Sci U S A. 119:e2115639118. 10.1073/pnas.2115639118. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Lawniczak MKN et al. 2025. Best-practice guidance for Earth BioGenome Project sample collection and processing: progress and challenges in biodiverse reference genome creation. Gigascience. 14:giaf041. 10.1093/gigascience/giaf041. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Lawson D et al. 2009. VectorBase: a data resource for invertebrate vector genomics. Nucleic Acids Res. 37:D583–D587. 10.1093/nar/gkn857. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Lécher P, Defaye D, Noel P. 1995. Chromosomes and nuclear DNA of crustacea. Invertebr Reprod Dev. 27:85–114. 10.1080/07924259.1995.9672440. [DOI] [Google Scholar]
- Lewin HA et al. 2018. Earth BioGenome Project: sequencing life for the future of life. Proc Natl Acad Sci U S A. 115:4325–4333. 10.1073/pnas.1720115115. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Lewin HA et al. 2022. The Earth BioGenome Project 2020: starting the clock. Proc Natl Acad Sci U S A. 119:e2115635118. 10.1073/pnas.2115635118. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Li H et al. 2022. Fly cell atlas: a single-nucleus transcriptomic atlas of the adult fruit fly. Science. 375:eabk2432. 10.1126/science.abk2432. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Li H, Durbin R. 2024. Genome assembly in the telomere-to-telomere era. Nat Rev Genet. 25:658–670. 10.1038/s41576-024-00718-w. [DOI] [PubMed] [Google Scholar]
- Lou RN, Jacobs A, Wilder AP, Therkildsen NO. 2021. A beginner's guide to low-coverage whole genome sequencing for population genomics. Mol Ecol. 30:5966–5993. 10.1111/mec.16077. [DOI] [PubMed] [Google Scholar]
- Martin FJ et al. 2023. Report on annotation standards v.1.0. [accessed 06.23.2026]. https://www.earthbiogenome.org/report-on-annotation-standards.
- McCartney A et al. 2022. Data sharing and management best practices v.1.0 [accessed 06.23.2026]. https://www.earthbiogenome.org/data-sharing-management-best-practices.
- McCartney A, Böhne A, Staunton C; ERGA ELSI Committee, ERGA Council . 2021. European Reference Genome Atlas sample code of practice (v1 27.09.2021) [accessed 03.21.2025]. https://github.com/ERGA-consortium/ERGA-sample-manifest/blob/main/ERGA_SamplingCode_BestPractice.pdf.
- Mei Y et al. 2022. InsectBase 2.0: a comprehensive gene resource for insects. Nucleic Acids Res. 50:D1040–D1045. 10.1093/nar/gkab1090. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Mölder F et al. 2021. Sustainable data analysis with snakemake. F1000Res. 10:33. 10.12688/f1000research.29032.2. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Mueller O, Lightfoot S, Schroeder A. 2016. RNA Integrity Number (RIN)–standardization of RNA quality control. Application Note - Agilent Technologies [accessed 06.23.2026]. https://www.agilent.com/cs/library/applications/5989-1165EN.pdf?srsltid=AfmBOorRhlqYnPVh_-S7TiKhctMI7Smu6oBhKt9CQzgOEWlji70zrBcK.
- Nevers Y et al. 2024. Quality assessment of gene repertoire annotations with OMArk. Nat Biotechnol. 43:124–133. 10.1038/s41587-024-02147-w [DOI] [PMC free article] [PubMed] [Google Scholar]
- Nurk S et al. 2022. The complete sequence of a human genome. Science. 376:44–53. 10.1126/science.abj6987. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Nuss AB et al. 2023. The highly improved genome of Ixodes scapularis with X and Y pseudochromosomes. Life Sci Alliance. 6:e202302109. 10.26508/lsa.202302109. [DOI] [PMC free article] [PubMed] [Google Scholar]
- O’Leary NA et al. 2024. Exploring and retrieving sequence and metadata for species across the tree of life with NCBI Datasets. Sci Data. 11:732. 10.1038/s41597-024-03571-y. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Pardo-Palacios FJ et al. 2024. Systematic assessment of long-read RNA-seq methods for transcript identification and quantification. Nat Methods. 21:1349–1363. 10.1038/s41592-024-02298-3. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Petersen M et al. 2019. Diversity and evolution of the transposable element repertoire in arthropods with particular reference to insects. BMC Ecol Evol. 19:11. 10.1186/s12862-018-1324-9. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Poelchau M et al. 2015. The i5k Workspace@NAL—enabling genomic data access, visualization and curation of arthropod genomes. Nucleic Acids Res. 43:D714–D719. 10.1093/nar/gku983. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Poelchau M et al. 2026. EBP IT and Informatics Standards v.2.3 [accessed 06.23.2026]. https://www.earthbiogenome.org/it-and-informatics-standards.
- Pruitt KD, Katz KS, Sicotte H, Maglott DR. 2000. Introducing RefSeq and Locuslink: curated human genome resources at the NCBI. Trends Genet. 16:44–47. 10.1016/S0168-9525(99)01882-X. [DOI] [PubMed] [Google Scholar]
- Ranallo-Benavidez TR, Jaron KS, Schatz MC. 2020. GenomeScope 2.0 and Smudgeplot for reference-free profiling of polyploid genomes. Nat Commun. 11:1432. 10.1038/s41467-020-14998-3. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Reichel K et al. 2026. From permits to samples: addressing key challenges for high-quality reference genome generation in Europe. Mol Ecol Resour. 26:e70100. 10.1111/1755-0998.70100. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Reiter T et al. 2021. Streamlining data-intensive biology with workflow systems. GigaScience. 10:giaa140. 10.1093/gigascience/giaa140. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Rhie A et al. 2021. Towards complete and error-free genome assemblies of all vertebrate species. Nature. 592:737–746. 10.1038/s41586-021-03451-0. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Rhie A et al. 2023. The complete sequence of a human y chromosome. Nature. 621:344–354. 10.1038/s41586-023-06457-y. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Saha S et al. 2022. Recommendations for extending the GFF3 specification for improved interoperability of genomic data [preprint]. arXiv, arXiv:220207782. 10.48550/arXiv.2202.07782. [DOI]
- Sember A et al. 2020. Patterns of sex chromosome differentiation in spiders: insights from comparative genomic hybridisation. Genes (Basel). 11:849. 10.3390/genes11080849. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Sherkow JS et al. 2022. Ethical, legal, and social issues in the Earth BioGenome Project. Proc Natl Acad Sci U S A. 119:e2115859119. 10.1073/pnas.2115859119. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Simão FA, Waterhouse RM, Ioannidis P, Kriventseva EV, Zdobnov EM. 2015. BUSCO: assessing genome assembly and annotation completeness with single-copy orthologs. Bioinformatics. 31:3210–3212. 10.1093/bioinformatics/btv351. [DOI] [PubMed] [Google Scholar]
- Smee C; Darwin Tree of Life Code of Practice Working Group . 2020. Sampling code of practice (v1 2020.04.21) [accessed 03-21-2025]. https://www.darwintreeoflife.org/wp-content/uploads/2023/10/Template-DToL-Sampling-Code-Of-Practice.pdf.
- Smith B et al. 2007. The OBO foundry: coordinated evolution of ontologies to support biomedical data integration. Nat Biotechnol. 25:1251–1255. 10.1038/nbt1346. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Sommer MJ, Zimin AV, Salzberg SL. 2025. PSAURON: a tool for assessing protein annotation across a broad range of species. NAR Genom Bioinform. 7:lqae189. 10.1093/nargab/lqae189. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Storer J, Hubley R, Rosen J, Wheeler TJ, Smit AF. 2021. The Dfam community resource of transposable element families, sequence models, and genome annotations. Mob DNA. 12:2. 10.1186/s13100-020-00230-y. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Sweeten A, Schatz MC, Phillippy AM. 2026. AniAnn's: Alignment-free annotation of tandem repeat arrays using fast average nucleotide identity estimates [preprint]. bioRxiv. 10.64898/2026.01.27.702063. [DOI] [PMC free article] [PubMed]
- Sweeten AP, Schatz MC, Phillippy AM. 2024. ModDotPlot—rapid and interactive visualization of tandem repeats. Bioinformatics. 40:btae493. 10.1093/bioinformatics/btae493. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Teltscher F, Lawniczak M. 2023. Squishing insects for preservation of HMW DNA in the field [accessed 03-21-2025]. https://www.protocols.io/view/squishing-insects-for-preservation-of-hmw-dna-in-t-4r3l2224jl1y/v1.
- Teytelman L, Stoliartchouk A, Kindler L, Hurwitz BL. 2016. Protocols.Io: virtual communities for protocol development and discussion. PLoS Biol. 14:e1002538. 10.1371/journal.pbio.1002538. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Thomas GWC, McKibben MTW, Hahn MW, Barker MS. 2024. A comprehensive examination of chelicerate genomes reveals no evidence for a whole genome duplication among spiders and scorpions [preprint]. bioRxiv 578966. 10.1101/2024.02.05.578966. [DOI]
- Threlfall J, Blaxter M. 2021. Launching the tree of life gateway. Wellcome Open Res. 6:125. 10.12688/wellcomeopenres.16913.1. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Thurmond J et al. 2018. Flybase 2.0: the next generation. Nucleic Acids Res. 47:D759–D765. 10.1093/nar/gky1003. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Tidwell JP et al. 2024. Identifying the sex chromosome and sex determination genes in the cattle tick, Rhipicephalus (Boophilus) microplus. G3 (Bethesda). 14:jkae234. 10.1093/g3journal/jkae234. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Uliano-Silva M et al. 2023. Mitohifii: a python pipeline for mitochondrial genome assembly from pacbio high fidelity reads. BMC Bioinformatics. 24:288. 10.1186/s12859-023-05385-y. [DOI] [PMC free article] [PubMed] [Google Scholar]
- UniProt Consortium . 2018. UniProt: the universal protein knowledgebase. Nucleic Acids Res. 46:2699–2699. 10.1093/nar/gky092. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Vancaester E, Blaxter M. 2023. Phylogenomic analysis of Wolbachia genomes from the Darwin Tree of Life biodiversity genomics project. PLoS Biol. 21:e3001972. 10.1371/journal.pbio.3001972. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Vancaester E, Blaxter ML. 2024. Markerscan: separation and assembly of cobionts sequenced alongside target species in biodiversity genomics projects. Wellcome Open Res. 9:33. 10.12688/wellcomeopenres.20730.1. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Vizueta J, Rozas J, Sánchez-Gracia A. 2018. Comparative genomics reveals thousands of novel chemosensory genes and massive changes in chemoreceptor repertories across chelicerates. Genome Biol Evol. 10:1221–1236. 10.1093/gbe/evy081. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Vizueta J, Sánchez-Gracia A, Rozas J. 2020. Bitacora: a comprehensive tool for the identification and annotation of gene families in genome assemblies. Mol Ecol Resour. 20:1445–1452. 10.1111/1755-0998.13202. [DOI] [PubMed] [Google Scholar]
- von Reumont BM et al. 2022. Modern venomics—current insights, novel methods, and future perspectives in biological and applied animal venom research. GigaScience. 11:giac048. 10.1093/gigascience/giac048. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Wang P, Wang F. 2023. A proposed metric set for evaluation of genome assembly quality. Trends Genet. 39:175–186. 10.1016/j.tig.2022.10.005. [DOI] [PubMed] [Google Scholar]
- Wetterstrand KA. 2025. DNA sequencing costs: Data [accessed 06.23.2026]. https://www.genome.gov/about-genomics/fact-sheets/DNA-Sequencing-Costs-Data.
- Wilkinson MD et al. 2016. The FAIR guiding principles for scientific data management and stewardship. Sci Data. 3:160018. 10.1038/sdata.2016.18. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Wilkinson SR et al. 2025. Applying the FAIR principles to computational workflows. Sci Data. 12:328. 10.1038/s41597-025-04451-9. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Wright CJ et al. 2025. Project psyche: reference genomes for all Lepidoptera in Europe. Trends Ecol Evol. 40:1234–1250. 10.1016/j.tree.2025.10.007. [DOI] [PubMed] [Google Scholar]
- Wu C, Lu J. 2019. Diversification of transposable elements in arthropods and its impact on genome evolution. Genes (Basel). 10:338. 10.3390/genes10050338. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Xing B, Yang L, Gulinuer A, Ye G. 2023. Research progress on horizontal gene transfer and its functions in insects. Tropical Plants. 2:1–12. 10.48130/TP-2023-0003. [DOI] [Google Scholar]
- Yoo D et al. 2025. Complete sequencing of ape genomes. Nature. 641:401–418. 10.1038/s41586-025-08816-3. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Zhou C et al. 2025. Oatk: a de novo assembly tool for complex plant organelle genomes. Genome Biol. 26:235. 10.1186/s13059-025-03676-6. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Ziemann M, Poulain P, Bora A. 2023. The five pillars of computational reproducibility: bioinformatics and beyond. Brief Bioinform. 24:bbad37. 10.1093/bib/bbad375. [DOI] [PMC free article] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Data Citations
- Goble C et al. 2021. Implementing FAIR digital objects in the EOSC-Life workflow collaboratory. Zenodo. 10.5281/zenodo.4605654. [DOI]
Supplementary Materials
Data Availability Statement
The authors affirm that all data necessary to confirm the article’s conclusions are included in the article, the figures, and the supplementary files.
Supplemental material available at GENETICS online.
