Abstract
DNA barcoding is a widely used tool for species identification, with its reliability heavily dependent on reference databases. While the quality of these databases has long been debated, a critical knowledge gap remains in their comprehensive evaluation and comparison at regional scales. Marine metazoan species in the western and central Pacific Ocean (WCPO), a region characterized by high biodiversity and limited sequencing efforts, are an example of this gap. This study developed a systematic workflow to assess mitochondrial cytochrome c oxidase subunit I (COI) barcode coverage and sequence quality in two commonly used reference databases for DNA barcoding: the nucleotide reference database from the National Center for Biotechnology Information (NCBI); and from the Barcode of Life Data System (BOLD). Comparative analyses across marine phyla and WCPO regions identified significant barcode gaps and quality problems, providing insights to guide future barcoding efforts. NCBI exhibited higher barcode coverage, but lower sequence quality compared to BOLD. Quality issues, including over- or under-represented species, short sequences, ambiguous nucleotides, incomplete taxonomic information, conflict records, high intraspecific distances, and low inter-specific distances were identified in both databases, likely resulting from contamination, cryptic species, sequencing errors, or inconsistent taxonomic assignment. The barcode identification number (BIN) system in BOLD demonstrated potential for identifying and addressing problematic records, highlighting the benefits of curated databases. Significant barcode deficiencies and quality issues were observed in the south temperate region of WCPO and phyla such as Porifera, Bryozoa, and Platyhelminthes. Additionally, the COI barcode showed limited species-level resolution for certain taxa, including Scombridae and Lutjanidae. Addressing barcode coverage gaps, improving taxonomic representation, and enhancing sequence quality will be essential for strengthening future barcoding initiatives and advancing biodiversity monitoring and conservation in the WCPO and beyond. This study highlights the need for standardized database curation and sequencing practices to improve the global reliability and applicability of DNA barcoding.
Keywords: Reference database, DNA barcoding, COI, Marine macrofauna, Western and Central Pacific Ocean
Introduction
Marine ecosystems are rich in species diversity and support multiple ecological and socio-economic services (Hills et al., 2011; Hoegh-Guldberg & Bruno, 2010; Pauly et al., 2002). Recent advancements in DNA barcoding, the use of a standardized short DNA fragment to identify species, have greatly enhanced biodiversity assessments by improving the efficiency and accuracy of taxonomic identification (Bourlat et al., 2013; Bucklin, Steinke & Blanco-Bercial, 2011; Hebert et al., 2003). DNA barcoding, when combined with high-throughput sequencing technologies, enables the efficient investigation of entire biological communities, an approach referred to as DNA metabarcoding (Bourlat et al., 2013; Bucklin, Steinke & Blanco-Bercial, 2011; Stoeckle, Das Mishu & Charlop-Powers, 2020). Both DNA barcoding and metabarcoding are widely used in ecological and biodiversity investigations, including species identification (Baetscher et al., 2023; Bucklin, Steinke & Blanco-Bercial, 2011), dietary analysis (Alberdi et al., 2019; Günther et al., 2021), trophic interaction assessment (Clare, 2014; Leray et al., 2012), and environmental biomonitoring (Bista et al., 2017; Trujillo-González, 2022).
DNA barcoding and metabarcoding require reliable reference databases to ensure the accurate assignment of DNA sequences to specific taxa (Leray et al., 2019; Marques et al., 2021; Ramirez et al., 2020). Global databases like the National Center for Biotechnology Information (NCBI) (Benson, Lipman & Ostell, 1993), the European Nucleotide Archive (ENA) (Kanz, 2004), and the DNA Data Bank of Japan (DDBJ) (Tateno & Gojobori, 1997) are widely used in DNA barcoding studies due to their extensive collections of publicly available sequence records contributed by researchers worldwide. However, the accuracy and reliability of these records have been the subject of debate (Ardura, 2019; Chen, Zobel & Verspoor, 2017; Leray et al., 2019; Marques et al., 2021; Shen, Chen & Murphy, 2013; Turanov & Kartavtsev, 2021). Criticisms of global databases primarily center on whether their curation and validation systems for user-submitted sequences and associated metadata are as robust as specialized curated databases (Ardura, 2019; Chorlton, 2024; Gonçalves & Musen, 2018). For example, some studies have shown that curated reference databases provide more reliable taxonomic identification for DNA barcoding analysis compared to global databases (Gold et al., 2021; Hou et al., 2018; Stoeckle, Das Mishu & Charlop-Powers, 2020). Additionally, redundancies and inconsistencies in global-scale databases can reduce alignment performance, introduce errors, and produce ambiguous results (Chen, Zobel & Verspoor, 2017; Chorlton, 2024).
In comparison, curated databases like the Barcoding of Life Data System (BOLD), which focuses on the mitochondrial cytochrome c oxidase subunit I (COI) barcode (Ratnasingham & Hebert, 2007), and Greengenes (DeSantis et al., 2006) and SILVA (Quast et al., 2012), which focus on ribosomal RNA genes, are generally considered to be more reliable due to their strict quality control protocols and standardized metadata system (Puillandre et al., 2012; Steinke & Hanner, 2011). For instance, the barcode index number (BIN) system (Ratnasingham & Hebert, 2013) is a feature of BOLD that automatically clusters sequences into operational taxonomic units (OTUs) based on genetic similarity, typically corresponding to species-level groupings. This system facilitates species delimitation, highlights potential cases of cryptic diversity, and assists in identifying problematic records, thereby enhancing the reliability of sequence and taxonomy data (Costa et al., 2012; Fontes et al., 2021; Oliveira et al., 2016). However, the lack of barcode records remains a significant challenge in curated databases, potentially reducing taxonomic resolution and increasing the likelihood of misidentification or failed taxonomic assignments (Ardura, 2019; Puillandre et al., 2012; Ramirez et al., 2020). BOLD, in particular, has been reported to exhibit lower public barcode coverage compared to NCBI, partly due to its stricter metadata requirements, voucher specimen standards, and sequence curation protocols, which can limit the immediate availability of sequence submissions (Heller et al., 2018; Hestetun et al., 2020).
The evaluation of databases is normally conducted on BOLD due to its BIN system which facilitates the assessment of record quality (Costa et al., 2012; Fontes et al., 2021; Leite et al., 2020; Ratnasingham & Hebert, 2013). In contrast, the absence of a quality evaluation system in NCBI often limits comparisons between NCBI and BOLD to barcode coverage alone (Mugnai et al., 2021; Ramirez et al., 2020; Stoeckle, Das Mishu & Charlop-Powers, 2020; Weigand et al., 2019). This creates a knowledge gap in the comprehensive evaluation of sequence quantity and quality of both databases. In this study, we developed an approach to simultaneously evaluate the quantity and quality of databases based on key features of DNA barcodes and the barcoding gap concept (Bucklin, Steinke & Blanco-Bercial, 2011; Hou et al., 2018; Shen, Chen & Murphy, 2013).
Here, we constructed an evaluation system to evaluate COI barcode records of marine metazoan species using data from the NCBI and BOLD databases. To do so, we focused on the western and central Pacific Ocean (WCPO) region, one of the world’s most heavily exploited ecosystems with the highest biodiversity worldwide (Costello et al., 2010; Hills et al., 2011; Nicol et al., 2013). DNA barcoding presents an excellent opportunity to improve biodiversity monitoring and assessment in the WCPO (Günther et al., 2021; Trujillo-González, 2022; Yeh et al., 2020). However, there is currently no systematic evaluation of the reliability of DNA barcode reference databases for marine species in the WCPO (Gold et al., 2021; Macheriotou et al., 2019). Of most interest is the mitochondrial COI region because its efficiency and ability to amplify a wide range of metazoans makes it a widely used target for many DNA metabarcoding studies (Leray et al., 2013; Lobo et al., 2013). Our objectives were to: (1) test the feasibility of the evaluation workflow for identifying issues in barcode records; (2) compare the reliability of NCBI and BOLD; and (3) assess barcode coverage and quality across phyla and regions within the WCPO.
Methods
All data retrieval and analyses were conducted in R (version 4.1.2) (R Core Team, 2019) using the RStudio environment (version 2022.2.0.443) (RStudio Team, 2022). Data manipulation utilized base R functions and the dplyr package (version 1.1.4) (Hadley et al., 2023), while data visualizations were performed using the ggplot2 package (version 3.5.1) (Wickham, 2016).
Species records retrieval
Species occurrence records in nine marine phylum (Annelida, Bryozoa, Chordata, Cnidaria, Arthropoda (crustacean subphylum only), Echinodermata, Mollusca, Platyhelminthes, and Porifera) from the WCPO were retrieved from the Ocean Biodiversity Information System (OBIS, https://obis.org) in October 2024 using the robis package (version 2.11.3) (Provoost & Bosch, 2022). Given the ongoing concerns regarding barcode deficiencies in the Southern Hemisphere and tropical regions (Marques et al., 2021; Ramirez et al., 2020), separate species checklists were generated for the tropical (23.5°N to 23.5°S, 140°E to 150°W), north temperate (23.5°N to 50°N, 140°E to 150°W), and south temperate (23.5°S to 50°S, 140°E to 150°W) regions of the WCPO. This regional separation was intended to better detect underrepresented areas and support region-specific recommendations for enhancing barcode coverage.
The checklists were initially filtered using metadata embedded in OBIS records, retaining only records with valid species names, accepted taxonomic status, and marine habitat classifications. Species names were further refined by removing records containing numbers or ambiguous strings, such as “sp.”, “complex.”, “cf.”. Duplicates were removed, and records were grouped by species names. Finally, the validity of species name and their synonyms was verified again using the World Register of Marine Species (WoRMS) through the worrms package (version 0.4.3) (Chamberlain & Vanhoorne, 2023).
Based on overlapping and distinct distributions among regions, species were categorized into seven distribution groups: Tropical only, North temperate only, South temperate only, All regions, North-tropical, South-tropical, and North-south. These categories were designed to facilitate comparisons of barcode coverage across WCPO regions and are not intended to reflect ecological endemism.
Barcode records retrieval
COI barcode records were retrieved from NCBI in October 2024 using the rentrez package (version 1.2.3) (Winter, 2017). Valid species names and synonyms from the WPCO species checklist were used as taxonomic group keywords, while keywords for the COI gene included “COI”, “CO1”, “cytochrome c oxidase subunit I”, “Cox1”, and “COXI”. Taxonomic information for each sequence record was retrieved using accession IDs.
Complete barcode record data, including sequences and specimen information for WPCO species in BOLD, were retrieved in October 2024 via the BOLD Web Services for Public Data Portal (https://v3.boldsystems.org/index.php/resources/api?type=webservices) using phylum names as taxa keyword. Sequences with species names matching the valid species names and/or synonyms from the WPCO species checklist, labeled as “COI-5P” and “COI-3P”, and with a BIN were retained.
Barcode databases evaluation workflow
A standardized process for barcode database evaluation and visualization was developed, comprising six basic components and three barcode taxonomic validations: (1) Barcode coverage assessment (proportion of species in the checklist with/ without barcode records). (2) Calculation of sequence counts per phylum. (3) Sequence length distribution analysis. (4) Detection of ambiguous nucleotide characters (e.g., N, R, Y, S, W, K, M, B, D, H, V), categorized by their positions (only in the at 5′ or 3′ ends, only in the middle region, and in both ends and middle). (5) Taxonomic information completeness analysis. (6) Species representation analysis, categorizing sequences into over-represented (species with > 100 barcode records), under-represented (species with < 3 barcode records), and normally represented (species with 3–100 barcode records) (Costa et al., 2012; Fontes et al., 2021; Leite et al., 2020).
The barcode taxonomic validations include: (1) Identification of conflicting records (sequences assigned to multiple taxonomic groups). (2) Barcoding gap analysis, comparing interspecific and intraspecific distances (Hou et al., 2018; Shen, Chen & Murphy, 2013; Zhang et al., 2017; Zhang & Zhang, 2014). (3) Sequence clustering based on genetic distances to identify problematic sequences and assess barcode resolution.
The script for the barcode evaluation workflow is available at GitHub (https://github.com/xyzzzeno/reference_evaluation).
Evaluation on NCBI and BOLD
Raw COI barcode records for WPCO species from NCBI and BOLD underwent the evaluations after retrieval, without prior cleanup or curation. Barcode coverage for species in the tropical, north temperate, and south temperate WCPO regions was analyzed separately. Each species was classified into four categories based on its barcode coverage: no records in either database (no barcode), records in both BOLD and NCBI (both databases), records only in BOLD (only BOLD), and records only in GenBank (only NCBI). The proportions of barcoded and non-barcoded species were also calculated across distribution categories (species distinct to a region or shared among regions). Two-proportions Z-tests were conducted to compare NCBI and BOLD, and two-way analysis of variance (two-way ANOVA) tests (Sthle & Wold, 1989) were conducted to compare WCPO regions and marine phyla.
The remaining five basic evaluations were performed on NCBI and BOLD records for all WPCO species without dividing them into climate regions or distribution categories. Two-proportions Z-tests were conducted to compare evaluation results between NCBI and BOLD.
Conflict records were initially identified by grouping identical sequences and flagging those sequences assigned to multiple taxonomy information. Barcoding gap analysis was conducted to further assess barcode taxonomic accuracy. Due to the computational intensity of this analysis, subsets of 100 species (at least three species per phylum) were randomly sampled from each database. Subsets were aligned using MAFFT (Katoh, 2002), and neighbor-joining (NJ) trees were constructed using the Jukes-Cantor (JC) model (Jukes & Cantor, 1969) in Geneious software (version 2023.2.1) (https://www.geneious.com). Pairwise distances between sequences were exported from Geneious, then analyzed and visualized in RStudio (version 2022.2.0.443) (RStudio Team, 2022). Sequences with high intra-specific distances (>0.2) or low inter-specific distances (<0.02) were flagged as problematic and re-examined to determine potential causes.
Lastly, sequences from two fish families of interests, Lutjanidae (snappers), and Scombridae (tunas), were evaluated for clustering. These families were selected due to their significant economic and ecological roles in the WCPO (Hare et al., 2020; Newman et al., 2016; NOAA Fisheries, 2018), and both have been reported to present challenges in species discrimination using COI barcodes (Fadli et al., 2024; Hou et al., 2018; Victor, Valdez-Moreno & Vásquez-Yeomans, 2015). Sequences were aligned using MAFFT (Katoh, 2002), NJ trees were constructed with the Jukes-Cantor (JC) model (Jukes & Cantor, 1969), and pairwise distances were calculated in Geneious (version 2023.2.1). Distance data were visualized in low-dimensional space using t-Distributed Stochastic Neighbor Embedding (t-SNE) with the Rtsne package (version 0.17) (Krijthe, van der Maaten & Krijthe, 2018).
Results
Species checklist and barcode coverage
A total of 46,620 species records for the tropical region, 10,717 records for the north temperate region, and 40,892 records for the south temperate region, spanning nine marine phyla, were retrieved from OBIS for the WCPO. After data cleaning and dereplication, 26,682 records remained for the tropical region, 6,280 for the north temperate region, and 24,290 for the south temperate region, resulting in a combined total of 42,419 unique marine species records (Fig. 1A). The tropical region had the highest number of regionally distinct species (14,366 species, 33.9%), followed by the south temperate region (12,791 species, 30.2%). The north temperate region had the fewest distinct species (2,593 distinct species, 6.1%) (Fig. 1B).
Figure 1. Barcode coverage across nine marine phyla and geographic regions in the western and central Pacific Ocean (WCPO).
(A) Barcode coverage and species counts for nine marine phyla across north temperate, south temperate, and tropical regions of the WCPO. Red: species with no barcode records, green: species present in both BOLD and NCBI, yellow: species present only in BOLD, and blue: species present only in NCBI. The percentages indicate the proportion of species with barcode relative to the total number of species per phylum. (B) Proportion and barcode coverage for species categorized by distinct and shared distribution patterns among the north temperate, south temperate, and tropical regions of the WCPO. The inner donut plot illustrates the proportions of species in each distribution category, the outer donut plot illustrates the proportions of barcoded and non-barcoded species within each distribution category.
NCBI exhibited higher barcode coverage (31.5%, 13,346 species) than BOLD (29.8%, 12,621 species) (Table 1). Although the difference was statistically significant, the absolute difference in coverage was relatively small (1.7%). Among barcoded species, 78.2% had records in both databases, 13.4% had records only in NCBI, and 8.4% had records only in BOLD. Barcode coverage varied significantly across the north temperate, south temperate, and tropical regions (two-way ANOVA: F(2,103) = 7.3e+29, p < 0.001,) (Fig. 1A). The north temperate region had the highest barcode coverage (61.1%), followed by the tropical (41.0%) and south temperate (35.5%) regions. Marine phyla also exhibited significant differences in barcode coverage (two-way ANOVA: F(8,97) = 1.6e+30, p < 0.001) (Fig. 1A). Chordata had the highest barcode coverage (64.1%), followed by Echinodermata (41.4%), and Cnidaria (30.7%). In contrast, Mollusca (27.8%), Annelida (26.5%), and Arthropoda (26.0%) had similar barcode coverage levels, while Bryozoa (17.9%), Platyhelminthes (14.6%), and Porifera (14.4%) showed the lowest barcode coverage. Species occurring across all three regions had the highest barcode coverage (83.5%). In contrast, species distinct to a single region had the lowest barcode coverages, and species distinct to the south temperate region exhibited the lowest barcode coverage (17.9%) (Fig. 1B). Detailed species checklist and barcode coverage is available in File S1.
Table 1. Barcode coverage and quality comparisons between NCBI and BOLD.
| NCBI | BOLD | Two-proportions Z-test comparison between NCBI and BOLD | ||||
|---|---|---|---|---|---|---|
| df | z-score | χ 2 | p-value | |||
| Barcode coverage | 31.5% | 29.8% | 1 | 5.3 | 29.0 | <0.01 |
| Over-represented species | 3.7% | 2.6% | 1 | 22.7 | 517.4 | <0.01 |
| Under-represented species | 40.7% | 38.8% | 1 | 14.2 | 202.0 | <0.01 |
| Long sequences (>1,000 bp) | 2.8% | 8.7% | 1 | 97.4 | 9,480.3 | <0.01 |
| Short sequences (<200 bp) | 0.2% | 0% | 1 | 21.1 | 444.4 | <0.01 |
| Sequences with ambiguous nucleotides | 3.1% | 3.9% | 1 | 16.1 | 258.8 | <0.01 |
| Records missing taxonomic information | 7.0% | 3.4% | 1 | 57.8 | 3,344 | <0.01 |
| Conflicting records | 0.27% | 0.31% | 1 | 3.0 | 9.14 | 0.02 |
| High intra-specific distances (>0.2) | 36.7% | 6.4% | 1 | 30.0 | 902.8 | <0.01 |
| Low inter-specific distances (<0.02) | 2.9% | 0% | 1 | 9.5 | 90.5 | <0.01 |
Barcode quality evaluation
A total of 321,997 sequences were retrieved from NCBI and 229,943 sequences from BOLD for WCPO species. NCBI contained more barcode sequences than BOLD for all marine phyla except Echinodermata (Figs. 2A, 2B).
Figure 2. Evaluation on NCBI (left) and BOLD (right) barcode reference databases for WCPO species.
(A–B) Number of sequences per phylum; (C–D) Proportion of species in each representation category; (E–F) Distribution of sequence length; (G–H) Proportion of sequences with ambiguous nucleotide characters; (I–J) Proportion of sequence records without taxonomic information at each taxonomic rank.
NCBI had a significantly higher proportion of over-represented and under-represented species (Fig. 2C), short sequences (Fig. 2E), and records with missing taxonomic information (Fig. 2I) compared to BOLD (Figs. 2D, 2F, 2J; Table 1). In contrast, BOLD had a higher proportion of long sequences (Fig. 2F) and sequences with ambiguous nucleotides (Fig. 2H) compared to NCBI (Figs. 2E, 2G; Table 1). However, NCBI had more sequences with ambiguous nucleotides located in the middle of the sequence (9,174) than BOLD (5,926). Notably, BOLD had no sequences shorter than 200 bp.
Among the phyla, Bryozoa had the highest proportion of under-represented species in both databases (67.9% in NCBI and 77.9% in BOLD), while Chordata had the lowest (26.3% in NCBI and 30.4% in BOLD) (Figs. 2C, 2D). Chordata had the highest number of short or long sequences (4,097 sequences) in NCBI (Fig. 2E) while Mollusca accounted for the largest number of long sequences (10,841 sequences) in BOLD (Fig. 2F). Porifera exhibited the highest proportion of sequences with ambiguous characters (11.7% in NCBI and 10.5% in BOLD), followed by Bryozoa (7.4% in NCBI and 6.9% in BOLD) (Figs. 2G, 2H). Mollusca had the highest proportion of records with missing order names in both databases (12.7% in NCBI and 11.5% in BOLD) (Figs. 2I, 2J).
Barcode taxonomic validation
The proportions of conflicting records were comparable between NCBI and BOLD (Fig. 3; Table 1). Both of NCBI and BOLD exhibited the highest number of conflicts at the species level. Upon further examination, conflicts caused by invalid names or synonyms accounted for 16.0% in NCBI and 22.5% in BOLD. Only five of these records (less than 1%) were submitted before 2010, suggesting that record age is unlikely to be a major contributor to the observed quality issues in this dataset. Among phyla, Chordata had the highest proportion of conflict records in both databases (29.4% in NCBI and 37.2% in BOLD). Platyhelminthes had the fewest conflict records, with only one instance in NCBI and none in BOLD. Detailed information about conflict records is provided in the File S2.
Figure 3. Number of sequence records with conflict taxonomic information, colored by phylum.
The y-axis indicates number of conflict records, the x-axis indicates taxonomic rank levels. Top, number of conflict records in NCBI; Bottom, number of conflict records in BOLD.
A total of 4,125 sequences from 100 species in NCBI and 3,119 sequences from the same 100 species in BOLD were selected for barcoding gap analysis. After calculating genetic distances, 19 outliers in NCBI and 20 outliers in BOLD exhibited extremely high genetic distances (≥ 2,000) with other sequences. After alignment and visual examination, these sequences were out of the typical barcode region, thus the outliers were excluded from further analysis.
NCBI sequences exhibited an average intra-specific distance of 0.10 ± 0.29 and an inter-specific distance of 0.46 ± 0.41 (Fig. 4A). BOLD sequences showed significantly lower average intra-specific distance (0.01 ± 0.04) and similar inter-specific distance (0.43 ± 0.11) (Fig. 4B). NCBI had significantly higher proportion of sequences with high intra-specific distances (i.e., sequences with family-level or higher taxonomic differences were labeled as the same species) and low inter-specific distances (i.e., sequences with lower than species-level differences were in different taxa) than BOLD (Table 1). Notably, BOLD had no case of low inter-specific distances.
Figure 4. Distribution of inter-specific distances and intra-specific distances in 100 species subsets of NCBI (left panel) and BOLD (right panel).
The y-axis indicates pairwise genetic distance between sequences, x-axis indicates frequency of distances, colored by inter-specific distance (red) and intra-specific distance (blue). (A–B) overall distance distribution in NCBI and BOLD; (C–D) phylum-focused distance distribution.
Among phyla, Bryozoa exhibited the highest proportion of sequences with large intra-specific distances (95.7% in NCBI and 96.0% in BOLD), followed by Cnidaria (56.8% in NCBI and 4% in BOLD) (Figs. 4C, 4D). At species level, Bugula neritina (Bryozoa) was most frequently associated with high intra-specific distance problems (File S2). Majority of sequences with low inter-specific distance problems showed conflicts at the phylum level (91.4%). Upon visual examination, these problematic sequences were likely to be false annotated or contamination sequences generated in the same project. Notably, 92.5% of BOLD records with high intra-specific distances were assigned to multiple BINs, while the remaining 7.5%, all from Bugula neritina, were still grouped within a single BIN. Detailed information about high intra-specific distances and low inter-specific distances records is provided in the File S2.
A total of 1,520 sequences from Lutjanidae and 1,962 sequences from Scombridae in NCBI, as well as 712 sequences from Lutjanidae and 1,190 sequences from Scombridae in BOLD were used for clustering analysis. The mean sequence distance for Lutjanidae was 0.18 ± 0.09 in NCBI (Fig. 5A) and 0.16 ± 0.09 in BOLD (Fig. 5B), and the mean distance for Scombridae was 0.15 ± 0.13 in both databases (Figs. 5C, 5D). Clustering plots for both databases revealed a consistent scatter pattern for Lutjanidae (Figs. 5A, 5B), indicating issues with high intra-specific distances. In contrast, Scombridae exhibited mixing pattern among species, suggesting low inter-specific distances in both NCBI and BOLD (Figs. 5C, 5D).
Figure 5. T-SNE clustering of Lutjanidae (A–B) and Scombridae (C–D) sequences in NCBI (left panel) and BOLD (right panel), colored by species.
Distribution of inter-specific distances and intra-specific distances for each group is showed on top-left of each figure. Detailed legend available in File S3.
Discussion
Barcode coverage and quality in NCBI and BOLD
In this study, NCBI demonstrated higher COI barcode coverage than BOLD, consistent with previous findings (Ardura, 2019; Duarte, Vieira & Costa, 2020; Hestetun et al., 2020). This difference can be attributed to NCBI’s status as a global-scale and public-accessible database, which allows users worldwide to freely upload sequences (Gold et al., 2021; Leray et al., 2019; Turanov & Kartavtsev, 2021). However, the absolute difference in barcode coverage between NCBI and BOLD was only 1.7%, which is unlikely to substantially impact biodiversity assessments, especially given that BOLD was observed to have fewer quality issues than NCBI (Table 1, Fig. 2). A more critical concern identified in this study is that more than two-thirds of marine species in the WCPO lack reference COI barcode records in both databases. This substantial gap currently limits the reliability of DNA barcoding for biodiversity monitoring in the WCPO, suggesting that further efforts to expand and curate reference databases are critical prerequisites for effective large-scale applications.
Our evaluation revealed key differences between BOLD and NCBI that users should consider when conducting metabarcoding studies. First, a larger proportion of WCPO species were either over-represented or under-represented in NCBI compared to BOLD. Second, NCBI contained a greater number of extremely short sequences (<200 bp), which may compromise barcode alignment accuracy. Third, missing taxonomic information was more prevalent in the NCBI records. Finally, while NCBI had a lower overall proportion of sequences with ambiguous nucleotide characters, a higher percentage of these problematic sequences had ambiguities in the central region of sequences, raising concerns about sequence quality. These issues likely stem from the original purposes of the two databases: NCBI was initially constructed for biomedical research (Benson, Lipman & Ostell, 1993), whereas BOLD was designed specifically for biodiversity applications (Ratnasingham & Hebert, 2007). A key contributor to BOLD’s higher reliability is its systematic use of voucher specimens, which are often lacking in NCBI records. Voucher-linked records not only enhance taxonomic accuracy but also facilitate curation and quality control, thereby improving the utility of BOLD for biomonitoring purposes (Heller et al., 2018; Ip et al., 2019; Valdez-Moreno et al., 2019). The utility of BOLD for large-scale biodiversity biomonitoring has been demonstrated by Valdez-Moreno et al. (2019), who employed thousands of voucher-linked BOLD records for eDNA-based assessments of fish community composition in a tropical oligotrophic lake.
The barcode evaluation workflow developed in this study highlighted several critical aspects affecting database reliability. For instance, species representation analysis identified disparities in barcoding and sequencing efforts with over-represented species indicating potential redundancy and under-represented species lacking sufficient sequence variation for reliable taxonomic assignment (Ardura, 2019; Bazinet et al., 2018; Costa et al., 2012; Weigand et al., 2019). Although NCBI exhibited broader barcode coverage overall, species representation was highly uneven, with some taxa disproportionately well-represented while others remained poorly represented. Targeted efforts to improve barcode coverage for underrepresented groups are essential to enhance the accuracy and completeness of both NCBI and BOLD databases (Ardura, 2019; Bazinet et al., 2018; Ramirez et al., 2020).
Sequence length evaluation is a crucial initial step in identifying problematic sequences that fall outside the target barcode region. Short sequences may result from sequencing artifacts and lead to ambiguous alignments (Nagai et al., 2022; Preston, Fritzsche & Woodcock, 2022), while excessively long sequences could contain non-target genomic regions or pseudogenes, which can interfere with barcode alignment (Guo et al., 2022; Song et al., 2008). BOLD, with a stricter quality control process (Ratnasingham & Hebert, 2007; Ratnasingham & Hebert, 2013), contained fewer short, low-quality sequences but a higher proportion of full-length COI gene sequences (1,500–1,600 bp) (Table 1, Fig. 2F). Although longer sequences may provide additional genetic information, trimming them to the barcode region is recommended to ensure consistency in analyses (Jeunen et al., 2023; Robeson et al., 2021).
The presence of ambiguous nucleotides, which can arise from sequencing errors, primer mismatches, or contamination is another critical issue (Redelings, 2014; Wheeler, 1994). Our analysis categorized ambiguities based on their positions (Figs. 2G–2H), which may provide insights into their potential causes and help guide appropriate handling strategies. For instance, ambiguities at sequence ends could be resolved by the users through trimming, while ambiguities in conserved regions require closer examination or exclusion from analyses depending on the purpose of the users. Although NCBI contained fewer ambiguous sequences overall, the higher proportion of centrally located ambiguities suggests potential sequencing quality issues compared to BOLD.
Missing taxonomic evaluations was another significant concern, reflecting potential issues such as outdated taxonomic classifications, unresolved species names, or user submission errors (Bouchet et al., 2017; Cunha & Giribet, 2019). NCBI exhibited more records lacking taxonomic hierarchy details, highlighting the need for improving curation and standardization. Addressing these deficiencies will require extensive database curation, ideally involving taxonomic specialists to ensure accurate species identification and consistent taxonomic frameworks across records.
Barcode taxonomic accuracy in NCBI and BOLD
Two major barcode taxonomic issues were identified in NCBI and BOLD: conflicting taxonomic records and high intra-specific genetic distances. Notably, all problematic records in BOLD lacked BIN numbers, reinforcing the BIN system’s utility in identifying and resolving errors.
Conflicting records, where identical or highly similar sequences were assigned to different taxonomic groups, were likely caused by (1) human errors or contamination during sequencing or data entry, particularly in datasets using broad-range primers designed for multiple taxa. Conflicts because of this reason were often observed between morphologically and/or ecologically distinct groups (Leray et al., 2019; Mugnai et al., 2021); (2) taxonomic inconsistencies, such as synonyms and invalid or outdated names (Fontes et al., 2021; Knebelsberger et al., 2014; Oliveira et al., 2016); (3) misidentifications, especially in morphologically similar taxa. For example, within Porifera, confusions between Axinellida and Bubarida was observed (File S2), likely due to their close morphologically resemblance and recent divergence history (Galitz et al., 2021).
The second major issue—high intra-specific distances—suggests the presence of cryptic species, sequencing contamination, or misidentifications. For example, sequences associated with the bryozoan Bugula neritina exhibited extreme genetic divergence, consistent with its known cryptic diversity and recent speciation events (Mackie, Keough & Christidis, 2006). The symbiotic relationships of Bugula neritina with other species (Linneman et al., 2014) may further contribute to the amplification of non-target DNA, potentially leading to ambiguous barcode assignments (Leray et al., 2019; Xie et al., 2025). These findings underscore the importance of involving taxonomic specialists to accurately resolve species boundaries. In addition, linking sequence records to voucher specimens would greatly support record validation and database curation efforts (Fontes et al., 2021; Ratnasingham & Hebert, 2007; Valdez-Moreno et al., 2019).
Sequence clustering plots provided a straightforward approach to visualize sequence distances, clusters, and outliers. Due to computational limitations, only two families were analyzed here. A more scattered pattern was observed in Lutjanidae, likely resulting from cross-phyla contamination, as some Lutjanidae sequences were identified as problematic in the barcoding gap analysis. For Scombridae, cross-species clustering highlighted inefficiencies in barcode resolution for distinguishing among Scombridae species, consistent with previous studies (Hou et al., 2018; Victor, Valdez-Moreno & Vásquez-Yeomans, 2015). These issues in Lutjanidae and Scombridae were consistent across both databases, underscoring the need for further optimization and improvement of DNA barcoding analyses for these taxa.
Barcode problems in the WCPO regions across nine marine phyla
The south temperate WCPO exhibited the lowest barcode coverage among the three regions, with over 80% of distinct species remaining unbarcoded. Severe barcode deficiencies were also noted in the tropical WCPO, particularly among distinct species. Given the high biodiversity in these regions, such barcode deficiencies could have serious consequences, including reduced ability to identify species, discover new or cryptic species, monitor biodiversity, or construct accurate phylogenetic profiles (Bucklin, Steinke & Blanco-Bercial, 2011). Notably, many countries in the Southern Hemisphere are economically disadvantaged (Sachs, Mellinger & Gallup, 2001), which limits their access to high-quality biodiversity investigation technologies. Collaborative initiatives between high-income and lower-income countries could play a critical role in expanding regional barcoding capacity and improving the representation of biodiversity hotspots within reference databases.
Additionally, the lack of barcode records for distinct species might suggest that sequencing resources and efforts may be disproportionately allocated to common and widely distributed species. Addressing this imbalance requires a critical increase in sequencing efforts targeting distinct and rare species. These species are essential for comprehensive biodiversity analyses and may play vital roles in maintaining ecosystem stability and functionality (Burlakova et al., 2011; Lamoreux et al., 2006).
Among phyla, Porifera, Bryozoa, and Platyhelminthes exhibited both low species counts and barcode coverage, requiring urgent attention to species inventories and molecular sequencing. Barcode quality issues were also notable in Porifera and Bryozoa, including a high proportion of sequences with ambiguous nucleotide characters and a lack of distinct COI barcoding gaps. These challenges may stem from high divergence rates (Linneman et al., 2014), cryptic diversity (Bucklin, Steinke & Blanco-Bercial, 2011; Mackie, Keough & Christidis, 2006), and symbiotic associations with other organisms (Linneman et al., 2014; Mugnai et al., 2021; Vargas et al., 2012), which might reduce barcode efficiency and reliability.
Despite the relatively high barcode coverage in Chordata, quality issues persisted, such as unreliable short sequences and conflicting records. Fish species, in particular, are often studied using broad-range primers (Iwasaki et al., 2013; Leray et al., 2013; Miya et al., 2015), increasing the risk of co-amplification and contamination. Additionally, some studies suggest that COI lacks sufficient species-level resolution for certain Scombridae species (Leray et al., 2019; Victor, Valdez-Moreno & Vásquez-Yeomans, 2015; Wangensteen et al., 2018), as observed in our clustering study. A significant issue in Mollusca was missing taxonomic information, likely due to taxonomic uncertainties and validations in certain Gastropoda groups (Bouchet et al., 2017; Cunha & Giribet, 2019). The absence of taxonomic information underscores the need for further database curation, phylogenetic analyses, and the consistent use of updated taxonomic names and classifications for molluscs, ideally involving collaboration with taxonomic experts and the use of authoritative resources such as WoRMS or MolluscaBase.
Our study is among the first to systematically evaluate both the coverage and quality of COI barcode records for marine metazoan species in the WCPO. Similar issues with reference databases have been reported in other regions (Mugnai et al., 2021; Ramirez et al., 2020), suggesting that these challenges are not region-specific. Therefore, our findings have broader implications beyond the WCPO, highlighting the need for global improvements in barcode database curation, quality control, and standardization to enhance the reliability of DNA barcoding for biodiversity research and conservation.
Recommendations for database curation
To improve reference database reliability, we propose the following curation strategies for database users:
-
1.
Remove or trim sequences that fall outside the standard COI barcode length.
-
2.
Discard sequences with ambiguous nucleotides in conserved regions and trim ambiguous sites and sequence ends.
-
3.
Standardize taxonomic metadata by replacing missing taxonomic information with meaningful strings or marks (e.g., uncertain order name), and fill gaps by checking authorized species list to ensure taxonomic completeness.
-
4.
Dereplicate sequences to minimize the effects of over-represented species and redundant sequences.
-
5.
Flag and remove sequences with conflicting taxonomy, particularly those with discrepancies at higher taxonomic levels. When relevant, also consider the year of the record, as older sequences may reflect outdated taxonomic concepts or low record quality.
-
6.
Conduct barcode gap analyses and clustering assessments to identify problematic sequences and assess barcode resolution.
Conclusions
Reliable reference databases are essential for the accuracy of DNA barcoding studies. However, a significant knowledge gap remains in the systematic evaluation of these databases, particularly in assessing both barcode coverage and quality of COI barcode sequences in NCBI and BOLD, focusing on marine metazoan species in the WCPO region.
Our findings indicate that NCBI exhibits higher barcode coverage than BOLD, which may be attributed to its broader research scope and less stringent submission requirements. In contrast, although BOLD is also open-access, its stricter quality control protocols and focus on biodiversity-specific data contribute to its relatively lower barcode coverage. However, this higher coverage comes at the cost of lower sequence quality, including a greater prevalence of short sequences, ambiguous nucleotides, and incomplete taxonomic information. In constrast, BOLD maintains a stricter quality control, but its reliance on curated metadata limits the number of submitted sequences, resulting in higher barcode deficiencies. The BIN system in BOLD demonstrates significant potential for identifying and addressing these problematic records, highlighting the advantages of curated databases.
Despite these differences, both databases contain a substantial proportion of problematic sequences, likely caused by contamination, cryptic species, sequencing errors, or inconsistent taxonomic assignments. Additionally, we identified severe barcode deficiencies in underrepresented marine regions—particularly in the south temperate WCPO—and among certain phyla such as the Bryozoa, Porifera, and Platyhelminthes. Even though the COI gene is widely regarded as a standard barcode for metazoan species, challenges remain in achieving clear species-level resolution for certain groups like Scombridae and Lutjanidae.
By addressing the critical gaps in barcode coverage, taxonomic representation, and sequence quality identified in this study, future barcoding initiatives can significantly enhance biodiversity monitoring and conservation research in the WCPO and beyond. This study underscores the importance of standardized protocols for database curation and sequencing practices, offering a roadmap for improving the reliability and applicability of DNA barcoding globally.
Supplemental Information
Acknowledgments
We thank Xing Quan for his assistance in analyzing data and creating figures.
Funding Statement
This work was supported by the University of Canberra and by the European Union “Pacific-European-Union-Marine-Partnership” Programme (agreement FED/2018/397-941) grant to the Pacific Community. This publication was produced with the financial support of the European Union and Sweden. Its contents are the sole responsibility of the authors and do not necessarily reflect the views of the European Union and Sweden. The funders had no role in study design, data collection and analysis, decision to publish, or preparation of the manuscript.
Additional Information and Declarations
Competing Interests
The authors declare there are no competing interests.
Author Contributions
Yufei Zhou conceived and designed the experiments, performed the experiments, analyzed the data, prepared figures and/or tables, authored or reviewed drafts of the article, and approved the final draft.
Alejandro Trujillo-González conceived and designed the experiments, authored or reviewed drafts of the article, and approved the final draft.
Simon Nicol conceived and designed the experiments, authored or reviewed drafts of the article, and approved the final draft.
Roger Huerlimann conceived and designed the experiments, authored or reviewed drafts of the article, and approved the final draft.
Stephen D. Sarre conceived and designed the experiments, authored or reviewed drafts of the article, and approved the final draft.
Dianne Gleeson conceived and designed the experiments, authored or reviewed drafts of the article, and approved the final draft.
Data Availability
The following information was supplied regarding data availability:
The script for the barcode evaluation workflow is available at GitHub and Zenodo:
- https://github.com/xyzzzeno/reference_evaluation.
- xyzzzeno. (2025). xyzzzeno/reference_evaluation: Code for Evaluation of DNA barcoding reference databases for marine species in the Western and Central Pacific Ocean (v1.0.0). Zenodo. https://doi.org/10.5281/zenodo.15644776
The raw barcode records are available in the Supplementary File.
References
- Alberdi et al. (2019).Alberdi A, Aizpurua O, Bohmann K, Gopalakrishnan S, Lynggaard C, Nielsen M, Gilbert MTP. Promises and pitfalls of using high-throughput sequencing for diet analysis. Molecular Ecology Resources. 2019;19(2):327–348. doi: 10.1111/1755-0998.12960. [DOI] [PubMed] [Google Scholar]
- Ardura (2019).Ardura A. Species-specific markers for early detection of marine invertebrate invaders through eDNA methods: gaps and priorities in GenBank as database example. Journal for Nature Conservation. 2019;47:51–57. doi: 10.1016/j.jnc.2018.11.005. [DOI] [Google Scholar]
- Baetscher et al. (2023).Baetscher DS, Locatelli NS, Won E, Fitzgerald T, McIntyre PB, Therkildsen NO. Optimizing a metabarcoding marker portfolio for species detection from complex mixtures of globally diverse fishes. Environmental DNA. 2023;5(6):1589–1607. doi: 10.1002/edn3.479. [DOI] [Google Scholar]
- Bazinet et al. (2018).Bazinet AL, Ondov BD, Sommer DD, Ratnayake S. BLAST-based validation of metagenomic sequence assignments. PeerJ. 2018;6:e4892. doi: 10.7717/peerj.4892. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Benson, Lipman & Ostell (1993).Benson D, Lipman DJ, Ostell J. GenBank. Nucleic Acids Research. 1993;21(13):2963–2965. doi: 10.1093/nar/21.13.2963. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Bista et al. (2017).Bista I, Carvalho GR, Walsh K, Seymour M, Hajibabaei M, Lallias D, Christmas M, Creer S. Annual time-series analysis of aqueous eDNA reveals ecologically relevant dynamics of lake ecosystem biodiversity. Nature Communications. 2017;8(1):14087. doi: 10.1038/ncomms14087. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Bouchet et al. (2017).Bouchet P, Rocroi J-P, Hausdorf B, Kaim A, Kano Y, Nützel A, Parkhaev P, Schrödl M, Strong EE. Revised classification, nomenclator and typification of gastropod and monoplacophoran families. Malacologia. 2017;61(1–2):1–526. doi: 10.4002/040.061.0201. [DOI] [Google Scholar]
- Bourlat et al. (2013).Bourlat SJ, Borja A, Gilbert J, Taylor MI, Davies N, Weisberg SB, Griffith JF, Lettieri T, Field D, Benzie J, Glöckner FO, Rodríguez-Ezpeleta N, Faith DP, Bean TP, Obst M. Genomics in marine monitoring: new opportunities for assessing marine health status. Marine Pollution Bulletin. 2013;74(1):19–31. doi: 10.1016/j.marpolbul.2013.05.042. [DOI] [PubMed] [Google Scholar]
- Bucklin, Steinke & Blanco-Bercial (2011).Bucklin A, Steinke D, Blanco-Bercial L. DNA barcoding of marine metazoa. Annual Review of Marine Science. 2011;3(1):471–508. doi: 10.1146/annurev-marine-120308-080950. [DOI] [PubMed] [Google Scholar]
- Burlakova et al. (2011).Burlakova LE, Karatayev AY, Karatayev VA, May ME, Bennett DL, Cook MJ. Endemic species: contribution to community uniqueness, effect of habitat alteration, and conservation priorities. Biological Conservation. 2011;144(1):155–165. doi: 10.1016/j.biocon.2010.08.010. [DOI] [Google Scholar]
- Chamberlain & Vanhoorne (2023).Chamberlain S, Vanhoorne B. worrms: world register of marine species (WoRMS) client. https://CRAN.R-project.org/package=worrms 2023
- Chen, Zobel & Verspoor (2017).Chen Q, Zobel J, Verspoor K. Duplicates, redundancies and inconsistencies in the primary nucleotide databases: a descriptive study. Database. 2017;2017:baw163. doi: 10.1093/database/baw163. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Chorlton (2024).Chorlton SD. Ten common issues with reference sequence databases and how to mitigate them. Frontiers in Bioinformatics. 2024;4:1278228. doi: 10.3389/fbinf.2024.1278228. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Clare (2014).Clare EL. Molecular detection of trophic interactions: emerging trends, distinct advantages, significant considerations and conservation applications. Evolutionary Applications. 2014;7(9):1144–1157. doi: 10.1111/eva.12225. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Costa et al. (2012).Costa FO, Landi M, Martins R, Costa MH, Costa ME, Carneiro M, Alves MJ, Steinke D, Carvalho GR. A ranking system for reference libraries of DNA barcodes: application to marine fish species from Portugal. PLOS ONE. 2012;7(4):e35858. doi: 10.1371/journal.pone.0035858. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Costello et al. (2010).Costello MJ, Coll M, Danovaro R, Halpin P, Ojaveer H, Miloslavich P. A census of marine biodiversity knowledge, resources, and future challenges. PLOS ONE. 2010;5(8):e12110. doi: 10.1371/journal.pone.0012110. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Cunha & Giribet (2019).Cunha TJ, Giribet G. A congruent topology for deep gastropod relationships. Proceedings of the Royal Society B: Biological Sciences. 2019;286(1898):20182776. doi: 10.1098/rspb.2018.2776. [DOI] [PMC free article] [PubMed] [Google Scholar]
- DeSantis et al. (2006).DeSantis TZ, Hugenholtz P, Larsen N, Rojas M, Brodie EL, Keller K, Huber T, Dalevi D, Hu P, Andersen GL. Greengenes, a Chimera-checked 16S rRNA gene database and workbench compatible with ARB. Applied and Environmental Microbiology. 2006;72(7):5069–5072. doi: 10.1128/AEM.03006-05. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Duarte, Vieira & Costa (2020).Duarte S, Vieira PE, Costa FO. Assessment of species gaps in DNA barcode libraries of non-indigenous species (NIS) occurring in European coastal regions. Metabarcoding and Metagenomics. 2020;4:e55162. doi: 10.3897/mbmg.4.55162. [DOI] [Google Scholar]
- Fadli et al. (2024).Fadli N, Rahayu SR, El Rahimi SA, Damora A, Muchlisin ZA, Ramadhaniaty M, Razi NM, Siregar AR, Ummamah ARA, Samad APA, Habib A, Siti-Azizah MN. DNA barcoding of commercially important snappers (Genus Lutjanus (Pisces: Lutjanidae)) from Aceh waters, Indonesia. International Journal of Design & Nature and Ecodynamics. 2024;19(3):787–794. doi: 10.18280/ijdne.190309. [DOI] [Google Scholar]
- Fontes et al. (2021).Fontes JT, Vieira PE, Ekrem T, Soares P, Costa FO. BAGS: an automated barcode, audit & grade system for DNA barcode reference libraries. Molecular Ecology Resources. 2021;21(2):573–583. doi: 10.1111/1755-0998.13262. [DOI] [PubMed] [Google Scholar]
- Galitz et al. (2021).Galitz A, Nakao Y, Schupp PJ, Wörheide G, Erpenbeck D. A soft spot for chemistry–current taxonomic and evolutionary implications of sponge secondary metabolite distribution. Marine Drugs. 2021;19(8):448. doi: 10.3390/md19080448. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Gold et al. (2021).Gold Z, Curd EE, Goodwin KD, Choi ES, Frable BW, Thompson AR, Walker HJ, Burton RS, Kacev D, Martz LD, Barber PH. Improving metabarcoding taxonomic assignment: a case study of fishes in a large marine ecosystem. Molecular Ecology Resources. 2021;21(7):2546–2564. doi: 10.1111/1755-0998.13450. [DOI] [PubMed] [Google Scholar]
- Gonçalves & Musen (2018).Gonçalves RS, Musen MA. The variable quality of metadata about biological samples used in biomedical experiments. Scientific Data. 2018;6(1):1–15. doi: 10.48550/ARXIV.1808.06907. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Günther et al. (2021).Günther B, Fromentin J-M, Metral L, Arnaud-Haond S. Metabarcoding confirms the opportunistic foraging behaviour of Atlantic bluefin tuna and reveals the importance of gelatinous prey. PeerJ. 2021;9:e11757. doi: 10.7717/peerj.11757. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Guo et al. (2022).Guo M, Yuan C, Tao L, Cai Y, Zhang W. Life barcoded by DNA barcodes. Conservation Genetics Resources. 2022;14(4):351–365. doi: 10.1007/s12686-022-01291-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Hadley et al. (2023).Hadley W, Romain F, Lionel H, Kirill M, Davis V. dplyr: a grammar of data manipulation (Version R package version 1.1.2) https://CRAN.R-project.org/package=dplyr 2023
- Hare et al. (2020).Hare SR, Williams PG, Ducharme-Barth ND, Hamer PA, Hampton WJ, Scott RD, Vincent MT, Pilling GH. The western and central Pacific tuna fishery: 2019 overview and status of stocks. Pacific Community, Noumea, New CaledoniaTuna Fisheries Assessment Report no. 20. 2020
- Hebert et al. (2003).Hebert PDN, Cywinska A, Ball SL, De Waard JR. Biological identifications through DNA barcodes. Proceedings of the Royal Society of London. Series B: Biological Sciences. 2003;270(1512):313–321. doi: 10.1098/rspb.2002.2218. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Heller et al. (2018).Heller P, Casaletto J, Ruiz G, Geller J. A database of metazoan cytochrome c oxidase subunit I gene sequences derived from GenBank with CO-ARBitrator. Scientific Data. 2018;5(1):180156. doi: 10.1038/sdata.2018.156. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Hestetun et al. (2020).Hestetun JT, Bye-Ingebrigtsen E, Nilsson RH, Glover AG, Johansen P-O, Dahlgren TG. Significant taxon sampling gaps in DNA databases limit the operational use of marine macrofauna metabarcoding. Marine Biodiversity. 2020;50(5):70. doi: 10.1007/s12526-020-01093-5. [DOI] [Google Scholar]
- Hills et al. (2011).Hills T, Brooks A, Atherton J, Rao N, James R, editors. Pacific island biodiversity, ecosystems, and climate change adaptation: building on nature’s resilience. Apia, Samoa: SPREP; 2011. [Google Scholar]
- Hoegh-Guldberg & Bruno (2010).Hoegh-Guldberg O, Bruno JF. The impact of climate change on the world’s marine ecosystems. Science. 2010;328(5985):1523–1528. doi: 10.1126/science.1189930. [DOI] [PubMed] [Google Scholar]
- Hou et al. (2018).Hou G, Chen W-T, Lu H-S, Cheng F, Xie S-G. Developing a DNA barcode library for perciform fishes in the South China Sea: species identification, accuracy and cryptic diversity. Molecular Ecology Resources. 2018;18(1):137–146. doi: 10.1111/1755-0998.12718. [DOI] [PubMed] [Google Scholar]
- Ip et al. (2019).Ip YCA, Tay YC, Gan SX, Ang HP, Tun K, Chou LM, Huang D, Meier R. From marine park to future genomic observatory? Enhancing marine biodiversity assessments using a biocode approach. Biodiversity Data Journal. 2019;7:e46833. doi: 10.3897/BDJ.7.e46833. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Iwasaki et al. (2013).Iwasaki W, Fukunaga T, Isagozawa R, Yamada K, Maeda Y, Satoh TP, Sado T, Mabuchi K, Takeshima H, Miya M, Nishida M. MitoFish and MitoAnnotator: a mitochondrial genome database of fish with an accurate and automatic annotation pipeline. Molecular Biology and Evolution. 2013;30(11):2531–2540. doi: 10.1093/molbev/mst141. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Jeunen et al. (2023).Jeunen G, Dowle E, Edgecombe J, Von Ammon U, Gemmell NJ, Cross H. crabs—a software program to generate curated reference databases for metabarcoding sequencing data. Molecular Ecology Resources. 2023;23(3):725–738. doi: 10.1111/1755-0998.13741. [DOI] [PubMed] [Google Scholar]
- Jukes & Cantor (1969).Jukes TH, Cantor CR. Mammalian protein metabolism. New York: Elsevier; 1969. Evolution of protein molecules; pp. 21–132. [DOI] [Google Scholar]
- Kanz (2004).Kanz C. The EMBL nucleotide sequence database. Nucleic Acids Research. 2004;33(Database issue):D29–D33. doi: 10.1093/nar/gki098. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Katoh (2002).Katoh K. MAFFT: a novel method for rapid multiple sequence alignment based on fast fourier transform. Nucleic Acids Research. 2002;30(14):3059–3066. doi: 10.1093/nar/gkf436. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Knebelsberger et al. (2014).Knebelsberger T, Landi M, Neumann H, Kloppmann M, Sell AF, Campbell PD, Laakmann S, Raupach MJ, Carvalho GR, Costa FO. A reliable DNA barcode reference library for the identification of the North European shelf fish fauna. Molecular Ecology Resources. 2014;14(5):1060–1071. doi: 10.1111/1755-0998.12238. [DOI] [PubMed] [Google Scholar]
- Krijthe, Van der Maaten & Krijthe (2018).Krijthe J, Van der Maaten L, Krijthe MJ. Package ‘Rtsne’. R Package Version 0.13https://github.com/jkrijthe/Rtsne 2018
- Lamoreux et al. (2006).Lamoreux JF, Morrison JC, Ricketts TH, Olson DM, Dinerstein E, McKnight MW, Shugart HH. Global tests of biodiversity concordance and the importance of endemism. Nature. 2006;440(7081):212–214. doi: 10.1038/nature04291. [DOI] [PubMed] [Google Scholar]
- Leite et al. (2020).Leite BR, Vieira PE, Teixeira MAL, Lobo-Arteaga J, Hollatz C, Borges LMS, Duarte S, Troncoso JS, Costa FO. Gap-analysis and annotated reference library for supporting macroinvertebrate metabarcoding in Atlantic Iberia. Regional Studies in Marine Science. 2020;36:101307. doi: 10.1016/j.rsma.2020.101307. [DOI] [Google Scholar]
- Leray et al. (2012).Leray M, Boehm JT, Mills SC, Meyer CP. Moorea BIOCODE barcode library as a tool for understanding predator–prey interactions: insights into the diet of common predatory coral reef fishes. Coral Reefs. 2012;31(2):383–388. doi: 10.1007/s00338-011-0845-0. [DOI] [Google Scholar]
- Leray et al. (2019).Leray M, Knowlton N, Ho S-L, Nguyen BN, Machida RJ. GenBank is a reliable resource for 21st century biodiversity research. Proceedings of the National Academy of Sciences of the United States of America. 2019;116(45):22651–22656. doi: 10.1073/pnas.1911714116. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Leray et al. (2013).Leray M, Yang JY, Meyer CP, Mills SC, Agudelo N, Ranwez V, Boehm JT, Machida RJ. A new versatile primer set targeting a short fragment of the mitochondrial COI region for metabarcoding metazoan diversity: application for characterizing coral reef fish gut contents. Frontiers in Zoology. 2013;10(1):34. doi: 10.1186/1742-9994-10-34. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Linneman et al. (2014).Linneman J, Paulus D, Lim-Fong G, Lopanik NB. Latitudinal variation of a defensive symbiosis in the Bugula neritina (Bryozoa) sibling species complex. PLOS ONE. 2014;9(10):e108783. doi: 10.1371/journal.pone.0108783. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Lobo et al. (2013).Lobo J, Costa PM, Teixeira MA, Ferreira MS, Costa MH, Costa FO. Enhanced primers for amplification of DNA barcodes from a broad range of marine metazoans. BMC Ecology. 2013;13(1):34. doi: 10.1186/1472-6785-13-34. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Macheriotou et al. (2019).Macheriotou L, Guilini K, Bezerra TN, Tytgat B, Nguyen DT, Phuong Nguyen TX, Noppe F, Armenteros M, Boufahja F, Rigaux A, Vanreusel A, Derycke S. Metabarcoding free-living marine nematodes using curated 18S and CO1 reference sequence databases for species-level taxonomic assignments. Ecology and Evolution. 2019;9(3):1211–1226. doi: 10.1002/ece3.4814. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Mackie, Keough & Christidis (2006).Mackie JA, Keough MJ, Christidis L. Invasion patterns inferred from cytochrome oxidase I sequences in three bryozoans, Bugula neritina, Watersipora subtorquata, and Watersipora arcuata. Marine Biology. 2006;149(2):285–295. doi: 10.1007/s00227-005-0196-x. [DOI] [Google Scholar]
- Marques et al. (2021).Marques V, Milhau T, Albouy C, Dejean T, Manel S, Mouillot D, Juhel J. GAPeDNA: assessing and mapping global species gaps in genetic databases for eDNA metabarcoding. Diversity and Distributions. 2021;27(10):1880–1892. doi: 10.1111/ddi.13142. [DOI] [Google Scholar]
- Miya et al. (2015).Miya M, Sato Y, Fukunaga T, Sado T, Poulsen JY, Sato K, Minamoto T, Yamamoto S, Yamanaka H, Araki H, Kondoh M, Iwasaki W. MiFish, a set of universal PCR primers for metabarcoding environmental DNA from fishes: detection of more than 230 subtropical marine species. Royal Society Open Science. 2015;2(7):150088. doi: 10.1098/rsos.150088. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Mugnai et al. (2021).Mugnai F, Meglécz E, CoMBoMed group. Abbiati M, Bavestrello G, Bertasi F, Bo M, Capa M, Chenuil A, Colangelo MA, De Clerck O, Gutiérrez JM, Lattanzi L, Leduc M, Martin D, Matterson KO, Mikac B, Plaisance L, Ponti M, Riesgo A, Costantini F. Are well-studied marine biodiversity hotspots still blackspots for animal barcoding? Global Ecology and Conservation. 2021;32:e01909. doi: 10.1016/j.gecco.2021.e01909. [DOI] [Google Scholar]
- Nagai et al. (2022).Nagai S, Sildever S, Nishi N, Tazawa S, Basti L, Kobayashi T, Ishino Y. Comparing PCR-generated artifacts of different polymerases for improved accuracy of DNA metabarcoding. Metabarcoding and Metagenomics. 2022;6:e77704. doi: 10.3897/mbmg.6.77704. [DOI] [Google Scholar]
- Newman et al. (2016).Newman SJ, Williams AJ, Wakefield CB, Nicol SJ, Taylor BM, O’Malley JM. Review of the life history characteristics, ecology and fisheries for deep-water tropical demersal fish in the Indo-Pacific region. Reviews in Fish Biology and Fisheries. 2016;26(3):537–562. doi: 10.1007/s11160-016-9442-1. [DOI] [Google Scholar]
- Nicol et al. (2013).Nicol SJ, Allain V, Pilling GM, Polovina J, Coll M, Bell J, Dalzell P, Sharples P, Olson R, Griffiths S, Dambacher JM, Young J, Lewis A, Hampton J, Jurado Molina J, Hoyle S, Briand K, Bax N, Lehodey P, Williams P. An ocean observation system for monitoring the affects of climate change on the ecology and sustainability of pelagic fisheries in the Pacific Ocean. Climatic Change. 2013;119(1):131–145. doi: 10.1007/s10584-012-0598-y. [DOI] [Google Scholar]
- NOAA Fisheries (2018).NOAA Fisheries . Washington: Department of Commerce; 2018. Western pacific non-commercial fisheries. [Google Scholar]
- Oliveira et al. (2016).Oliveira LM, Knebelsberger T, Landi M, Soares P, Raupach MJ, Costa FO. Assembling and auditing a comprehensive DNA barcode reference library for European marine fishes: DNA barcode library for european marine fishes. Journal of Fish Biology. 2016;89(6):2741–2754. doi: 10.1111/jfb.13169. [DOI] [PubMed] [Google Scholar]
- Pauly et al. (2002).Pauly D, Christensen V, Guénette S, Pitcher TJ, Sumaila UR, Walters CJ, Watson R, Zeller D. Towards sustainability in world fisheries. Nature. 2002;418(6898):689–695. doi: 10.1038/nature01017. [DOI] [PubMed] [Google Scholar]
- Preston, Fritzsche & Woodcock (2022).Preston M, Fritzsche M, Woodcock P. Peterborough: Joint Nature Conservation Committee; 2022. Understanding and mitigating errors and biases in metabarcoding: an introduction for non-specialists. [Google Scholar]
- Provoost & Bosch (2022).Provoost P, Bosch S. robis: ocean biodiversity information system (OBIS) client. https://CRAN.R-project.org/package=robis 2022
- Puillandre et al. (2012).Puillandre N, Bouchet P, Boisselier-Dubayle M-C, Brisset J, Buge B, Castelin M, Chagnoux S, Christophe T, Corbari L, Lambourdière J, Lozouet P, Marani G, Rivasseau A, Silva N, Terryn Y, Tillier S, Utge J, Samadi S. New taxonomy and old collections: integrating DNA barcoding into the collection curation process. Molecular Ecology Resources. 2012;12(3):396–402. doi: 10.1111/j.1755-0998.2011.03105.x. [DOI] [PubMed] [Google Scholar]
- Quast et al. (2012).Quast C, Pruesse E, Yilmaz P, Gerken J, Schweer T, Yarza P, Peplies J, Glöckner FO. The SILVA ribosomal RNA gene database project: improved data processing and web-based tools. Nucleic Acids Research. 2012;41(D1):D590–D596. doi: 10.1093/nar/gks1219. [DOI] [PMC free article] [PubMed] [Google Scholar]
- R Core Team (2019).R Core Team . R Foundation for Statistical Computing; 2019. [Google Scholar]
- Ramirez et al. (2020).Ramirez JL, Rosas-Puchuri U, Cañedo RM, Alfaro-Shigueto J, Ayon P, Zelada-Mázmela E, Siccha-Ramirez R, Velez-Zuazo X. DNA barcoding in the Southeast Pacific marine realm: low coverage and geographic representation despite high diversity. PLOS ONE. 2020;15(12):e0244323. doi: 10.1371/journal.pone.0244323. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Ratnasingham & Hebert (2007).Ratnasingham S, Hebert PDN. BARCODING: bold: the barcode of life data system (http://www.barcodinglife.org) Molecular Ecology Notes. 2007;7(3):355–364. doi: 10.1111/j.1471-8286.2007.01678.x. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Ratnasingham & Hebert (2013).Ratnasingham S, Hebert PDN. A DNA-based registry for all animal species: the barcode index number (BIN) system. PLOS ONE. 2013;8(7):e66213. doi: 10.1371/journal.pone.0066213. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Redelings (2014).Redelings B. Erasing errors due to alignment ambiguity when estimating positive selection. Molecular Biology and Evolution. 2014;31(8):1979–1993. doi: 10.1093/molbev/msu174. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Robeson et al. (2021).Robeson MS, O’Rourke DR, Kaehler BD, Ziemski M, Dillon MR, Foster JT, Bokulich NA. RESCRIPt: reproducible sequence taxonomy reference database management for the masses. Bioinformatics. 2021;7(11):e1009581. doi: 10.1101/2020.10.05.326504. [DOI] [PMC free article] [PubMed] [Google Scholar]
- RStudio Team (2022).RStudio Team . RStudio, PBC; Boston, MA: 2022. [Google Scholar]
- Sachs, Mellinger & Gallup (2001).Sachs JD, Mellinger AD, Gallup JL. The geography of poverty and wealth. Scientific American. 2001;284(3):70–75. doi: 10.1038/scientificamerican0301-70. [DOI] [PubMed] [Google Scholar]
- Shen, Chen & Murphy (2013).Shen Y-Y, Chen X, Murphy RW. Assessing DNA barcoding as a tool for species identification and data quality control. PLOS ONE. 2013;8(2):e57125. doi: 10.1371/journal.pone.0057125. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Song et al. (2008).Song H, Buhay JE, Whiting MF, Crandall KA. Many species in one: DNA barcoding overestimates the number of species when nuclear mitochondrial pseudogenes are coamplified. Proceedings of the National Academy of Sciences of the United States of America. 2008;105(36):13486–13491. doi: 10.1073/pnas.0803076105. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Steinke & Hanner (2011).Steinke D, Hanner R. The FISH-BOL collaborators’ protocol. Mitochondrial DNA. 2011;22(sup1):10–14. doi: 10.3109/19401736.2010.536538. [DOI] [PubMed] [Google Scholar]
- Sthle & Wold (1989).Sthle L, Wold S. Analysis of variance (ANOVA) Chemometrics and Intelligent Laboratory Systems. 1989;6(4):259–272. doi: 10.1016/0169-7439(89)80095-4. [DOI] [Google Scholar]
- Stoeckle, Das Mishu & Charlop-Powers (2020).Stoeckle MY, Das Mishu M, Charlop-Powers Z. Improved environmental DNA reference library detects overlooked marine fishes in New Jersey, United States. Frontiers in Marine Science. 2020;7:226. doi: 10.3389/fmars.2020.00226. [DOI] [Google Scholar]
- Tateno & Gojobori (1997).Tateno Y, Gojobori T. DNA data Bank of Japan in the age of information biology. Nucleic Acids Research. 1997;25(1):14–17. doi: 10.1093/nar/25.1.14. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Trujillo-González (2022).Trujillo-González A. Can stomach content and microbiomes of tuna provide near real-time detection of ecosystem composition in the Pacific Ocean? Frontiers in Marine Science. 2022;9:14. doi: 10.3389/fmars.2022.811532. [DOI] [Google Scholar]
- Turanov & Kartavtsev (2021).Turanov SV, Kartavtsev Y Ph. A complement to DNA barcoding reference library for identification of fish from the Northeast Pacific. Genome. 2021;64(10):927–936. doi: 10.1139/gen-2020-0192. [DOI] [PubMed] [Google Scholar]
- Valdez-Moreno et al. (2019).Valdez-Moreno M, Ivanova NV, Elías-Gutiérrez M, Pedersen SL, Bessonov K, Hebert PDN. Using eDNA to biomonitor the fish community in a tropical oligotrophic lake. PLOS ONE. 2019;14(4):e0215505. doi: 10.1371/journal.pone.0215505. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Vargas et al. (2012).Vargas S, Schuster A, Sacher K, Büttner G, Schätzle S, Läuchli B, Hall K, Hooper JNA, Erpenbeck D, Wörheide G. Barcoding sponges: an overview based on comprehensive sampling. PLOS ONE. 2012;7(7):e39345. doi: 10.1371/journal.pone.0039345. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Victor, Valdez-Moreno & Vásquez-Yeomans (2015).Victor BC, Valdez-Moreno M, Vásquez-Yeomans L. Status of DNA barcoding coverage for the tropical Western Atlantic shorefishes and reef fishes. DNA Barcodes. 2015;3(1):85–93. doi: 10.1515/dna-2015-0011. [DOI] [Google Scholar]
- Wangensteen et al. (2018).Wangensteen OS, Palacín C, Guardiola M, Turon X. DNA metabarcoding of littoral hard-bottom communities: high diversity and database gaps revealed by two molecular markers. PeerJ. 2018;6:e4705. doi: 10.7717/peerj.4705. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Weigand et al. (2019).Weigand H, Beermann AJ, Geiger MF, Weigand AM, Willassen E. DNA barcode reference libraries for the monitoring of aquatic biota in Europe: gap-analysis and recommendations for future work. Science of the Total Environment. 2019;26:499–524. doi: 10.1016/j.scitotenv.2019.04.247. [DOI] [PubMed] [Google Scholar]
- Wheeler (1994).Wheeler WC. Sources of ambiguity in nucleic acid sequence alignment. In: Schierwater B, Streit B, Wagner GP, De Salle R, editors. Molecular ecology and evolution: approaches and applications. Vol. 69. Basel, Switzerland: Birkhäuser Basel; 1994. pp. 323–352. [DOI] [PubMed] [Google Scholar]
- Wickham (2016).Wickham H. New York: Springer; 2016. [Google Scholar]
- Winter (2017).Winter DJ. rentrez: an R package for the NCBI eUtils API. PeerJ Preprints. 2017 doi: 10.7287/peerj.preprints.3179v2. [DOI] [Google Scholar]
- Xie et al. (2025).Xie J, Zhang Y, Wang L, Deng Y. DNA barcode contamination screen (DBCscreen): a pipeline to rapidly detect DNA barcode contamination for biodiversity research. Diversity. 2025;17(3):186. doi: 10.3390/d17030186. [DOI] [Google Scholar]
- Yeh et al. (2020).Yeh HD, Questel JM, Maas KR, Bucklin A. Metabarcoding analysis of regional variation in gut contents of the copepod Calanus finmarchicus in the North Atlantic Ocean. Deep Sea Research Part II: Topical Studies in Oceanography. 2020;180:104738. doi: 10.1016/j.dsr2.2020.104738. [DOI] [Google Scholar]
- Zhang et al. (2017).Zhang A, Hao M, Yang C, Shi Z. BarcodingR: an integrated r package for species identification using DNA barcodes. Methods in Ecology and Evolution. 2017;8(5):627–634. doi: 10.1111/2041-210X.12682. [DOI] [Google Scholar]
- Zhang & Zhang (2014).Zhang RL, Zhang B. Prospects of using DNA barcoding for species identification and evaluation of the accuracy of sequence databases for ticks (Acari: Ixodida) Ticks and Tick-Borne Diseases. 2014;5(3):352–358. doi: 10.1016/j.ttbdis.2014.01.001. [DOI] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Data Availability Statement
The following information was supplied regarding data availability:
The script for the barcode evaluation workflow is available at GitHub and Zenodo:
- https://github.com/xyzzzeno/reference_evaluation.
- xyzzzeno. (2025). xyzzzeno/reference_evaluation: Code for Evaluation of DNA barcoding reference databases for marine species in the Western and Central Pacific Ocean (v1.0.0). Zenodo. https://doi.org/10.5281/zenodo.15644776
The raw barcode records are available in the Supplementary File.





