Skip to main content
Wiley Open Access Collection logoLink to Wiley Open Access Collection
. 2022 Nov 21;25(5):e13733. doi: 10.1111/1755-0998.13733

Fast‐tracking bespoke DNA reference database generation from museum collections for biomonitoring and conservation

Andrew Dopheide 1, Talia Brav‐Cubitt 1, Anastasija Podolyan 2, Richard A B Leschen 1, Darren Ward 1,3, Thomas R Buckley 1,3, Manpreet K Dhami 2,
PMCID: PMC12142715  PMID: 36345645

Abstract

Despite recent advances in high‐throughput DNA sequencing technologies, a lack of locally relevant DNA reference databases limits the potential for DNA‐based monitoring of biodiversity for conservation and biosecurity applications. Museums and national collections represent a compelling source of authoritatively identified genetic material for DNA database development, yet obtaining DNA barcodes from long‐stored specimens may be difficult due to sample degradation. Here we demonstrate a sensitive and efficient laboratory and bioinformatic process for generating DNA barcodes from hundreds of invertebrate specimens simultaneously via the Illumina MiSeq system. Using this process, we recovered full‐length (334) or partial (105) COI barcodes from 439 of 450 (98%) national collection‐held invertebrate specimens. This included full‐length barcodes from 146 specimens which produced low‐yield DNA and no visible PCR bands, and which produced as little as a single sequence per specimen, demonstrating high sensitivity of the process. In many cases, the identity of the most abundant sequences per specimen were not the correct barcodes, necessitating the development of a taxonomy‐informed process for identifying correct sequences among the sequencing output. The recovery of only partial barcodes for some taxa indicates a need to refine certain PCR primers. Nonetheless, our approach represents a highly sensitive, accurate and efficient method for targeted reference database generation, providing a foundation for DNA‐based assessments and monitoring of biodiversity.

Keywords: conservation, DNA‐based monitoring, invertebrate barcoding, molecular taxonomy, museum collection, taxonomy‐informed bioinformatics pipeline

1. INTRODUCTION

Recent advances in high‐throughput sequencing technology are enabling a shift towards environmental DNA (eDNA)‐based methods for biodiversity assessment and biosecurity monitoring (Bohmann et al., 2014). While still in their infancy, these tools offer great promise for rapid and accessible biodiversity monitoring applications in terrestrial, aquatic and marine ecosystems (Deiner et al., 2021; Harper et al., 2019). However, a scarcity of accurately identified reference DNA sequence data from local biota (McGee et al., 2019) remains a significant obstacle to the application of these eDNA tools to biomonitoring, preventing the confident identification and interpretation of detected organisms.

DNA‐based identification methods, regardless of application, rely on the determination of similarity between newly detected sequences and existing reference sequence data (McGee et al., 2019; Schenekar et al., 2020). Depending on target taxa, this requires representative taxonomically validated data on established marker genes such as the ~650‐bp region of cytochrome c oxidase subunit I (COI) for metazoans (Folmer et al., 1994), or combinations of plastid regions (rbcl, matK, trnH–psbA) and the ribosomal internal transcribed spacer region (ITS) for plants (China Plant BOL Group et al., 2011). Large open data sources, such as the GenBank nr database or BOLD (Ratnasingham & Hebert, 2007), are typically employed as reference databases, but rely on data submitters for fidelity of sequence to organism and suffer from geographical sampling biases (Marques et al., 2021). Furthermore, this reliance on large pre‐existing databases also limits the emergence of new or taxon‐specific markers, with individual studies utilizing previously established markers even if they may be suboptimal for certain taxa (Luo et al., 2011). Curated databases containing only sequences from a targeted ecosystem may result in improved accuracy of sequence identifications compared to a global database (Gold et al., 2021). However, the current sparse database coverage of biodiversity from most ecosystems means that targeted reference databases typically must be populated with newly generated and locally relevant reference sequences.

Taxonomically validated reference sequences are difficult to generate. Not only do they require high levels of sequence accuracy (traditionally achieved via Sanger sequencing, more recently possible via PacBio hifi technology; D'Ercole et al., 2021), but also accurate taxonomic identification of specimens. The former may be time‐consuming, contingent on sample quality, and expensive, especially when applied to large numbers of specimens, while the latter requires specialist taxonomic expertise across taxa. For example, generating a reliable reference database for a previously uncharacterized insect fauna may require taxonomic skills spanning 24 distinct insect orders. Natural history museums and national biological collections, however, are unparalleled repositories of both invaluable taxonomic knowledge (Winker, 2004) and authoritatively identified genetic source material (Wandeler et al., 2007), with the potential to allow the efficient generation of taxonomically comprehensive and locally relevant reference DNA sequence databases (Hebert et al., 2013). Generating full‐length DNA barcodes via Sanger sequencing from dried or historical specimens stored over long periods may be difficult due to DNA degradation and low sensitivity of the sequencing approach, often resulting in only partial barcodes (Boyer et al., 2012; Hebert et al., 2013; Lindahl, 1993; Shokralla et al., 2011). Furthermore, museum samples are often indispensable permanent records, and therefore unavailable for destructive DNA extraction. Nondestructive extraction (Batovska et al., 2021; Carew et al., 2018) and PCR (Wong et al., 2014) approaches can be effective, however, depending on the taxa being analysed, and multiplex PCR coupled with high‐throughput DNA sequencing technologies has allowed the efficient recovery of barcodes from 50‐ to 100‐year‐old museum samples (D'Ercole et al., 2021; Prosser et al., 2016), as well as recently collected specimens (Shokralla et al., 2015).

There is a pressing need to utilize museum collections for rapid and cost‐effective generation of reference databases, in order to aid eDNA‐based biodiversity monitoring (de Santana et al., 2021). Here, we present a fast, cost‐effective and efficient method for developing a reference COI database from a diverse selection of terrestrial invertebrates sourced from the New Zealand Arthropod Collection (NZAC). These taxonomically validated specimens represent a variety of field‐collected methods, specimen treatment and storage conditions, as well as variable accessibility for destructive sampling. We demonstrate the use of a dual indexing approach, in combination with a pair of overlapping short PCR amplicons suitable for sequencing on the Illumina MiSeq platform, for generating full‐length barcodes from hundreds of invertebrate specimens simultaneously. We provide a taxonomy‐informed bioinformatics pipeline for processing and filtering the sequence data and the rapid assembly of successful barcodes. Together, our approach represents a highly sensitive, accurate and efficient method for targeted reference database generation, providing a foundation for DNA‐based assessments and monitoring of biodiversity.

2. MATERIALS AND METHODS

2.1. Specimen sampling and DNA extraction

Ethanol‐stored and pinned‐dry invertebrate specimens were selected from a variety of taxa, primarily earthworms (132 specimens) and insects (315 specimens, mainly beetles and wasps), as well as individual millipede, spider and mite specimens: 450 specimens in total (Table 1; Table S1). We randomly selected 1–4 individuals per species, depending on availability. Earthworm specimens were variously collected between 2004 and 2018, with exactly half (66) collected in 2010 and 2011. Arthropod samples were collected from 1997 to 2019, with the most from any one year being 68 in 2016, using a range of techniques, including malaise traps, pitfall traps, sweeping and hand collection. Most specimens (443 of 450) were collected in New Zealand from a diverse range of locations, including 105 from Northland, 29 from Southland and 23 from outlying islands (Rangitahua/Kermadecs, Manawatawhi/Three Kings, Chathams and Campbell). Six of the remaining specimens (five earthworms and one arthropod) were collected in Australia, while one earthworm was collected in New Caledonia. Most specimens were stored in 95% ethanol at room temperature.

TABLE 1.

PCR outcomes and DNA barcode recovery from 450 invertebrate specimens

Phylum Class Order N PCRs Alignments Barcodes
FC BR Full FC only BR only
Annelida Clitellata Haplotaxida 132 120 109 11.7 (0–105) 116 11 2
Arthropoda Arachnida Arachnida 2 0 2 1 (0–2) 0 0 1
Insecta Coleoptera 160 41 144 15.1 (0–133) 106 1 50
Diptera 21 8 16 9.8 (0–42) 17 1 3
Hemiptera 5 4 5 66.8 (10–122) 5 0 0
Hymenoptera 114 46 94 10.6 (0–180) 75 4 31
Lepidoptera 11 11 10 17.1 (1–125) 11 0 0
Others 4 3 3 48.2 (0–102) 3 0 1
Myriapoda Diplopoda 1 0 1 10 1 0 0
Totals 450 233 384 334 17 88

Note: PCR values are the numbers of FC and BR PCRs that resulted in a visible PCR product. Alignment values are the mean numbers of pairwise FC–BR alignments with 85‐bp overlap and 100% sequence identity with ranges in parentheses, among the ≤20 most abundant filtered and denoised FC and BR sequences per specimen (or all sequences for 12 specimens).

Total genomic DNA of specimens was extracted in a sterile environment with the following variations, depending on the size and availability of specimens for destructive sampling. Briefly, specimens were either soaked whole (112 specimens: 93 wasps, 10 beetles and nine months) or crushed whole (48 small specimens: 25 beetles, 12 flies, nine other insects, one spider and one mite) in lysis buffer, or a piece of tissue was sampled (290 specimens; all 132 earthworms, 125 beetles, 32 other insects and one millipede) and added to the lysis buffer. DNA extraction was carried out using one of the following kits following the manufacturer's instructions: DX reagents kit on the X‐tractor Gene (Qiagen) (211 specimens; 115 earthworms, 96 insects, and one mite, spider and millipede); QIAamp 96 DNA QIAcube HT Kit on the QIAcube HT (Qiagen) (222 insects, including the 112 soaked specimens); or the AquaPure Genomic DNA Isolation kit (BioRad) (17 earthworms). Extracted DNA was stored at −20°C until amplicon library construction.

2.2. Library construction and sequencing

We amplified two short overlapping fragments, FC (235 bp) and BR (428 bp), that together form the standard 658‐bp DNA barcode region, located at the 5′ end of the cytochrome c oxidase subunit I (COI) gene (Shokralla et al., 2015). These short fragments are expected to enhance PCR success rates and ensure compatibility with read length constraints of the Illumina MiSeq system. The FC and BR fragments were amplified using the primer pairs Ill_LCO1490 and Ill_C_R, and Ill_B_F and Ill_HCO2198, respectively (Shokralla et al., 2015). Forward and reverse primers were tagged with Nextera XT forward or reverse adapters (Illumina), respectively, padded with 0–4 mer nucleotide spacers to increase sequence heterogeneity, as described in Moinet et al. (2021). PCRs for both amplicons were carried out using the FastStart Taq DNA Polymerase kit (Roche), with final concentrations of 1× PCR buffer with MgCl2, 0.2 nm dNTP mix, 1 μg ml−1 bovine serum albumin (BSA), 250 nm of primer Ill_LCO1490 or Ill_HCO2198 and 750 nm of primer Ill_C_R or Ill_B_F, 1U Taq polymerase, and PCR‐grade water up to a total reaction volume of 20 μl. PCR cycle conditions were 95°C for 5 min; 40 cycles of 95°C for 45 s, 50°C for 45 s and 72°C for 45 s; and 72°C for 5 min (Veriti 96‐Well Thermal Cycler; Applied Biosystems). Following this, 4 μl of PCR product for each sample was run on a 1.5% agarose gel to check amplification success. First‐step PCR amplicons were cleaned, normalized for maximum concentration, and size selected using SerraMag SpeedBeads magnetic carboxylate modified particles (Cytiva) with 0.8× bead: buffer concentration (Toju et al., 2018) to remove primer dimers before the second‐step PCR.

Second‐step PCRs were performed using the KAPA3G Plant PCR kit (Kapa Biosystems), with Fusion primers with custom indices synthesized with P5/P7 Illumina adapters (forward, 5′‐AATGATACGGCGACCACCGAGATCTACAC – [8‐mer tag] – TCGTCGGCAGCGTC‐3′; reverse, 5′‐CAAGCAGAAGACGGCATACGAGAT‐ [8‐mer tag] – GTCTCGTGGGCTCGG‐3′) (Hamady et al., 2008), using 400 nm of each primer and 1.44 μl of first‐step PCR product in 18 μl volume. The second‐step PCR cycle was 95°C for 2 min, then 5 cycles of 95°C for 20 s, 50°C for 20 s and 72°C for 30 s, followed by 72°C for 2 min (Bio‐Rad T100TM Thermal Cycler). Then 2.5 μl of second‐step PCR products were run on a 2% agarose gel to check amplification success.

The amplicons were again cleaned, normalized, and size selected using SerraMag SpeedBeads magnetic carboxylate modified particles (Cytiva) with 1.5× bead: buffer concentration. Amplicons were eluted in 15 μl of 0.1× TE buffer, and pooled together at equal volumes to form the amplicon library, which was analysed on the Automated Bioanalysis System LabChip GX Touch HT (Perkin‐Elmer). DNA concentration was estimated with a Qubit 2.0 Fluorometer (Invitrogen), and adjusted to 4 nm. The final library was processed on an Illumina MiSeq 3000 sequencer with 2 × 300 paired‐end mode at the Genomics Facility at the University of Auckland, with 10% PhiX spiked in.

2.3. Bioinformatics pipeline

We obtained DNA barcodes from raw sequence data with the following process, using bcl2fastq version 2.20 (Illumina), claident version 2018.05.08 (Tanabe & Toju, 2013), vsearch 2.14 (Rognes et al., 2016) and emboss merger version 6.6.0 (Rice et al., 2000). Steps 1–8 were executed in Bash scripts, and steps 9 and 10 in a Python script:

  1. Raw sequence data in bcl format were converted into fastq format, without demultiplexing, using bcl2fastq.

  2. The forward and reverse fastq sequences were each trimmed of sequencing adapters and primers and demultiplexed into separate FC and BR fastq files for each of the 450 specimens, using the clsplitseq script from claident. This resulted in a pair of (adapter‐ and primer‐trimmed) R1 and R2 fastq files for the FC amplicon, and also for the BR amplicon, for each of the 450 specimens.

  3. For each amplicon (FC and BR) and specimen, the R1 and R2 sequences contained in each pair of fastq files were merged using ‐fastq_mergepairs in vsearch with minimum overlap (−fastq_minovlen), minimum merged length (−fastq_minmergelen) and maximum merged length (−fastq_maxmergelen) options of 200, 250 and 400 respectively for FC, and 100, 350 and 500 respectively for BR, with simultaneous filtering of merged sequences for errors (−fastq_maxee 1) and ambiguous bases (−fastq_maxn 0).

  4. The filtered sequences from each amplicon and specimen were then dereplicated using ‐derep_fulllength, denoised using ‐cluster_unoise with ‐minsize 1, and filtered for chimeras using ‐uchime3_denovo, all in vsearch. Most of the filtered and denoised sequence files still contained multiple sequences, necessitating additional filtering to recover the correct barcode sequences, based on taxonomy, sequence abundance and attributes of pairwise FC–BR sequence alignments, using subsequent steps.

  5. A taxonomic identification was obtained for each filtered and denoised FC and BR sequence using blast against the GenBank nr database, accepting the top match in each case.

  6. Up to n = 20 most abundant denoised FC and BR sequences (if any) were output as fasta files for each specimen, using the vsearch command ‐derep_fulllength with option ‐topn n. The FC and BR sequence files were concatenated together into a single fasta file per specimen.

  7. All possible pairwise sequence alignments with the expected overlap between FC and BR amplicons, namely an overlap of 85 bp with 100% pairwise sequence identity, were identified between the sequences in the concatenated FC and BR fasta file for each specimen using the vsearch command ‐allpairs_global with options ‐id 1 and ‐mincols 85.

  8. If no alignments were detected among the 20 most abundant sequences from a given specimen, steps 6 and 7 were repeated with all sequences (if >20) from that specimen.

  9. Putatively correct full‐length or partial barcodes were then identified per specimen based on comparing FC and BR sequence identifications and FC–BR alignment characteristics. An optimal FC sequence, BR sequence and aligned FC–BR sequence pair (if any) was selected for each specimen by matching their sequence identities to the a priori specimen taxonomy, prioritizing the lowest taxonomic rank level. The sequences were also required to have lengths between 324 and 326 bp for FC and between 417 and 419 bp for BR, and blast identification bitscores ≥200 for FC and ≥250 for BR, to avoid spurious matches. If more than one FC sequence, BR sequence or aligned FC–BR sequence pair per specimen met these criteria, the most abundant sequence or sequence pair was selected.

  10. If an FC–BR sequence pair with expected taxonomy was detected, and the lowest taxonomic identification rank of this sequence pair was equal to or lower than that of the selected FC and/or BR sequences for that specimen, that aligned FC–BR sequence pair was accepted as the probable correct barcode components and combined by emboss merger. This carries out a pairwise alignment according to the Needleman–Wunsch algorithm, outputting a merged barcode sequence, along with an alignment summary that was checked to ensure the alignment attributes were as expected (overlap of 85 bp and score of 425).

  11. To avoid accepting a merged sequence from a nontarget organism, if either of the selected FC or BR sequences had a lower taxonomic identification rank than that of the selected aligned FC–BR sequence pair, that FC or BR sequence was accepted as a probable correct partial barcode, instead of the aligned FC–BR sequence pair.

  12. If an FC or BR sequence (but not an aligned FC–BR sequence pair) with expected taxonomy was detected, that sequence was accepted as a probable correct partial barcode. If both an FC and BR sequence (but no aligned FC–BR sequence pair) with expected taxonomy were detected, the FC or BR sequence with the lowest taxonomic rank was accepted as a probable correct partial barcode. If these FR and BR sequences had the same lowest taxonomic rank, the most abundant sequence was accepted as the probable correct partial barcode.

The merged FC and BR barcodes, plus any partial FC‐ or BR‐only barcodes, were concatenated into a single fasta file, and aligned using mafft (Katoh & Standley, 2013). An approximately maximum‐likelihood phylogeny was generated from the alignment using fasttree 2 (Price et al., 2010), and visualized using the R package ggtree (Yu et al., 2017).

The resulting barcode data set was examined for any effects of DNA extraction methodology, PCR primer or taxonomy on successful barcode recovery. For this purpose, any orders and families represented by fewer than five specimens were pooled together as “Others.”

To assess the accuracy of recovered Illumina barcodes, 96 of the 450 specimen DNA extracts (47 beetles, 21 flies, 14 wasps, 11 other insects, and individual mite, millipede and spider specimens) were also subjected to PCR using primers LCO1490 and HCO2198 (Folmer et al., 1994), followed by Sanger sequencing of the amplicons using standard methods. The resulting Sanger barcode sequences were compared with Illumina‐derived barcodes for the same specimens via generation of pairwise sequence alignments using mafft (Katoh & Standley, 2013) and determination of sequence identity between each sequence pair.

3. RESULTS

3.1. Overall DNA barcoding success

Our barcoding process resulted in full‐length COI barcodes for 334 specimens (74.2%), plus FC‐only and BR‐only barcodes for a further 17 (3.78%) and 88 (19.6%) specimens, respectively. In total, full‐length or partial COI barcodes were recovered for 439 of 450 specimens (97.6%) (Tables 1, S2, Figure 1). This included full‐length/partial barcodes for 87.9%/9.85% of earthworm specimens, 66.2%/31.9% of Hymenoptera specimens and 65.8%/30.7% of Coleoptera specimens, respectively.

FIGURE 1.

FIGURE 1

An approximately maximum‐likelihood phylogeny of 334 full‐length (FC + BR) and 105 partial (FC only or BR only) COI barcodes recovered from 450 invertebrate specimens using Illumina sequencing, generated using fasttree 2 (Price et al., 2010)

3.2. PCR and sequencing outcomes

Initial PCR success rates differed between the FC and BR amplicons and among different taxa (Tables 1, S2). Visible PCR products were amplified from 233 of 450 specimens (51.7%) in FC PCRs, compared to 384 specimens (85.3%) in BR PCRs. Both FC and BR PCRs visibly succeeded for 205 specimens (45.6%). Visible FC PCR success rates were lowest for Coleoptera (25.6%) followed by Diptera (38.1%), and highest for Annelida (90.1%). In contrast, visible BR PCR success rates exceeded 70% for all orders except for Psocoptera (with only two specimens), including 90% for Coleoptera, 72.6% for Diptera and 86.2% for Annelida.

The numbers of sequences per amplicon and specimen after filtering and denoising varied widely, from zero (for 67 FC PCRs and five BR PCRs) to several thousand, with means of 72 in FC PCRs and 161 in BR PCRs (Table S1). High numbers of denoised FC sequences per specimen were strongly correlated with high numbers of denoised BR sequences per specimen (Pearson correlation coefficient = .85, p < .001). The numbers of pairwise alignments between FC and BR amplicons with expected characteristics (overlap of 85 bp and 100% identity) per specimen ranged between zero (for 92 specimens) and 180 (for a Hymenoptera specimen), with a mean of 14. Numbers of pairwise alignments were not obviously correlated with numbers of denoised FC and BR sequences. Low to moderate numbers of alignments were detected for 140 specimens from which FC PCRs did not produce visible products, and 39 specimens from which BR PCRs did not produce visible products (Table S1).

After taxonomic filtering of sequences and pairwise alignments to identify optimal barcodes, full‐length COI barcodes were recovered for 334 of 450 specimens (74%). This included full‐length barcodes for 146 specimens from which FC PCRs (109), BR PCRs (22) or both (15) did not result in a visible PCR product. Partial barcodes in the form of BR sequences only were recovered for a further 88 specimens, and FC sequences only for another 17 specimens.

No evidence of DNA extraction methodology effects on barcoding outcomes was observed. Rather, the most obvious factor affecting successful barcode detection was a combination of PCR amplicon and taxonomy (Table 2). FC sequences (either full‐length or FC‐only barcodes) were successfully detected for 75% of specimens on average across 19 different orders and families considered, compared to >94% for BR sequences. Rates of FC sequence detection were lower than rates of BR sequence detection in 13 groups, the same in five and higher in only one (Annelida, by 7%). Among insect taxa, FC sequence detection rates were the same as BR sequence detection rates for two orders (Hemiptera and Lepidoptera) and one family (Ichneumonidae), and lower for all other insect groups. The largest difference between successful FC and BR sequence detection rates was observed for Staphylinidae (10 specimens, −100%), with disparities ≥−20% observed for a further five insect groups including Chrysomelidae (107 specimens, −30%) and Braconidae (87 specimens, −23%).

TABLE 2.

Disparities between FC and BR sequence detection rates among different taxonomic families. FC and BR recovery rates represent the proportions of specimens for which full‐length, plus either FC‐only or BR‐only barcodes, respectively, were detected

Phylum Class Order Family Specimens FC recovery BR recovery Disparity (%)
Annelida Clitellata Haplotaxida 132 96 89 7
Arthropoda Arachnida Arachnida 2 0 50 −50
Insecta Coleoptera Carabidae 12 83 100 −17
Cerambycidae 6 83 100 −17
Chrysomelidae 107 67 97 −30
Coccinellidae 7 86 100 −14
Curculionidae 5 80 100 −20
Staphylinidae 10 0 100 −100
Others 13 77 92 −15
Diptera Ephydridae 5 80 80 0
Others 16 88 100 −12
Hemiptera 5 100 100 0
Hymenoptera Bethylidae 6 83 100 −17
Braconidae 87 69 92 −23
Ichneumonidae 7 100 100 0
Others 14 50 93 −43
Lepidoptera 11 100 100 0
Others 4 75 100 −25
Myriapoda Diplopoda 1 100 100 0
Mean 24 75 94 −20

3.3. PCR amplification specificity

Sequences identified as the most likely partial barcodes by taxonomic filtering were the maximally abundant sequence per specimen in 345 cases for the FC amplicon and in 365 cases for the BR amplicon. Combined, the selected FC and BR sequences were both the maximally abundant sequences per specimen in only 286 cases. A further 29 selected FC sequences and 50 selected BR sequences each had abundance ranks between two and 10. Two selected FC sequences had abundance ranks of 18 and 30, respectively, and 10 selected BR sequences each had abundance ranks between 11 and 101. In 37 cases where the maximally abundant FC sequence was not selected, the maximally abundant sequences were identified as deriving from insects (25), annelids (eight), arachnids (two), algae (one) and amoebae (one). Similarly, in 74 cases where the maximally abundant BR sequence was not selected, these were identified as deriving from insects (44), other hexapods (two), annelids (two), or gastropods (one); and as Homo sapiens (five), and eukaryote (two) or prokaryote (18) micro‐organisms, including 10 cases of Wolbachia.

To investigate the origins of maximally abundant but nontarget sequences (i.e., those with unexpected taxonomic identifications), the allpairs_global function in vsearch was used to identify any identical sequences among the maximally abundant sequences per specimen, plus the selected (presumed correct) sequences (if not maximally abundant) from each specimen. Out of 37 specimens with maximally abundant but nontarget FC sequences, only three of those sequences were identical to another maximally abundant but nontarget FC sequence, in two cases from adjacent PCR wells and in all three cases from within the same PCR plates. Similarly, out of 74 specimens with maximally abundant but nontarget BR sequences, 10 of those sequences were each identical to one or more other maximally abundant but nontarget BR sequence, in only four cases from adjacent PCR wells but in all cases from within the same PCR plates.

3.4. Comparison of MiSeq barcodes with sanger barcodes

Full‐length barcodes were successfully recovered by both MiSeq and Sanger barcoding approaches for 68 of 96 specimens subjected to both methods. Full‐length Illumina barcodes were obtained from seven specimens, and partial Illumina barcodes (one FC‐only and eight BR‐only) from a further nine specimens, that each failed to produce Sanger barcodes; 15 of these 16 specimens were beetles. On the other hand, only partial barcodes (one FC‐only and nine BR‐only) were recovered using the Illumina method from 10 specimens from which full‐length Sanger barcodes were obtained. No barcodes were obtained using either method from just two specimens (Ephutomorpha bivulnerata, a wasp; and the single mite specimen).

Among 68 specimens from which both Sanger and Illumina barcodes were recovered, pairwise sequence identities between these barcodes were from 75.5% to 85.7% for three specimens, 97.6–98.6% for six specimens, 99.1–99.9% for 18 specimens, and 100% for 40 specimens.

4. DISCUSSION

A lack of high‐quality and location‐specific reference sequence data appears to limit the potential for DNA‐based monitoring of terrestrial biodiversity, despite the great promise of these techniques. A scarcity of invertebrate taxonomic expertise and an associated lack of authoritatively identified specimens, and the high costs of Sanger sequencing, pose significant barriers to reference database generation. We present an efficient strategy that helps to overcome these barriers, with the potential to lessen the need for reliance on publicly available databases with inadequate local relevance for biodiversity monitoring. By utilizing a set of taxonomically identified specimens from a national collection coupled with a sensitive high‐throughput sequencing approach, we rapidly generated a reference COI sequence database consisting of full‐length barcodes for 334 specimens and partial (FC or BR) barcodes for a further 105 specimens, representing a wide range of invertebrate taxa from a diverse range of locations and with varied storage conditions. We observed no obvious effects of sampling or DNA extraction methods on barcoding success, indicating that a variety of protocols and specimen types can provide acceptable outcomes using this process. Furthermore, BR sequences were recovered from nearly all specimens, highlighting the potential for this process to achieve exceptionally high rates of DNA barcoding success, apparent deficiencies with the FC PCR primers notwithstanding.

While this analysis represents only a small portion of the source collection, the number of specimens included was arbitrarily limited, and there is considerable scope to greatly increase the throughput of this process. We recovered an average of over 10,000 sequences per amplicon and specimen (albeit with considerable variance). Given that the theoretical capacity of the MiSeq system exceeds 20 million sequence reads, and that two correct sequences (one FC and one BR sequence) are required to form a complete barcode, this suggests that 10,000 specimens could be sequenced in a single MiSeq run using this process at an average sequencing depth of 2000 reads per specimen (or 1000 reads per amplicon per specimen), although to our knowledge this remains to be tested. Typically, as the number of samples pooled together increases, the read depth per sample decreases, which may influence the detection of sequences from specimens that were difficult to amplify. Increasing the number of specimens will increase the costs of DNA extractions and sequencing library preparation, while the costs of the sequencing run, and subsequent data processing, tend to be fixed. The choice of DNA extraction method is thus an important determinant of costs. Possible methods include relatively expensive tissue extraction kits producing high‐quality DNA extracts (used in this study), nondestructive methods that aim to preserve intact specimens (Rowley et al., 2007; Wong et al., 2014), and the quick and inexpensive hot NaOH‐based “HotSHOT” method (Truett et al., 2000), which results in more degradation‐prone extracts (Srivathsan et al., 2021). Previous attempts to obtain DNA barcodes from multiple invertebrate specimens have used a variety of sequencing approaches. In one example, Sanger sequencing was used to obtain DNA barcodes from 86% of over 40,000 museum‐held Lepidoptera specimens, demonstrating a profound effect of specimen age on barcoding success using this method (Hebert et al., 2013). However, this effort required 6 months of molecular work by five people, illustrating the inefficiencies/impracticality of Sanger sequencing applied to large numbers of specimens. Invertebrate DNA barcoding efforts utilizing high‐throughput sequencing technologies typically report greater efficiency, lower costs and higher barcoding success rates than equivalent Sanger sequencing‐based efforts (D'Ercole et al., 2021; Hebert et al., 2018; Prosser et al., 2016). The two‐amplicon PCR approach used in this study was previously used to obtain barcodes from 97% of >1000 freshly trapped arthropod specimens (Shokralla et al., 2015). We achieved comparable success rates from older specimens, from a wide range of locations, including a diverse selection of earthworms, confirming the utility of this MiSeq approach for efficiently barcoding numerous specimens from diverse lineages and sources. On the other hand, the same approach applied to barcoding of dried saproxylic beetle specimens achieved a lower success rate of 55%, perhaps due to specimen collection methods being suboptimal for DNA preservation (Sire et al., 2019). The MiSeq system has also been used in a multilocus metabarcoding approach for detecting insect pests in bulk trap catches, which confirmed the importance of taxonomic information for confirming metabarcoding outcomes (Batovska et al., 2021).

Single molecule real‐time (SMRT) sequencing on the Pacific Biosciences Sequel platform has recently been used to recover DNA barcodes from about 85% of some 10,000 insect specimens (Hebert et al., 2018), and to recover barcodes from hundreds of ~50‐year‐old butterfly specimens (D'Ercole et al., 2021). This system provides sequence reads long enough to encompass the full COI barcode region but with higher error rates than those from Illumina platforms, with accurate “consensus” sequences obtainable by repeated reads of individual circularized molecules. This system has been argued to be a more economic high‐throughput barcoding system than Illumina MiSeq, due to the ability to obtain barcodes from 10,000 or more input specimens from one SMRT cell, and not requiring the amplification and processing of two amplicons per specimen (Hebert et al., 2018). However, our taxonomy‐informed pipeline provides a solution for barcode assembly from the two amplicons; furthermore, the Illumina MiSeq system has a theoretically comparable (and flexible) barcoding capacity, and is arguably more accessible in terms of platform availability and sequencing run costs than the PacBio Sequel system. This may make the MiSeq approach more suitable for biomonitoring programmes with limited budgets or modest numbers of specimens (hundreds to low thousands) requiring identification. The Illumina NovaSeq system, meanwhile, could be the most economic platform for barcode generation due to enormous sequencing depth (Srivathsan et al., 2021), but the maximum read length of 250 bp precludes merging of the R1 and R2 reads for the BR fragment.

Another rapidly developing approach to sequencing is Oxford Nanopore technology, which can provide long single‐molecule sequence reads carried out on low‐cost and portable devices (Leggett & Clark, 2017). These platforms enable rapid species identification and real‐time biomonitoring in the field, especially in locations without access to molecular sequencing facilities (Krehenwinkel et al., 2019; Pomerantz et al., 2018; Srivathsan et al., 2019). Much like the Sequel, this system has higher error rates than Illumina platforms, but accurate “consensus” barcodes can be obtained by assembly of multiple sequences per specimen (Srivathsan et al., 2018). Recent improvements to base‐calling accuracy and data processing tools have now made it feasible to recover barcodes from hundreds to thousands of specimens, using Oxford Nanopore Flongle and MinION devices, respectively (Srivathsan et al., 2021).

Illumina MiSeq, PacBio Sequel and Oxford Nanopore sequencing platforms therefore all offer viable high‐throughput barcoding options, with broadly comparable capacity and scalability but varying accessibility, and data processing methods and requirements specific to each case. These systems will probably provide the foundations of the next generation of biomonitoring tools. Further research to evaluate the relative sensitivities of these platforms, especially Oxford Nanopore, for generating accurate barcodes from old or degraded specimens and from a diverse range of taxonomic lineages, will help to inform the use of these systems in biomonitoring programmes. Below, we discuss some of the salient features and limitations of our MiSeq‐based approach.

4.1. Sensitivity, specificity and accuracy of Illumina barcode generation

The balance between sensitivity and specificity of high‐throughput sequencing is often difficult to maintain (Sint et al., 2012), and in our case is tilted in favour of sensitivity to increase the rate of barcode recovery. This high sensitivity of Illumina sequencing enabled the recovery of numerous complete DNA barcodes from PCRs that did not work well enough to produce a product visible by gel electrophoresis. Such specimens, accounting for almost half of our samples for the FC fragment, would probably fail to produce Sanger sequences, and be relegated to the “difficult to sequence” set of preserved specimens (Hajibabaei et al., 2006). Furthermore, full‐length barcodes were recovered from various specimens with clearly suboptimal sequencing outcomes, including eight specimens with a single FC sequence (after error filtering), another with a single BR sequence, and a further 14 specimens with between two and five FC or BR sequences. On the other hand, this sensitivity also allowed the amplification of nontarget sequences from well‐known sources of contamination, such as extraction buffers (“kitome”) (Paniagua Voirol et al., 2021), human specimen handling or bacterial symbionts (Sicard et al., 2019), as well as apparent cross‐contaminant sequences from other samples. Indeed, many of the maximally abundant sequences that were not selected as part of correct barcodes were similar or identical to those from other specimens included in the analysis, suggesting that these may variously result from PCR errors (Potapov & Ong, 2017), co‐amplification of numts (Song et al., 2008), cross‐contamination during library preparation (Minich et al., 2019) or index switching during sequencing (Schnell et al., 2015). Similar issues were observed in a study using the same sequencing approach to barcode a saproxylic beetle collection (Sire et al., 2019), indicating such contaminants may be an inevitable consequence of applying a highly sensitive method to specimens that may not have been collected with DNA analyses in mind. These issues might be mitigated in future analyses by stringent laboratory protocols to limit contamination (Eisenhofer et al., 2019), alternative library generation workflows (Bohmann et al., 2022) and fewer PCR cycles (Sze & Schloss, 2019), although the last might be at the expense of sensitivity. Additionally, bioinformatic tools can be used to identify correct barcodes among sequencing output. In this case, for example, the presence of contaminants necessitated the development of a semi‐automated barcode assembly pipeline to accurately resolve barcode sequences.

To assess recovered barcode sequence accuracy, 96 of the specimens included in this analysis were also subjected to DNA barcoding via Sanger sequencing. Full‐length barcodes were recovered from most of these specimens (68) by both methods. However, only Illumina barcodes (seven full‐length and nine partial) were obtained from 16 specimens that failed to produce Sanger barcodes. This included 15 beetles, which can be challenging DNA barcoding subjects due to their tough exoskeletons, further illustrating the sensitivity of the Illumina approach. On the other hand, only partial barcode sequences were obtained via Illumina sequencing from 10 specimens from which full Sanger barcodes were obtained. Some of these required multiple PCR optimization attempts to obtain Sanger barcodes, however, along with examination and manual editing of sequence chromatograms. High levels of pairwise sequence identity were observed between Sanger and Illumina barcodes from most specimens (99%–100% for 59 of 68 specimens), indicating generally high levels of accuracy for both sequencing approaches. Three specimens (all beetles) had pairwise sequence identities of between 75.5% and 85.7%, suggesting that either the Illumina approach or the Sanger approach recovered a sequence from a nontarget organism in these cases. These FC–BR sequence pairs were each selected due to their constituent FC and/or BR fragments being the most abundant sequences identified to the correct taxonomic families, with no lower rank taxonomic information available among the blast results to further guide correct barcode selection.

4.2. Using taxonomy to assemble barcodes

Because the maximally abundant sequences from 20% to 30% of specimens were not from the expected taxa according to blast, we developed a taxonomy‐weighted barcode‐assembly approach. For each specimen, we considered the most abundant FC and/or BR sequences with the expected taxonomic identifications—prioritizing the lowest identifiable taxonomic rank in each case—to be correct and considered merged FC and BR sequences to be correct barcodes only if the contributing sequences had 100% identity across the expected overlap length. This approach typically identified correct FC and BR sequences among the 20 most abundant sequences per specimen, but in a small number of cases, the correct sequences were identified at abundance ranks between 21 and 101. The ability to examine multiple sequences for correct identity is a key advantage of this process over Sanger sequencing, in which only a single sequence per specimen can typically be examined. This simple yet effective filtering approach greatly enhanced successful barcode recovery and provided evidence against relying solely on sequence abundances to select barcode sequences. Directly utilizing a priori taxonomic data from validated specimens allowed accurate identification of nontarget contaminant sequences, further stressing the value that taxonomically validated specimens can confer towards barcode generation. Similarly, taxonomic information was considered important for confirming the identity of insect pests detected in bulk trap catches by multilocus metabarcoding (Batovska et al., 2021).

4.3. Limitations and potential improvements

The lower rate of FC sequence recovery compared to BR sequence recovery implies that factors associated with FC PCRs, rather than sampling or DNA extraction, were the main cause of failures to obtain complete barcode sequences. The most probable explanation for this is that one or both primers used in FC PCRs have suboptimal matches with the specimens in question (Elbrecht et al., 2019). While it is unclear which of the FC primers (Ill_LCO1490 and Ill_C_R) might cause this problem, deficiencies of the LCO1490/HCO2198 primer pair have been noted previously (Geller et al., 2013; Lobo et al., 2013), pointing to LCO1490 as problematic. These failures were concentrated in certain Coleoptera and Hymenoptera families, suggesting that the primer sequences may need adjustment to improve outcomes for these groups.

Improvements to our bioinformatic process may be possible. We separately identified FC and BR amplicons before attempting to align and merge those with expected taxonomic identities. It may seem intuitively simpler to merge all detected FC and BR sequences into putative barcodes, and then to identify the correct barcode among those based on taxonomy. However, there were unexpectedly high numbers of FC and BR sequences for many specimens after filtering and denoising, which would result in exceedingly high numbers of pairwise combinations of sequences requiring examination for correct taxonomic identity.

The museum specimens included in this study spanned a diverse range of taxa and were collected from a wide range of locations. Neither of these factors affected our results, other than the limitations of the LCO1490 primer for certain insect families. Therefore, our approach could equally be applied to a set of specimens from one location to generate a reference data set of local relevance for use in biomonitoring. In principle, our approach may also be adaptable to different sets of amplicons, such as shorter fragments in order to amplify more degraded DNA from older specimens, or different taxonomic lineages. This would require adjustments to overlapping sequence alignment characteristics, but few other changes to the process.

5. CONCLUSION

As the need for DNA‐based biodiversity assessment continues to grow (D'Ercole et al., 2021; Gibson et al., 2015), the development of fit‐for‐purpose and reliable reference databases follows. Further, the development of methods for nondestructive DNA extraction from museum/stored samples has created an opportunity to develop rapid barcoding technology. Along with other recent attempts at using Illumina, PacBio and Oxford Nanopore technologies (D'Ercole et al., 2021; Shokralla et al., 2015), we provide another approach for rapid, sensitive and high‐throughput barcode generation. Our taxonomy‐informed pipeline utilizes the benefits and value of keeping specimens in taxonomic collections in the long‐term and helps develop targeted reference databases to support regional and national biodiversity surveys.

AUTHOR CONTRIBUTIONS

M.K.D., A.D. and T.R.B. conceived and designed the study, and M.K.D. and A.D. wrote the first draft of the manuscript with input from all the authors. T.B.C. and A.P. performed the laboratory experiments and A.D. performed the analysis. Taxonomic identification of voucher specimens was undertaken by RABL (Coleoptera), D.F.W. (Hymenoptera) and T.R.B. (Annelida).

CONFLICTS OF INTEREST

The authors declare no conflicts of interest.

BENEFIT SHARING STATEMENT

Benefits Generated: Benefits from this research accrue from the sharing of our data and results on public databases as described above. Manaaki Whenua Landcare Research continually engages with relevant iwi and tangata whenua representatives (indigenous Māori peoples of New Zealand) to ensure collection materials are sourced respectfully and with permission.

Supporting information

Table S1

MEN-25-e13733-s001.xlsx (163.9KB, xlsx)

ACKNOWLEDGEMENTS

This work is supported via the Strategic Science Investment Fund from the Ministry of Business Innovation and Employment and supported via the B3 (Better Border Biosecurity) Science Collaboration Project #D17.22. The authors would like to acknowledge the use of the New Zealand eScience Infrastructure (NeSI) for data analysis. Open access publishing facilitated by Landcare Research New Zealand, as part of the Wiley ‐ Landcare Research New Zealand agreement via the Council of Australian University Librarians.

Dopheide, A. , Brav‐Cubitt, T. , Podolyan, A. , Leschen, R. A. B. , Ward, D. , Buckley, T. R. , & Dhami, M. K. (2025). Fast‐tracking bespoke DNA reference database generation from museum collections for biomonitoring and conservation. Molecular Ecology Resources, 25, e13733. 10.1111/1755-0998.13733

Handling Editor: Catherine E Grueber

DATA AVAILABILITY STATEMENT

Raw sequence reads are available on Manaaki‐Whenua DataStore: https://doi.org/10.7931/m0pt‐np59. The bioinformatics pipeline is available on github: https://github.com/manaakiwhenua/Fast_DNA_barcoding

REFERENCES

  1. Batovska, J. , Piper, A. M. , Valenzuela, I. , Cunningham, J. P. , & Blacket, M. J. (2021). Developing a non‐destructive metabarcoding protocol for detection of pest insects in bulk trap catches. Scientific Reports, 11(1), 7946. 10.1038/s41598-021-85855-6 [DOI] [PMC free article] [PubMed] [Google Scholar]
  2. Bohmann, K. , Elbrecht, V. , Carøe, C. , Bista, I. , Leese, F. , Bunce, M. , Yu, D. W. , Seymour, M. , Dumbrell, A. J. , & Creer, S. (2022). Strategies for sample labelling and library preparation in DNA metabarcoding studies. Molecular Ecology Resources, 22(4), 1231–1246. 10.1111/1755-0998.13512 [DOI] [PMC free article] [PubMed] [Google Scholar]
  3. Bohmann, K. , Evans, A. , Gilbert, M. T. P. , Carvalho, G. R. , Creer, S. , Knapp, M. , Douglas, W. Y. , & De Bruyn, M. (2014). Environmental DNA for wildlife biology and biodiversity monitoring. Trends in Ecology & Evolution, 29(6), 358–367. [DOI] [PubMed] [Google Scholar]
  4. Boyer, S. , Brown, S. D. J. , Collins, R. A. , Cruickshank, R. H. , Lefort, M.‐C. , Malumbres‐Olarte, J. , & Wratten, S. D. (2012). Sliding window analyses for optimal selection of mini‐barcodes, and application to 454‐pyrosequencing for specimen identification from degraded DNA. PLoS One, 7(5), e38215. 10.1371/journal.pone.0038215 [DOI] [PMC free article] [PubMed] [Google Scholar]
  5. Carew, M. E. , Coleman, R. A. , & Hoffmann, A. A. (2018). Can non‐destructive DNA extraction of bulk invertebrate samples be used for metabarcoding? PeerJ, 6, e4980. 10.7717/peerj.4980 [DOI] [PMC free article] [PubMed] [Google Scholar]
  6. China Plant BOL Group , Li, D.‐Z. , Gao, L.‐M. , Li, H.‐T. , Wang, H. , Ge, X.‐J. , Liu, J.‐Q. , Chen, Z.‐D. , Zhou, S.‐L. , Chen, S.‐L. , Yang, J.‐B. , Fu, C.‐X. , Zeng, C.‐X. , Yan, H.‐F. , Zhu, Y.‐J. , Sun, Y.‐S. , Chen, S.‐Y. , Zhao, L. , Wang, K. , … Duan, G.‐W. (2011). Comparative analysis of a large dataset indicates that internal transcribed spacer (ITS) should be incorporated into the core barcode for seed plants. Proceedings of the National Academy of Sciences, 108(49), 19641–19646. 10.1073/pnas.1104551108 [DOI] [PMC free article] [PubMed] [Google Scholar]
  7. de Santana, C. D. , Parenti, L. R. , Dillman, C. B. , Coddington, J. A. , Bastos, D. A. , Baldwin, C. C. , Zuanon, J. , Torrente‐Vilara, G. , Covain, R. , Menezes, N. A. , Datovo, A. , Sado, T. , & Miya, M. (2021). The critical role of natural history museums in advancing eDNA for biodiversity studies: A case study with Amazonian fishes. Scientific Reports, 11(1), 18159. 10.1038/s41598-021-97128-3 [DOI] [PMC free article] [PubMed] [Google Scholar]
  8. Deiner, K. , Yamanaka, H. , & Bernatchez, L. (2021). The future of biodiversity monitoring and conservation utilizing environmental DNA. Environmental DNA, 3(1), 3–7. [Google Scholar]
  9. D'Ercole, J. , Prosser, S. W. J. , & Hebert, P. D. N. (2021). A SMRT approach for targeted amplicon sequencing of museum specimens (Lepidoptera)—Patterns of nucleotide misincorporation. PeerJ, 9, e10420. 10.7717/peerj.10420 [DOI] [PMC free article] [PubMed] [Google Scholar]
  10. Eisenhofer, R. , Minich, J. J. , Marotz, C. , Cooper, A. , Knight, R. , & Weyrich, L. S. (2019). Contamination in low microbial biomass microbiome studies: Issues and recommendations. Trends in Microbiology, 27(2), 105–117. 10.1016/j.tim.2018.11.003 [DOI] [PubMed] [Google Scholar]
  11. Elbrecht, V. , Braukmann, T. W. A. , Ivanova, N. V. , Prosser, S. W. J. , Hajibabaei, M. , Wright, M. , Zakharov, E. V. , Hebert, P. D. N. , & Steinke, D. (2019). Validation of COI metabarcoding primers for terrestrial arthropods. PeerJ, 7, e7745. 10.7717/peerj.7745 [DOI] [PMC free article] [PubMed] [Google Scholar]
  12. Folmer, O. , Black, M. , Hoeh, W. , Lutz, R. , & Vrijenhoek, R. (1994). DNA primers for amplification of mitochondrial cytochrome c oxidase subunit I from diverse metazoan invertebrates. Molecular Marine Biology and Biotechnology, 3(5), 294–299. [PubMed] [Google Scholar]
  13. Geller, J. , Meyer, C. , Parker, M. , & Hawk, H. (2013). Redesign of PCR primers for mitochondrial cytochrome c oxidase subunit I for marine invertebrates and application in all‐taxa biotic surveys. Molecular Ecology Resources, 13(5), 851–861. 10.1111/1755-0998.12138 [DOI] [PubMed] [Google Scholar]
  14. Gibson, J. F. , Shokralla, S. , Curry, C. , Baird, D. J. , Monk, W. A. , King, I. , & Hajibabaei, M. (2015). Large‐scale biomonitoring of remote and threatened ecosystems via high‐throughput sequencing. PLoS One, 10(10), e0138432. 10.1371/journal.pone.0138432 [DOI] [PMC free article] [PubMed] [Google Scholar]
  15. Gold, Z. , Curd, E. E. , Goodwin, K. D. , Choi, E. S. , Frable, B. W. , Thompson, A. R. , Walker, H. J., Jr. , Burton, R. S. , Kacev, D. , Martz, L. D. , & Barber, P. H. (2021). Improving metabarcoding taxonomic assignment: A case study of fishes in a large marine ecosystem. Molecular Ecology Resources, 21(7), 2546–2564. 10.1111/1755-0998.13450 [DOI] [PubMed] [Google Scholar]
  16. Hajibabaei, M. , Smith, M. A. , Janzen, D. H. , Rodriguez, J. J. , Whitfield, J. B. , & Hebert, P. D. N. (2006). A minimalist barcode can identify a specimen whose DNA is degraded. Molecular Ecology Notes, 6(4), 959–964. 10.1111/j.1471-8286.2006.01470.x [DOI] [Google Scholar]
  17. Hamady, M. , Walker, J. J. , Harris, J. K. , Gold, N. J. , & Knight, R. (2008). Error‐correcting barcoded primers for pyrosequencing hundreds of samples in multiplex. Nature Methods, 5(3), 235–237. 10.1038/nmeth.1184 [DOI] [PMC free article] [PubMed] [Google Scholar]
  18. Harper, L. R. , Buxton, A. S. , Rees, H. C. , Bruce, K. , Brys, R. , Halfmaerten, D. , Read, D. S. , Watson, H. V. , Sayer, C. D. , Jones, E. P. , Priestley, V. , Machler, E. , Murria, C. , Garces‐Pastor, S. , Medupin, C. , Burgess, K. , Benson, G. , Boonham, N. , Griffiths, R. A. , … Hanfling, B. (2019). Prospects and challenges of environmental DNA (eDNA) monitoring in freshwater ponds. Hydrobiologia, 826(1), 25–41. 10.1007/s10750-018-3750-5 [DOI] [Google Scholar]
  19. Hebert, P. D. N. , Braukmann, T. W. A. , Prosser, S. W. J. , Ratnasingham, S. , DeWaard, J. R. , Ivanova, N. V. , Janzen, D. H. , Hallwachs, W. , Naik, S. , Sones, J. E. , & Zakharov, E. V. (2018). A sequel to sanger: Amplicon sequencing that scales. BMC Genomics, 19, 219. 10.1186/s12864-018-4611-3 [DOI] [PMC free article] [PubMed] [Google Scholar]
  20. Hebert, P. D. N. , DeWaard, J. R. , Zakharov, E. V. , Prosser, S. W. J. , Sones, J. E. , McKeown, J. T. A. , Mantle, B. , & Salle, J. L. (2013). A DNA ‘barcode blitz’: Rapid digitization and sequencing of a natural history collection. PLoS One, 8(7), e68535. 10.1371/journal.pone.0068535 [DOI] [PMC free article] [PubMed] [Google Scholar]
  21. Katoh, K. , & Standley, D. M. (2013). MAFFT multiple sequence alignment software version 7: Improvements in performance and usability. Molecular Biology and Evolution, 30(4), 772–780. 10.1093/molbev/mst010 [DOI] [PMC free article] [PubMed] [Google Scholar]
  22. Krehenwinkel, H. , Pomerantz, A. , Henderson, J. B. , Kennedy, S. R. , Lim, J. Y. , Swamy, V. , Shoobridge, J. D. , Graham, N. , Patel, N. H. , Gillespie, R. G. , & Prost, S. (2019). Nanopore sequencing of long ribosomal DNA amplicons enables portable and simple biodiversity assessments with high phylogenetic resolution across broad taxonomic scale. GigaScience, 8(5), giz006. 10.1093/gigascience/giz006 [DOI] [PMC free article] [PubMed] [Google Scholar]
  23. Leggett, R. M. , & Clark, M. D. (2017). A world of opportunities with nanopore sequencing. Journal of Experimental Botany, 68(20), 5419–5429. 10.1093/jxb/erx289 [DOI] [PubMed] [Google Scholar]
  24. Lindahl, T. (1993). Instability and decay of the primary structure of DNA. Nature, 362(6422), 709–715. 10.1038/362709a0 [DOI] [PubMed] [Google Scholar]
  25. Lobo, J. , Costa, P. M. , Teixeira, M. A. L. , Ferreira, M. S. G. , Costa, M. H. , & Costa, F. O. (2013). Enhanced primers for amplification of DNA barcodes from a broad range of marine metazoans. BMC Ecology, 13(1), 34. 10.1186/1472-6785-13-34 [DOI] [PMC free article] [PubMed] [Google Scholar]
  26. Luo, A. , Zhang, A. , Ho, S. Y. , Xu, W. , Zhang, Y. , Shi, W. , Cameron, S. L. , & Zhu, C. (2011). Potential efficacy of mitochondrial genes for animal DNA barcoding: A case study using eutherian mammals. BMC Genomics, 12, 84. 10.1186/1471-2164-12-84 [DOI] [PMC free article] [PubMed] [Google Scholar]
  27. Marques, V. , Milhau, T. , Albouy, C. , Dejean, T. , Manel, S. , Mouillot, D. , & Juhel, J. B. (2021). GAPeDNA: Assessing and mapping global species gaps in genetic databases for eDNA metabarcoding. Diversity and Distributions., 27, 1880–1892. 10.1111/ddi.13142 [DOI] [Google Scholar]
  28. McGee, K. M. , Robinson, C. V. , & Hajibabaei, M. (2019). Gaps in DNA‐based biomonitoring across the globe [perspective]. Frontiers in Ecology and Evolution, 7, 337. 10.3389/fevo.2019.00337 [DOI] [Google Scholar]
  29. Minich, J. J. , Sanders, J. G. , Amir, A. , Humphrey, G. , Gilbert, J. A. , & Knight, R. (2019). Quantifying and understanding well‐to‐well contamination in microbiome research. mSystems, 4(4), e00186–e00119. 10.1128/mSystems.00186-19 [DOI] [PMC free article] [PubMed] [Google Scholar]
  30. Moinet, G. Y. K. , Dhami, M. K. , Hunt, J. E. , Podolyan, A. , Liáng, L. L. , Schipper, L. A. , Whitehead, D. , Nuñez, J. , Nascente, A. , & Millard, P. (2021). Soil microbial sensitivity to temperature remains unchanged despite community compositional shifts along geothermal gradients. Global Change Biology, 27, 6217–6231. 10.1111/gcb.15878 [DOI] [PMC free article] [PubMed] [Google Scholar]
  31. Paniagua Voirol, L. R. , Valsamakis, G. , Yu, M. , Johnston, P. R. , & Hilker, M. (2021). How the ‘kitome’ influences the characterization of bacterial communities in lepidopteran samples with low bacterial biomass. Journal of Applied Microbiology, 130(6), 1780–1793. 10.1111/jam.14919 [DOI] [PubMed] [Google Scholar]
  32. Pomerantz, A. , Peñafiel, N. , Arteaga, A. , Bustamante, L. , Pichardo, F. , Coloma, L. A. , Barrio‐Amorós, C. L. , Salazar‐Valenzuela, D. , & Prost, S. (2018). Real‐time DNA barcoding in a rainforest using nanopore sequencing: opportunities for rapid biodiversity assessments and local capacity building. GigaScience, 7(4), giy033. https://10.1093/gigascience/giy033 [DOI] [PMC free article] [PubMed] [Google Scholar]
  33. Potapov, V. , & Ong, J. L. (2017). Examining sources of error in PCR by single‐molecule sequencing. PLoS One, 12(1), e0169774. 10.1371/journal.pone.0169774 [DOI] [PMC free article] [PubMed] [Google Scholar]
  34. Price, M. N. , Dehal, P. S. , & Arkin, A. P. (2010). FastTree 2 – approximately maximum‐likelihood trees for large alignments. PLoS One, 5(3), e9490. 10.1371/journal.pone.0009490 [DOI] [PMC free article] [PubMed] [Google Scholar]
  35. Prosser, S. W. J. , DeWaard, J. R. , Miller, S. E. , & Hebert, P. D. N. (2016). DNA barcodes from century‐old type specimens using next‐generation sequencing. Molecular Ecology Resources, 16(2), 487–497. 10.1111/1755-0998.12474 [DOI] [PubMed] [Google Scholar]
  36. Ratnasingham, S. , & Hebert, P. D. N. (2007). BOLD: The barcode of life data system. Molecular Ecology Notes, 7, 355–364. 10.1111/j.1471-8286.2006.01678.x [DOI] [PMC free article] [PubMed] [Google Scholar]
  37. Rice, P. , Longden, I. , & Bleasby, A. (2000). EMBOSS: The European molecular biology open software suite. Trends in Genetics, 16(6), 276–277. [DOI] [PubMed] [Google Scholar]
  38. Rognes, T. , Flouri, T. , Nichols, B. , Quince, C. , & Mahé, F. (2016). VSEARCH: A versatile open source tool for metagenomics. PeerJ, 4, e2584. 10.7717/peerj.2584 [DOI] [PMC free article] [PubMed] [Google Scholar]
  39. Rowley, D. L. , Coddington, J. A. , Gates, M. W. , Norrbom, A. L. , Ochoa, R. A. , Vandenberg, N. J. , & Greenstone, M. H. (2007). Vouchering DNA‐barcoded specimens: test of a nondestructive extraction protocol for terrestrial arthropods. Molecular Ecology Notes, 7(6), 915–924. 10.1111/j.1471-8286.2007.01905.x [DOI] [Google Scholar]
  40. Schenekar, T. , Schletterer, M. , Lecaudey, L. A. , & Weiss, S. J. (2020). Reference databases, primer choice, and assay sensitivity for environmental metabarcoding: Lessons learnt from a re‐evaluation of an eDNA fish assessment in the Volga headwaters. River Research and Applications, 36(7), 1004–1013. 10.1002/rra.3610 [DOI] [Google Scholar]
  41. Schnell, I. B. , Bohmann, K. , & Gilbert, M. T. P. (2015). Tag jumps illuminated – Reducing sequence‐to‐sample misidentifications in metabarcoding studies. Molecular Ecology Resources, 15(6), 1289–1303. 10.1111/1755-0998.12402 [DOI] [PubMed] [Google Scholar]
  42. Shokralla, S. , Porter, T. M. , Gibson, J. F. , Dobosz, R. , Janzen, D. H. , Hallwachs, W. , Golding, G. B. , & Hajibabaei, M. (2015). Massively parallel multiplex DNA sequencing for specimen identification using an Illumina MiSeq platform. Scientific Reports, 5(1), 9687. 10.1038/srep09687 [DOI] [PMC free article] [PubMed] [Google Scholar]
  43. Shokralla, S. , Zhou, X. , Janzen, D. H. , Hallwachs, W. , Landry, J.‐F. , Jacobus, L. M. , & Hajibabaei, M. (2011). Pyrosequencing for mini‐barcoding of fresh and old museum specimens. PLoS One, 6(7), e21252. 10.1371/journal.pone.0021252 [DOI] [PMC free article] [PubMed] [Google Scholar]
  44. Sicard, M. , Bonneau, M. , & Weill, M. (2019). Wolbachia prevalence, diversity, and ability to induce cytoplasmic incompatibility in mosquitoes. Current Opinion in Insect Science, 34, 12–20. 10.1016/j.cois.2019.02.005 [DOI] [PubMed] [Google Scholar]
  45. Sint, D. , Raso, L. , & Traugott, M. (2012). Advances in multiplex PCR: Balancing primer efficiencies and improving detection success. Methods in Ecology and Evolution, 3(5), 898–905. 10.1111/j.2041-210x.2012.00215.x [DOI] [PMC free article] [PubMed] [Google Scholar]
  46. Sire, L. , Gey, D. , Debruyne, R. , Noblecourt, T. , Soldati, F. , Barnouin, T. , Parmain, G. , Bouget, C. , Lopez‐Vaamonde, C. , & Rougerie, R. (2019). The challenge of DNA barcoding Saproxylic beetles in natural history collections—Exploring the potential of parallel multiplex sequencing with Illumina MiSeq. Frontiers in Ecology and Evolution, 7, 495. 10.3389/fevo.2019.00495 [DOI] [Google Scholar]
  47. Song, H. , Buhay, J. E. , Whiting, M. F. , & Crandall, K. A. (2008). Many species in one: DNA barcoding overestimates the number of species when nuclear mitochondrial pseudogenes are coamplified. Proceedings of the National Academy of Sciences, 105(36), 13486–13491. 10.1073/pnas.0803076105 [DOI] [PMC free article] [PubMed] [Google Scholar]
  48. Srivathsan, A. , Baloğlu, B. , Wang, W. , Tan, W. X. , Bertrand, D. , Ng, A. H. Q. , Boey, E. J. H. , Koh, J. J. Y. , Nagarajan, N. , & Meier, R. (2018). A MinION™‐based pipeline for fast and cost‐effective DNA barcoding. Molecular Ecology Resources, 18(5), 1035–1049. 10.1111/1755-0998.12890 [DOI] [PubMed] [Google Scholar]
  49. Srivathsan, A. , Hartop, E. , Puniamoorthy, J. , Lee, W. T. , Kutty, S. N. , Kurina, O. , & Meier, R. (2019). Rapid, large‐scale species discovery in hyperdiverse taxa using 1D MinION sequencing. BMC Biology, 17(1). 10.1186/s12915-019-0706-9 [DOI] [PMC free article] [PubMed] [Google Scholar]
  50. Srivathsan, A. , Lee, L. , Katoh, K. , Hartop, E. , Kutty, S. N. , Wong, J. , Yeo, D. , & Meier, R. (2021). ONTbarcoder and MinION barcodes aid biodiversity discovery and identification by everyone, for everyone. BMC Biology, 19, 217. 10.1186/s12915-021-01141-x [DOI] [PMC free article] [PubMed] [Google Scholar]
  51. Sze, M. A. , & Schloss, P. D. (2019). The impact of DNA polymerase and number of rounds of amplification in PCR on 16 S rRNA gene sequence data. mSphere, 4(3), e00163–e00119. 10.1128/mSphere.00163-19 [DOI] [PMC free article] [PubMed] [Google Scholar]
  52. Tanabe, A. S. , & Toju, H. (2013). Two new computational methods for universal DNA barcoding: A benchmark using barcode sequences of bacteria, archaea, animals, fungi, and land plants. PLoS One, 8(10), e76910. 10.1371/journal.pone.0076910 [DOI] [PMC free article] [PubMed] [Google Scholar]
  53. Toju, H. , Vannette, R. L. , Gauthier, M. P. L. , Dhami, M. K. , & Fukami, T. (2018). Priority effects can persist across floral generations in nectar microbial metacommunities. Oikos, 127(3), 345–352. 10.1111/oik.04243 [DOI] [Google Scholar]
  54. Truett, G. E. , Heeger, P. , Mynatt, R. L. , Truett, A. A. , Walker, J. A. , & Warman, M. L. (2000). Preparation of PCR‐Quality Mouse Genomic DNA with Hot Sodium Hydroxide and Tris (HotSHOT). BioTechniques, 29(1), 52–54. 10.2144/00291bm09 [DOI] [PubMed] [Google Scholar]
  55. Wandeler, P. , Hoeck, P. E. , & Keller, L. F. (2007). Back to the future: Museum specimens in population genetics. Trends in Ecology & Evolution, 22(12), 634–642. 10.1016/j.tree.2007.08.017 [DOI] [PubMed] [Google Scholar]
  56. Winker, K. (2004). Natural history museums in a Postbiodiversity era. Bioscience, 54(5), 455. 10.1641/0006-3568(2004)054[0455:nhmiap]2.0.co;2 [DOI] [Google Scholar]
  57. Wong, W. H. , Tay, Y. C. , Puniamoorthy, J. , Balke, M. , Cranston, P. S. , & Meier, R. (2014). ‘Direct PCR’ optimization yields a rapid, cost‐effective, nondestructive and efficient method for obtaining DNA barcodes without DNA extraction. Molecular Ecology Resources, 14(6), 1271–1280. 10.1111/1755-0998.12275 [DOI] [PubMed] [Google Scholar]
  58. Yu, G. , Smith, D. K. , Zhu, H. , Guan, Y. , & Lam, T. T.‐Y. (2017). Ggtree: An r package for visualization and annotation of phylogenetic trees with their covariates and other associated data. Methods in Ecology and Evolution, 8(1), 28–36. 10.1111/2041-210X.12628 [DOI] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Table S1

MEN-25-e13733-s001.xlsx (163.9KB, xlsx)

Data Availability Statement

Raw sequence reads are available on Manaaki‐Whenua DataStore: https://doi.org/10.7931/m0pt‐np59. The bioinformatics pipeline is available on github: https://github.com/manaakiwhenua/Fast_DNA_barcoding


Articles from Molecular Ecology Resources are provided here courtesy of Wiley

RESOURCES