Skip to main content
Wiley Open Access Collection logoLink to Wiley Open Access Collection
. 2025 Sep 9;25(8):e70042. doi: 10.1111/1755-0998.70042

Robust, Open‐Source and Automation‐Friendly DNA Extraction Protocol for Hologenomic Research

Jonas G Lauritsen 1, Christian Carøe 1, Nanna Gaun 1, Garazi Martin‐Bideguren 1, Aoife Leonard 1,2, Raphael Eisenhofer 1, Iñaki Odriozola 1, M Thomas P Gilbert 1,3, Ostaizka Aizpurua 1, Antton Alberdi 1,, Carlotta Pietroni 1,
PMCID: PMC12550468  PMID: 40923293

ABSTRACT

Global efforts to standardise methodologies benefit greatly from open‐source procedures that enable the generation of comparable data. Here, we present a modular, high‐throughput nucleic acid extraction protocol standardised within the Earth Hologenome Initiative to generate both genomic and microbial metagenomic data from faecal samples of vertebrates. The procedure enables the purification of either RNA and DNA in separate fractions (DREX1) or as total nucleic acids (DREX2). We demonstrate their effectiveness across faecal samples from amphibians, reptiles and mammals, with reduced performance observed on bird guano. Despite some variation in laboratory performance metrics, both DREX1 and DREX2 yielded highly similar microbial community profiles, as well as comparable depth and breadth of host genome coverages. Benchmarking against a commercial kit widely used in microbiome research showed comparable recovery of host genomic data and microbial community complexity. Our open‐source method offers a robust, cost‐effective, scalable and automation‐friendly nucleic acid extraction procedure to generate high‐quality hologenomic data across vertebrate taxa. The method enhances research comparability and reproducibility by providing standardised, high‐throughput, open‐access protocols with fully transparent reagents. It is designed to integrate automatised pipelines, and its modular structure also supports continuous development and improvement.

Keywords: automation, isolation, laboratory protocol, metagenomics, nucleic acid

1. Introduction

Hologenomics, the joint analysis of animal genomic and associated microbial metagenomic data, is a systemic approach to studying host–microbiota interactions using DNA sequencing information (Nyholm et al. 2020; Theis 2018). This approach involves combining data on the population structure, evolutionary history and natural selection patterns retrieved from animal genomes with information on the community composition and functional traits of associated microbial communities (Alberdi et al. 2021).

The generation of paired animal genomic and microbial metagenomic sequencing data requires the extraction of DNA from their respective cells. This can be achieved through processing sample types enriched for different sources of DNA. For instance, blood or tissue samples can be used to obtain host DNA (Chacon‐Cortes et al. 2012), while intestinal contents are suitable for obtaining gut microbiome DNA (Tang et al. 2020). However, collecting such samples often involves invasive methods, which may not be feasible or ethical when dealing with wildlife (Banks and Piggott 2022; Zemanova 2020). Alternatively, faecal samples can serve as a non‐invasive source of both animal genomic and microbial metagenomic DNA (Srivathsan et al. 2016). Animal genotyping from faecal samples is a common practice in wildlife research (Idaghdour et al. 2003; Rodgers and Janečka 2013), while faecal samples are the indirect source for characterising intestinal microbiomes (Hernández et al. 2023; Thomas et al. 2015).

The Earth Hologenome Initiative (EHI, www.earthhologenome.org) was conceived as an international collaborative platform to generate and analyse hologenomic data collected worldwide from wild animals, following standardised methodologies (Leonard et al. 2024). The EHI's strategy for data generation relies on shotgun sequencing of faecal DNA, which enables the simultaneous generation of whole‐genome genotyping of host animals as well as the reconstruction of microbial genomes (Pietroni et al. 2025). Although numerous commercial kits are available for extracting DNA from faecal samples (Claassen et al. 2013), they may have limitations such as cost, accessibility and reproducibility. Additionally, comprehensive data regarding the chemical makeup of commercial kits is frequently unavailable due to the proprietary nature of the reagents or limited disclosure of information. This impedes the capacity to document scientific procedures, restricts the ability to replicate experiments, and hides critical differences and similarities between methodologies when comparing results across studies. Consequently, a primary objective for the EHI is to develop cost‐effective, open‐source procedures that allow the generation of hologenomic data without relying on specific commercial options.

Here, we present a custom‐made, open‐source, magnetic bead‐based DNA extraction procedure (DREX), benchmarked against the widely employed commercial alternative (hereafter REF). DREX can be used as a high‐throughput extraction protocol that separates RNA and DNA in its workflow (DREX1), or as a simplified protocol to recover the total nucleic acid content (DREX2). Both protocols selectively isolate nucleic acids by exploiting the affinity between the negatively charged DNA and the silica‐coated purification beads in the presence of chaotropic salts such as guanidinium thiocyanate (Boom et al. 1990). Dual RNA/DNA extraction is achieved with lysis/binding buffers containing guanidinium thiocyanate and citrate buffer (Casas et al. 1995). These components help preserve RNA by inhibiting RNase enzyme activity and preventing nucleic acid hydrolysis through pH stabilisation, respectively (Chheda et al. 2024; Weidner et al. 2022). While not yet benchmarked, DREX1 has already been used in multiple organisms, including fish (Bozzi et al. 2021), mammals (Koziol et al. 2023), and birds (Marcos et al. 2022). DREX2 was developed as part of the EHI, to streamline shotgun DNA sequencing data production from diverse taxa by decreasing processing time, maximising DNA yields from sub‐optimal sample types, and allowing the semi‐automated implementation of the protocol on a basic commercial pipetting and liquid handling robot.

We present the first comprehensive benchmarking of both DREX protocols, employing faecal samples from 12 species representing four classes of terrestrial vertebrates. We demonstrate the suitability of the two custom DREX protocols for hologenomic data generation from faecal samples, showing comparable performance and minimal variation compared to the commercial option.

2. Methods

2.1. Sample Origin and Handling

Following standardised EHI procedures, sampling tubes were prepared by dispensing 1 mL DNA Shield (Zymo, cat. number R1100‐250) to either a 2 mL Micro tube (Sarstedt, cat. Number 72.694.005) or a 2 mL Brand cryogenic tube (Sigma‐Aldrich, cat. Number BR114832), and labelling them with machine‐readable 2D barcodes using a Sci‐print VX2 (Scinomix). Tubes were arranged in sampling kits and shipped along with sampling guidelines to EHI participants for sample collection.

Two samples from 12 vertebrate species were collected following standardised procedures, namely the amphibians Calotriton asper , Lissotriton helveticus and Salamandra atra , the reptiles Chalcides striatus, Natrix astreptophora and Podarcis muralis , the birds Geospizopsis unicolor, Perisoreus infaustus and Zonotrichia capensis , and the mammals Plecotus auritus , Sciurus carolinensis and Trichosurus vulpecula . A maximum of 100 mg of freshly produced faecal material, corresponding to the amount required to maintain the manufacturer recommended 1:10 preservation ratio in DNA/RNA Shield (Zymo, cat. no. Biosite‐R1100‐250), was stored at −20°C until sample lysis and DNA extraction.

2.2. Sample Homogenisation and Lysis

Samples were mechanically lysed (Figure 1a) using Lysing Matrix E 96‐tube Rack (1.2 mL, MP Biomedicals, cat. number 116984010) before aliquoting for parallel extraction workflows. The lysis matrix was added to the sample tubes under sterile conditions and subsequently vortexed to assess the movement of the beads for each sample. Samples were bead‐beaten on a TissueLyser II (Qiagen) for two 6‐min intervals at 30 Hz, inverted between TissueLyser runs. The lysed materials were transferred to XL1000 tubes (LVL Technologies, cat. Number 2DNC‐X10‐BL‐NS‐SLC‐L).

FIGURE 1.

FIGURE 1

Overview of the main steps of the three tested DNA extraction protocols. (a) Common preparation steps that in the benchmarking experiment sourced the lysed extract that was split into the three purification protocols. (b) Protocol‐specific steps to isolate DNA. (c) Common final steps to purify and collect a clean DNA extract.

2.3. DNA Extraction

Homogenised lysates were diluted with DNA/RNA Shield, as needed, to reach a final volume of 600 μL—sufficient for comparison across the three extraction protocols and ensuring representative subsamples in each extraction. Equal aliquots of 200 μL were then dispensed into three deep‐well plates using a single randomised layout generated with Excel's RAND function, applied uniformly to every plate. Two negative extraction controls were included per treatment. The reference extraction (REF) was performed using the ZymoBIOMICS MagBead DNA/RNA kit (Zymo, cat. no. R2135), following the manufacturer's instructions for ‘Total Nucleic Acid Purification (III).’ Detailed protocols for DREX1 and DREX2 are provided in the Supporting Information. An overview of the steps involved in each protocol for obtaining clean extracts is illustrated in Figure 1b.

2.4. Library Preparation and Sequencing

The concentration of the DNA extracts was measured using the Qubit dsDNA Quantification Assay (ThermoFisher, cat. number Q32851) on a Qubit Flex Fluorometer (Thermofisher, cat. number Q33327). The ratios of the absorbance values of 260 nm vs. 280 nm (A260/A280) and the 260 nm vs. 230 nm (A260/A230) were determined using 1 μL with a spectrophotometer (DS11+, DeNovix Inc.) as indicators of sample purity. The integrity of the DNA starting material was assessed using a 2200 TapeStation system (Agilent) in conjunction with the Genomic DNA ScreenTape assay (Agilent, cat. Number 5067‐5365/5067‐5366).

Positions of samples and controls from all three extractions were randomised together, and afterwards transferred to a single 96 AFA‐TUBE TPX Plate (Covaris, cat. number 520291). The DNA content for library preparation was adjusted to a maximum of 200 ng per well, with a final total volume of 24 μL. Two negative library controls were further included for each treatment. The samples were fragmented for 103 s using a Covaris LE220 focused‐ultrasonication platform (Covaris, USA), aiming for fragments of approximately 350 bp.

DNA sequencing libraries were generated using the Blunt‐End Single Tube (BEST) protocol with BEDC3 adaptors (Carøe et al. 2017) and optimised conditions following (Mak et al. 2017). Adaptor molarities were tailored to the input DNA content: 20 μM for 50–200 ng, 10 μM for 10–50 ng, 5 μM for < 10 ng, and 2 μM for samples below the detection limit of the Qubit Assay.

Libraries were purified by adding 75 μL (around 1.7 times the library volume) of homemade Solid Phase Reversible Immobilisation (SPRI) beads prepared as described in (Rohland and Reich 2012). The mixture was incubated for 5 min at room temperature (ca. 21°C), followed by a double wash with 80% ethanol on a 96S Super Magnet (Alpaca, SKU: A001322). The purified DNA was eluted in Elution Buffer Tween (EBT; Buffer EB, Qiagen, cat. Number 19086, and TWEEN 20, Sigma‐Aldrich, cat. Number P9416‐50ML) solution by incubating at 37°C for 10 min and transferred to XSX200 tubes (LVL Technologies, cat. number 2DNC‐X02‐BL‐NS‐SLC‐L). The sequencing library total DNA amount (ng) was assessed using a Fragment Analyzer (Agilent) and HS NGS Fragment Kit (Agilent, cat. Number DNF‐474‐0500).

Library preparation performance was assessed using Cts (cycle threshold) by means of a qPCR assay conducted on a Mx3005p qPCR System (Agilent, USA) using 1:20 diluted libraries. qPCR mixes were composed of 2.5 μL 10× PCR Gold Taq buffer, 2.5 μL MgCl2 (25 mM), 0.2 μL dNTP mix (10 mM each), 1 μL (10 μM) for each IS7 (ACACTCTTTCCCTACACGAC) and IS8 (GTGACTGGAGTTCAGACGTGT) primers, 0.5 μL AmpliTaq GOLD DNA polymerase, 1 μL SYBR green dye, 14.3 μl sterile dH2O, and 2 μL diluted sample library for a total reaction volume of 25 μL. Amplification was carried out following this programme: (1) 1 cycle of 95°C for 12 min, (2) 40 cycles of (a) 95°C for 20 s, (b) 60°C for 30 s, and (c) 72°C for 40 s, (3) dissociation curve (55°C–95°C). To minimise the risk of clonality, qPCR amplification curves were analysed to estimate the optimal number of PCR amplification cycles for each library, determined as one cycle above the midpoint along the exponential phase curve. Subsequently, sequencing libraries were PCR‐amplified using unique dual‐index Illumina primers per sample. PCR reaction was 5 μL 10× PCR Gold Taq buffer, 5 μL MgCl2 (25 mM), 0.4 μL dNTP mix (10 mM each), 1 μL (10 μM) for each P7 and P5 primers, 1 μL AmpliTaq GOLD DNA polymerase, 26.6 μL sterile dH2O, and 10 μL sample library for a total reaction volume of 50 μL. The PCR reaction was (1) 1 cycle of 95°C for 12 min, (2) 7–20 cycles of (a) 95°C for 20 s, (b) 60°C for 30 s, and (c) 72°C for 40 s, (3) 1 cycle of 72°C for 5 min and (4) hold at 4°C. Each sample was amplified for a different number of cycles (7–20), as determined by the aforementioned qPCR assay.

The indexed libraries were purified using 1.2 times (60 μL) their volume of homemade SPRI beads (Rohland and Reich 2012). The mixture was incubated for 5 min at room temperature (ca. 21°C), followed by a double wash with 80% ethanol on a 96S Super Magnet (Alpaca, SKU: A001322). The purified DNA was eluted in EBT solution after an incubation at 37°C for 10 min and transferred to SXS200 tubes. The final molarity of the indexed libraries was assessed using HS NGS Fragment Kit (Agilent, cat. Number DNF‐474‐0500) on a Fragment Analyzer (Agilent). The purified indexed libraries were then pooled and sequenced on an Illumina NovaSeq X Plus platform (10B flow cell), employing 150 base pair (bp) read lengths and targeting for 5 GB of data per sample.

2.5. Bioinformatics Data Processing

Sequencing data were processed using the standardised genome‐resolved metagenomics pipeline of the EHI (Eisenhofer and Alberdi 2023). In short, the snakemake‐based pipeline first quality‐filters sequencing reads using fastp (Chen et al. 2018), before mapping them against the reference genome of the host, or the closest one available, using Bowtie2 (Langmead and Salzberg 2012). Estimations of host genome coverage were carried out with Picard (“Picard Toolkit” 2019). Unmapped reads were first screened for complexity using nonpareil3 (Rodriguez et al. 2018) and SingleM (Woodcroft et al. 2024). Subsequently, the reads were assembled using MEGAHIT (Li et al. 2015), and binned using an ensemble approach consisting of MaxBin2 (Wu et al. 2016), MetaBAT2 (Kang et al. 2019) and CONCOCT (Alneberg et al. 2014), followed by a bin refinement step based on MetaWRAP (Uritskiy et al. 2018). Bins were quality‐checked using CheckM2 (Chklovski et al. 2023), taxonomically annotated using GTDB‐Tk (Chaumeil et al. 2022), functionally annotated using DRAM (Shaffer et al. 2020) and dereplicated using dRep (Olm et al. 2017) at 98% average nucleotide identity to generate a metagenome‐assembled genome (MAG) catalogue. For quantification, quality‐filtered reads were mapped against the MAG catalogue using Bowtie2, and mapping statistics were retrieved using CoverM (Aroney et al. 2025). The pipeline outputs a count table consisting of the number of reads mapped to each MAG in each sample, a coverage table consisting of the proportion of each genome covered by at least one read in each sample, a genome metadata table containing quality and taxonomic information, a tree indicating the phylogenetic relationships among genomes and a functional annotation table containing information of all genes predicted across all genomes. To ensure reliable genome detection, a minimum threshold of 30% genome coverage was applied (Acheampong et al. 2024; Hakim et al. 2022); i.e., a genome was considered present only if at least 30% of its sequence was covered by reads, thereby minimising false positives arising from cross‐mapping of conserved regions.

2.6. Data Analysis

To evaluate the performance of the three examined DNA extraction methods, we analysed various parameters concerning laboratory procedures, data integrity, genomic complexity, and microbial community analytics. Three metrics were employed to gauge the proficiency of laboratory steps: DNA extraction yield (ng DNA), sample purity based on ratios of absorbance (A260/A280 and A260/A230), and library performance, as determined by qPCR analysis (cycle threshold, Cts). Nucleic acid sample purity is generally considered acceptable for most downstream applications when A260/A280 ratio and A260/A230 ratio > 1.8. Therefore, sample purity was tested using the absolute deviation from the minimal threshold value of 1.8. Data integrity was evaluated by computing the percentage of sequencing reads discarded due to low quality. Genomic complexity was assessed by analysing the sequence duplication rate and the breadth of coverage across the reference genome. Metagenomic complexity was appraised using Nonpareil completeness, while microbial community composition was examined through Hill numbers‐based alpha and beta diversity metrics, encompassing multiple components of diversity (Alberdi and Gilbert 2019; Chao et al. 2016; Hill 1973; Jost 2007). Birds were excluded from community composition analysis, as most samples failed to yield a meaningful microbial community across any of the extraction methods tested.

For continuous response variables, we used linear mixed effect models (LMM) with gaussian distribution through the lmer function from lmerTest package (Kuznetsova et al. 2017). Response variables were log‐transformed or rank‐transformed when necessary to meet the assumptions of normal distribution and homoscedasticity of residuals. For proportional response variables, we used generalised linear mixed modelling (GLMM) with a quasibinomial distribution (Papke and Wooldridge 1996) with the glmmPQL function from the MASS package (James et al. 1996), using an approximate Wald F‐test where the F‐statistics were derived by dividing the Wald chi‐squared values by the factor degrees of freedom.

Comparisons across extraction methods and taxonomic groups were conducted using the extraction method (categorical factor with three levels: REF, DREX1, DREX2) and taxonomic group (categorical factor with four levels: amphibians, reptiles, mammals, birds) as fixed effects. The total DNA amount (log10[ng]) of the library was included as a fixed effect in library performance comparisons. Since three species represented each taxonomic group and two samples were taken from each species, we included species_id and sample_id as nested random intercepts. With this data, testing for the interactions between extraction methods and taxonomic groups and species identities was theoretically possible, which would have provided information of, e.g., whether the adequacy of extraction protocols differed between taxonomic groups or species. However, we refrained from doing it, given the small size of the experiment and the associated risk of overfitting and overinterpreting the results. Pairwise community dissimilarities were calculated using the Jaccard‐type turnover metric (S) derived from the Hill numbers beta diversity (Alberdi and Gilbert 2019; Chao et al. 2016; Hill 1973; Jost 2007). To evaluate the influence of experimental factors on microbial community composition, we performed permutational multivariate analysis of variance (PERMANOVA) using the adonis2 function from the vegan package (Oksanen et al. 2013), and the importance of the variables was assessed by the R 2 associated with each term (Anderson 2001). Non‐metric multidimensional scaling (NMDS) was employed to visualise and interpret patterns in microbial community composition. All analyses were conducted in the R environment, and data analysis code that enables full reproduction of the above described procedures is available on GitHub (https://github.com/earthhologenome/EHI_extraction_test).

3. Results

3.1. Laboratory Performance

DREX2 outperformed the other two extraction protocols (LMM: F (2,46) = 8.29, p < 0.001) yielding average DNA extraction amounts (Figure 2a) that were considerably higher (446.34 ng [118.77, 1677.29]) as compared with REF (205.25 ng [54.62, 771.31]) and DREX1 (160.39 ng [42.68, 602.74]). Birds tended to yield lower amounts of DNA, followed by reptiles, but the evidence for a significant difference across taxonomic groups was weak (LMM: F (3,8) = 3.94, p = 0.054) (Figure 2a).

FIGURE 2.

FIGURE 2

Laboratory and sequencing performance metrics of the three extraction protocols. (a) DNA extraction performance in terms of yield. (b) Library preparation performance based on qPCR Ct (cycle threshold) values. (c) Sequencing performance in terms of percentage of nucleic acid bases retained after quality filtering and adaptor trimming. (d) Data complexity performance in terms of percentage of duplicated reads.

Most of the samples fell outside the Tapestation's specified functional range (5–300 ng/μL) required to provide a DNA Integrity Number (DIN) for numerical assessment. However, overlaid electropherograms for samples within the functional range across all protocols showed comparable DNA integrity between the extracts (Figure S1).

There was a significant difference in sample purity among the three extraction methods, as measured by the absolute deviation from the expected value of 1.8 of the A260/A280 (LMM: F (2,58) = 3.64, p = 0.032) and A260/A230 (LMM: F (2,40.45) = 5.44, p = 0.008) ratios, while no significant difference was detected across taxonomic groups. Across all protocols, the A260/A280 and especially the A260/A230 ratios were appreciably lower than the recommended threshold of 1.8 (Figure S2a,b) suggesting a significant level of contamination especially from compounds that absorb at 230 nm, which include chaotropic salts such as guanidine thiocyanate (GTC), EDTA, non‐ionic detergents like Tween 20 and proteins. Higher A260/A280 ratios (1.63 ± 0.35) were detected for DREX1 extracts compared to DREX2 (1.44 ± 0.31) and REF (1.35 ± 0.18). In contrast, A260/A230 ratios were highest in DREX2 extracts (0.20 ± 0.27), followed by DREX1 (0.17 ± 0.13) and REF (0.09 ± 0.11). However, the validity of ratio measurements for certain samples may be compromised due to a concentration falling below the manufacturer's minimum recommended level. Amplification performance (Figure 2b) was comparable among the three extraction methods, as measured by the Ct values (LMM: F (2,45.42) = 2.28, p = 0.114) normalised to the amount of DNA input. Although library preparation performance was similar across amphibians, mammals, and reptiles, avian samples performed significantly worse, with cycle threshold values closer to the controls (Figure 2b).

3.2. Sequencing Data Performance

A total amount of 429.1 GB (gigabases) of data from 1.43 billion sequencing reads was generated, yielding an average of 5.36 ± 2.7 GB (17.87 million reads) per sample. Despite the slight variation in sequencing depth per sample, the sequencing effort was balanced for all taxonomic groups and extraction methods (Figure S2c). The fraction of high‐quality reads was significantly lower in birds than that of other taxa (GLMMquasibinomial: F (3,64) = 14.4, p < 0.001) (Figure 2c), but there was no significant difference in the overall proportion of quality‐filtered reads across extraction methods (GLMMquasibinomial: F (2,64) = 0.13, p = 0.875), indicating that sequencing quality was not affected by the extraction method.

3.3. Host Genome Performance

We tested multiple parameters to assess performance regarding host genomes. REF tended to have higher average duplication rates than the other methods, but the evidence for a difference was weak (GLMMquasibinomial: F (2,64) = 2.59, p = 0.083). On the other hand, duplication rates were highest in birds, with significantly lower values in amphibians, mammals and reptiles (GLMMquasibinomial: F (3,64) = 6.29, p < 0.001) (Figure 2d). Depth of coverage was similar across extraction methods (LMM: F (2,46) = 0.02, p = 0.984), while DREX1 and, particularly, DREX2 exhibited a significant improvement over REF (GLMMquasibinomial: F (2,64) = 10, p < 0.001) in terms of breadth of coverage. While birds exhibited the highest depth of coverage (LMM: F (3,8) = 6.82, p = 0.014), their breadth of coverage was not as high (GLMMquasibinomial: F (3,64) = 2.73, p = 0.051) due to the high duplication rates (Figure 3a).

FIGURE 3.

FIGURE 3

Overview of host genomic and microbial metagenomic features derived from the three extraction methods. (a) Relationship between log‐transformed breadth and depth of coverage of host genomes. (b) Richness, neutral (Shannon) diversity and phylogenetic diversity metrics yielded by the three extraction methods. Birds are omitted because the majority of samples failed to recover a meaningful microbial community after filtering and normalisation of the raw data. (c) Microbial community composition recovered by the three extraction methods. Each consecutive trio represents a biological sample processed with REF, DREX1 and DREX2 methods, respectively. Note the barplots are generated after mapping all reads against a common MAG catalogue derived from the dereplication of all MAGs produced within each species, without applying a minimum genome coverage filter of 30% to avoid birds losing most of their reads. (d) Ordination of the community composition differences measured based on neutral beta diversity, with communities derived from the biological sample linked with connection lines. (e) The proportion of compositional variance explained by the DNA extraction method, taxonomic group, species identity and samples in terms of richness, neutral diversity and phylogenetic diversity (as measured by the R 2 values of a PERMANOVA).

3.4. Microbial Community Performance

The complexity of metagenomic data after removing host DNA was measured by Nonpareil completeness and was significantly different among extraction methods (GLMMquasibinomial: F (2,64) = 7.11, p = 0.002), with DREX2 obtaining lower (indicating higher complexity) values as compared with DREX1 and REF methods (Figure S2d). Given the identical biological sample and the comparable sequencing depth employed, the higher complexity was interpreted as an improved capacity to recover microbial diversity. In contrast, the estimation of the recovered microbial fraction did not differ significantly across extraction methods (GLMMquasibinomial: F (2,64) = 2.3, p = 0.08) (Figure S2e). All different metrics (richness, neutral [Shannon], and phylogenetic diversity) of microbial alpha diversity were nearly identical between extraction methods (LMM: p > 0.05) (Figure 3b), whereas metrics of beta diversity significantly differed between them (PERMANOVA, permutations = 999, p < 0.001) (Figure 3c,d). A visual evaluation of the compositional differences between samples through NMDS showed that the communities generated by DREX1 and DREX2 were more similar than the one generated by REF for most of the samples (Figure 3d). However, the variance explained by the extraction method was small for the three metrics of beta diversity, with R 2 values between 0.02 and 0.03, while between sample technical variation obtained R 2 values of 0.14–0.2 (Figure 3e). As expected, taxonomic group and species identity explained most of the variance in microbial community composition (Figure 3e).

3.5. Estimation of Cost, Time and Environmental Impact

For an extraction batch of 96 samples, we compared the three protocols based on (1) cost based on the Danish market, (2) processing time and (3) environmental impact, assessed through plastic consumption of consumables (and their immediate packaging) and long‐term storage used within the laboratory confines following the extraction methods. DREX2 proved to be the most cost‐effective option for DNA isolation, while the slightly higher cost of DREX1 was offset by its enhanced performance, enabling the separate isolation of RNA. Both protocols were significantly more affordable than REF, with cost reductions ranging from 72% to 75% (Table 1). Including the time required for buffer preparation, we estimated that DREX2 and REF had comparable processing times (4.5 h), whereas DREX1 required an additional 3.5 h (55.5% longer; Table 1). However, this extended time accounts for the additional separation and purification of RNA. In terms of plastic consumption, DREX2 had the lowest plastic footprint, generating 1710 g, while DREX1 generated 4193.95 g and REF used 2653.98 g (Table 1). It is important to note that our evaluation only considered the plastic discarded during the extraction process, as there is currently no system in place to track plastic consumption related to the preparation, storage, or shipping of purchased reagents.

TABLE 1.

Cost‐efficiency, processing times and environmental impact estimates for the three extraction methods per 96‐sample batch. Costs were calculated using stock prices (excluding VAT) in Denmark as of June 2024. The estimates encompass buffer preparation, setup and extractions; for DREX1, both RNA and DNA extractions are included. The plastic items are a rounded measure based on packaging, storage, and themselves.

Method 1. Cost efficiency 2. Processing time 3. Environmental impact
Cost Difference (%) Time (h) Difference (%) Plastic usage (g) Difference (%)
REF 910.08 € 4.5 2653.98
DREX1 251.52 € −72.30 7.0 +55.50 4193.95 +58.00
DREX2 227.52 € −75 4.5 0 1710 −35.60

4. Discussion

The adoption of standard operating procedures (SOPs) is a crucial initial step for any research endeavour involving multiple parties. SOPs are particularly pertinent in microbiome research because metagenomic data generation is significantly influenced by both field and laboratory methodologies (Aizpurua et al. 2023). Unlike genomic analyses, which focus on the DNA of a single organism, metagenomic analyses aim to quantify DNA from various organisms within a complex mixture (Thomas et al. 2012). These organisms often exhibit different structural properties that affect the ease of DNA extraction. For example, Gram‐positive bacteria have thick peptidoglycan cell walls that are more difficult to lyse compared to Gram‐negative bacteria, potentially introducing biases in the recovered signal (Frostegård et al. 1999).

Many companies have developed DNA extraction kits designed to minimise biases and maximise DNA recovery (Claassen et al. 2013). However, we argue that standardising a commercial kit as the procedure for generating data in an international research project raises ethical and practical concerns regarding openness and reproducibility. The composition of many reagents in commercial kits is not disclosed, making procedures reproducible only through the use of the exact same kit. Furthermore, commercial limitations, production issues, or the discontinuance of a product can hinder the continuity of standardised research procedures. Consequently, as part of the EHI, we adapted a well‐established chemistry into a high‐throughput, custom open‐source procedure, which is benchmarked in this study.

4.1. Robust Performance of Open‐Source Procedures

Our custom open‐source DNA extraction protocol DREX1 yielded comparable DNA amounts to the commercial option REF. In contrast, DREX2 yielded significantly higher DNA quantities compared to the other two protocols. DREX2 likely outperformed DREX1 and REF due to its streamlined procedure, which involved fewer binding and washing steps, known to inevitably cause some degree of DNA loss during bead‐based purifications. Moreover, the absence of a binding incubation step in DREX1 during the purification of the DNA fraction may contribute, to a lesser extent, to the reduced DNA binding efficiency to the beads compared to DREX2. DNA yields were satisfactory for downstream analyses in all taxa analysed except birds, confirming the challenging nature of avian faecal DNA (Eriksson et al. 2017).

However, higher DNA yields do not ensure successful downstream processes if they come at the expense of other quality estimates such as extract purity (Schiebelhut et al. 2017) and integrity (Jansson et al. 2024). This occurs because inhibitory compounds, such as humic acids and acidic polysaccharides, are not successfully removed during DNA isolation (Monteiro et al. 1997; Opel et al. 2010), which can inhibit enzymatic reactions during library preparation (Murray et al. 2015; Schrader et al. 2012). Moreover, although short‐read sequencing requires DNA to be fragmented to specific sizes, the degree of the fragmentation of the initial DNA input can result in uneven insert size distributions and inefficient library construction, particularly if post‐fragmentation size selection is not performed before library preparation. The choice of specific DNA extraction protocols can exacerbate DNA fragmentation issues related to sample types and preservation, especially in specific applications such as genomics (Jansson et al. 2024). In metagenomics, the observed community composition may be skewed when protocols biased towards recovering low‐yield or highly fragmented DNA are applied (Costea et al. 2017). The manufacturer describes the REF kit as a reliable bead‐based method for purifying high‐quality, inhibitor‐free DNA. In our study, the DREX protocols produced extracts with comparable, and in some cases higher, purity values. However, as some sample concentrations fell below the optimal threshold for accurate purity assessment, the validity of this comparison may be affected. Our evaluation identified considerable contaminants and protein contamination in all DNA extracts, which is expected given the high presence of inhibitors and interfering substances in faeces. We found no evidence to suggest that any of the protocols are better suited to enrich high‐molecular‐weight DNA.

While purity ratios and integrity can serve as important indicators, the most reliable measure of DNA quality is its performance in the intended downstream application. Faecal samples are known to contain various inhibitors that can interfere with downstream applications, and nucleic acid isolation procedures that deal with the problem of coextraction of inhibitors have been proposed (Eggert et al. 2005; McKee et al. 2015). We did not detect any significant difference in amplification performance between methods when the amount of DNA input was taken into account. This suggests that the primary factor determining amplification performance and efficiency of library preparation was simply the yield of DNA available, rather than extraction‐related factors such as lower DNA quality due to the presence of contaminants, likely due to the enzyme/reagent robustness. Furthermore, these results may partly be attributed to the limited amount of stool used as starting material (≤ 100 mg) for nucleic acid isolation, which could have mitigated the concentration of inhibitors typically present in faecal samples, thus minimising their negative impact on the library preparation process regardless of the extraction method utilised (McKee et al. 2015).

Additionally, sequencing performance, assessed by the proportion of high‐quality reads, was consistent across methods, reinforcing the conclusion that library preparation efficiency was comparable between protocols. Similarly, different microbial alpha diversity metrics, which serve as reliable indicators of overall protocol performance and accuracy of the recovered microbial profile, were nearly identical across protocols. Metagenomic completeness was significantly lower in DREX2 extracts, suggesting higher metagenomic complexity in these extracts. This implies that the method is capable of producing data of the same high quality as DREX1 and REF, while potentially offering a more comprehensive representation of the microbial content at comparable sequencing effort. Microbial community compositions reconstructed from the DNA extracts displayed some variation across the tested DNA extraction protocols. However, when the extraction procedure effect size (Costea et al. 2017) was compared to other biological (taxonomic group and species) and other technical effects, the impact was minimal, accounting for only 2%–3% of the variance in microbial community compositions. Most of the variance was attributed to the taxonomic group and species identity, with the remaining between‐sample technical variation explained by the individual or other unmeasured factors that were not independently assessed in the study.

The three DNA extraction methods recovered microbial communities consistent with those reported in previous studies across amphibians, reptiles and mammals (Ley et al. 2008; Song et al. 2020). In contrast, bird DNA extractions performed worse than the rest of the taxa based on several performance metrics, yielding significantly lower quality sequencing reads, higher duplication rates, and higher estimates of metagenomic completeness, which indicates lower complexity. In the case of birds, REF recovered slightly more complex metagenomes, but most detected taxa were removed when applying the filtering of a minimum genome coverage of 30%. The consistently lower performance in avian samples is expected and likely reflects the biochemical properties of cloacal excreta, where the admixture of faeces and urine both dilutes bacterial concentrations and, due to high uric acid levels, renders DNA extraction and downstream enzymatic processes particularly challenging (Eriksson et al. 2017; Russell et al. 2024). Future work could quantify the relative contributions of bacterial biomass and uric acid‐mediated inhibition to the observed lower performance, thereby guiding the targeted optimisation of the DNA extraction protocols for avian faecal samples, with a focus on either minimising biomass loss or enhancing inhibitor removal through improved lysis or buffer formulations.

Host genome depth of coverage was similar across methods. However, DREX1—and especially DREX2—showed improvements in breadth of coverage compared to REF. These enhancements were likely due to greater complexity in the extracted genomic DNA, consistent with the higher complexity observed in metagenomic DNA recovered through DREX2, as well as the slightly different duplication rates measured across methods. However, regardless of the protocol used, the depth and breadth of coverage metrics were, in general, low due to a variety of reasons. The lack of species‐specific reference genomes precluded amphibian reads from aligning against the host genome. Although reptiles exhibited slightly larger depth of genomic coverage, the small fraction of host DNA in faeces prevented the recovery of a larger fraction of host signal. The larger overall depth of coverage in mammals was driven by the bat Plecotus auritus , whose faeces contained a significantly larger amount of host DNA compared to the other species.

4.2. Cost‐Effectiveness of Custom‐Made Procedures

An advantage of commercial kits is that they typically include most of the elements required, such as reagents and plasticware, sourced from a single provider. These kits are usually ready to use and undergo quality‐control tests by the manufacturers. In contrast, custom‐made procedures often involve sourcing materials from multiple providers, preparing working reagents in‐house, and implementing in‐house quality‐testing procedures. However, this extra work is offset by significantly reduced costs and/or the items required can be purchased from a variety of vendors, making custom‐made methods far more universally accessible. According to our calculations for a sample batch of 96 samples, both DREX protocols are 72%–75% cheaper than the commercial solution, allowing for a fourfold increase in sample size at the same cost. Moreover, the streamlined version DREX2 offers the same processing time as REF, but with significantly reduced plastic consumption, resulting in a lower environmental impact.

Consequently, while commercial kits might be the best option for small batches, the benefits of using custom‐made DNA extraction procedures are greatly amplified in large‐scale projects. Additionally, custom‐made open‐source procedures provide detailed information about the exact formulations of all reagents used, which is usually unavailable in commercial kits. This transparency enables researchers to optimise procedures for different samples or adjust them to different instruments. Transparent research practices are crucial for advancing knowledge and ensuring credibility in research, as they enable independent laboratories to reproduce results, verify findings, compare across studies, and detect potential errors or biases.

4.3. Conclusions

The benchmarking of our custom‐made, open‐source DNA extraction protocol, DREX, confirms its effectiveness for hologenomic data generation across diverse faecal sample types. The minimal variance in microbiome composition attributable to DNA extraction methods highlights the robustness of DREX and supports its interchangeable use without significantly influencing downstream results. However, further optimisation is warranted to enhance its performance on bird guano samples.

Conflicts of Interest

The authors declare no conflicts of interest.

Supporting information

Data S1: men70042‐sup‐0001‐DataS1.pdf.

MEN-25-e70042-s003.pdf (32.3KB, pdf)

Data S2: men70042‐sup‐0002‐DataS2.pdf.

MEN-25-e70042-s002.pdf (31.5KB, pdf)

Figure S1: Integrity of the DNA extracted using the three methods in four randomly selected samples. The first peak corresponds with the marker, while the second peak indicates the proportion of DNA fragments with distinct sizes.

Figure S2: Laboratory and sequencing performance metrics of the three extraction protocols. (A) 260/280 nm absorbance ratio indicating the purity of the DNA extraction. (B) 260/230 nm absorbance ratio indicating the purity of the DNA extraction. (C) Sequencing depth in gigabases (thousands of millions of nucleic acid bases). (D) Kmer‐based estimation of the recovered metagenomic complexity. (E) Proportion of the estimated microbial community recovered in each sample.

MEN-25-e70042-s001.docx (937.8KB, docx)

Acknowledgements

We would like to thank the EHI participants Carlos Cabido, Joana Fernandes, Peter Hosner, Emina Šunje, Alex Sutton, and Luc Wauters for kindly providing the samples they collected for this study. We would also like to extend a special gratitude to the EHI manager, Ella Lattenkamp, for the organisation of the research. This study was funded by the Carlsberg foundation's grant CF20‐0460 awarded to A.A., as well as the Danish National Research Foundation grant DNRF143 and the Danish Dairy Association and the Danish Agriculture and Food Sector Dairy award ‘Metacheese’ awarded to M.T.P.G.

Handling Editor: Kin‐Ming (Clement) Tsui

Funding: This work was supported by Carlsbergfondet (CF20‐0460), Danmarks Grundforskningsfond (DNRF143), The Danish Dairy Association and the Danish Agriculture and Food Sector Dairy, Metacheese.

Contributor Information

Antton Alberdi, Email: antton.alberdi@sund.ku.dk.

Carlotta Pietroni, Email: carlotta.pietroni@sund.ku.dk.

Data Availability Statement

The raw data tables containing the quantitative information, data accession codes, and bioinformatic code scripts used for the analysis are available in a dedicated GitHub repository (https://github.com/earthhologenome/EHI_extraction_test), both rendered into a HTML webbook (https://www.earthhologenome.org/EHI_extraction_test) and frozen in Zenodo with doi:10.5281/zenodo.16757704. Raw sequencing data and microbial genome sequences, along with their respective sampling, export, ABS and import permits, were published as part of the 1st EHI Data Release (Gaun et al. 2025), under Bioproject PRJEB76898. The samples used for the validation of the extraction methodologies were procured following the standardised administrative procedures of the Earth Hologenome Initiative, which has strict rules regarding local and international Access and Benefit Sharing regulations, including the Nagoya Protocol. The relevant authorisations for this work include permits 168E/2022 and 254E/2023 by the regional Government of Navarra; permit 2023FAUNA0000010501 by the regional Government of Gipuzkoa; permit 04–23‐550/16 ZM by the Federal Ministry of Environment and Tourism of Serbia; permit MAAE‐ARSFC‐2022‐2083 by the Ministry of Environment, Water and Ecological Transition of Ecuador; permit 23‐20 by the Agricultural Department of Sweden; permit DGMNB/SEN/avp_21_187 by the regional Government of Castilla‐La Mancha; permit DD‐1015 by the Region of Lombardia; permit DBCA‐AEC 2021‐01A by the Australian Wildlife Conservancy; and permit 648/2023/CAPT by the Institute of Nature Conservation of Portugal. The benefits of this research include sharing our code, data, and results on public databases as outlined above.

References

  1. Acheampong, D. A. , Jenjaroenpun P., Wongsurawat T., et al. 2024. “CAIM: Coverage‐Based Analysis for Identification of Microbiome.” Briefings in Bioinformatics 25, no. 5: bbae424. [DOI] [PMC free article] [PubMed] [Google Scholar]
  2. Aizpurua, O. , Dunn R. R., Hansen L. H., Gilbert M. T. P., and Alberdi A.. 2023. “Field and Laboratory Guidelines for Reliable Bioinformatic and Statistical Analysis of Bacterial Shotgun Metagenomic Data.” Critical Reviews in Biotechnology 44: 1–1182. [DOI] [PubMed] [Google Scholar]
  3. Alberdi, A. , Andersen S. B., Limborg M. T., Dunn R., and Gilbert M. T. P.. 2021. “Disentangling Host‐Microbiota Complexity Through Hologenomics.” Nature Reviews. Genetics 23: 281–297. [DOI] [PubMed] [Google Scholar]
  4. Alberdi, A. , and Gilbert M. T. P.. 2019. “A Guide to the Application of Hill Numbers to DNA‐Based Diversity Analyses.” Molecular Ecology Resources 19, no. 4: 804–817. [DOI] [PubMed] [Google Scholar]
  5. Alneberg, J. , Bjarnason B. S., de Bruijn I., et al. 2014. “Binning Metagenomic Contigs by Coverage and Composition.” Nature Methods 11, no. 11: 1144–1146. [DOI] [PubMed] [Google Scholar]
  6. Anderson, M. J. 2001. “A New Method for Non‐Parametric Multivariate Analysis of Variance.” Austral Ecology 26, no. 1: 32–46. [Google Scholar]
  7. Aroney, S. T. N. , Newell R. J. P., Nissen J. N., Camargo A. P., Tyson G. W., and Woodcroft B. J.. 2025. “CoverM: Read Alignment Statistics for Metagenomics.” Bioinformatics 41: btaf147 arXiv. http://arxiv.org/abs/2501.11217. [DOI] [PMC free article] [PubMed] [Google Scholar]
  8. Banks, S. C. , and Piggott M. P.. 2022. “Non‐Invasive Genetic Sampling Is One of Our Most Powerful and Ethical Tools for Threatened Species Population Monitoring: A Reply to Lavery et al.” Biodiversity and Conservation 31, no. 2: 723–728. [Google Scholar]
  9. Boom, R. , Sol C. J., Salimans M. M., Jansen C. L., Wertheim‐van Dillen P. M., and van der Noordaa J.. 1990. “Rapid and Simple Method for Purification of Nucleic Acids.” Journal of Clinical Microbiology 28, no. 3: 495–503. [DOI] [PMC free article] [PubMed] [Google Scholar]
  10. Bozzi, D. , Rasmussen J. A., Carøe C., et al. 2021. “Salmon Gut Microbiota Correlates With Disease Infection Status: Potential for Monitoring Health in Farmed Animals.” Animal Microbiome 3, no. 1: 30. [DOI] [PMC free article] [PubMed] [Google Scholar]
  11. Carøe, C. , Gopalakrishnan S., Vinner L., et al. 2017. “Single‐Tube Library Preparation for Degraded DNA.” Methods in Ecology and Evolution 9: 410–419. 10.1111/2041-210X.12871. [DOI] [Google Scholar]
  12. Casas, I. , Powell L., Klapper P. E., and Cleator G. M.. 1995. “New Method for the Extraction of Viral RNA and DNA From Cerebrospinal Fluid for Use in the Polymerase Chain Reaction Assay.” Journal of Virological Methods 53, no. 1: 25–36. [DOI] [PubMed] [Google Scholar]
  13. Chacon‐Cortes, D. , Haupt L. M., Lea R. A., and Griffiths L. R.. 2012. “Comparison of Genomic DNA Extraction Techniques From Whole Blood Samples: A Time, Cost and Quality Evaluation Study.” Molecular Biology Reports 39, no. 5: 5961–5966. [DOI] [PubMed] [Google Scholar]
  14. Chao, A. , Chiu C.‐H., and Jost L.. 2016. “Phylogenetic Diversity Measures and Their Decomposition: A Framework Based on Hill Numbers.” In Biodiversity Conservation and Phylogenetic Systematics: Preserving Our Evolutionary Heritage in an Extinction Crisis, edited by Pellens R. and Grandcolas P., 141–172. Springer International Publishing. [Google Scholar]
  15. Chaumeil, P.‐A. , Mussig A. J., Hugenholtz P., and Parks D. H.. 2022. “GTDB‐Tk v2: Memory Friendly Classification With the Genome Taxonomy Database.” Bioinformatics 38, no. 23: 5315–5316. [DOI] [PMC free article] [PubMed] [Google Scholar]
  16. Chen, S. , Zhou Y., Chen Y., and Gu J.. 2018. “Fastp: An Ultra‐Fast All‐In‐One FASTQ Preprocessor.” Bioinformatics 34, no. 17: i884–i890. [DOI] [PMC free article] [PubMed] [Google Scholar]
  17. Chheda, U. , Pradeepan S., Esposito E., Strezsak S., Fernandez‐Delgado O., and Kranz J.. 2024. “Factors Affecting Stability of RNA—Temperature, Length, Concentration, pH, and Buffering Species.” Journal of Pharmaceutical Sciences 113, no. 2: 377–385. [DOI] [PubMed] [Google Scholar]
  18. Chklovski, A. , Parks D. H., Woodcroft B. J., and Tyson G. W.. 2023. “CheckM2: A Rapid, Scalable and Accurate Tool for Assessing Microbial Genome Quality Using Machine Learning.” Nature Methods 20: 1203–1212. 10.1038/s41592-023-01940-w. [DOI] [PubMed] [Google Scholar]
  19. Claassen, S. , du Toit E., Kaba M., Moodley C., Zar H. J., and Nicol M. P.. 2013. “A Comparison of the Efficiency of Five Different Commercial DNA Extraction Kits for Extraction of DNA From Faecal Samples.” Journal of Microbiological Methods 94, no. 2: 103–110. [DOI] [PMC free article] [PubMed] [Google Scholar]
  20. Costea, P. I. , Zeller G., Sunagawa S., et al. 2017. “Towards Standards for Human Fecal Sample Processing in Metagenomic Studies.” Nature Biotechnology 35, no. 11: 1069–1076. [DOI] [PubMed] [Google Scholar]
  21. Eggert, L. S. , Maldonado J. E., and Fleischer R. C.. 2005. “Nucleic Acid Isolation From Ecological Samples—Animal Scat and Other Associated Materials.” In Methods in Enzymology, 73–82. Elsevier. [DOI] [PubMed] [Google Scholar]
  22. Eisenhofer, R. , and Alberdi A.. 2023. “The Earth Hologenome Initiative Bioinformatics Workflow. The Earth Hologenome Initiative Bioinformatics Workflow.” https://www.earthhologenome.org/bioinformatics.
  23. Eriksson, P. , Mourkas E., González‐Acuna D., Olsen B., and Ellström P.. 2017. “Evaluation and Optimization of Microbial DNA Extraction From Fecal Samples of Wild Antarctic Bird Species.” Infection Ecology & Epidemiology 7, no. 1: 1386536. [DOI] [PMC free article] [PubMed] [Google Scholar]
  24. Frostegård, A. , Courtois S., Ramisse V., et al. 1999. “Quantification of Bias Related to the Extraction of DNA Directly From Soils.” Applied and Environmental Microbiology 65, no. 12: 5409–5420. [DOI] [PMC free article] [PubMed] [Google Scholar]
  25. Gaun, N. , Pietroni C., Martin‐Bideguren G., et al. 2025. “The Earth Hologenome Initiative: Data Release 1. GigaScience.” [DOI] [PMC free article] [PubMed]
  26. Hakim, D. , Wandro S., Zengler K., et al. 2022. “Zebra: Static and Dynamic Genome Cover Thresholds With Overlapping References.” mSystems 7, no. 5: e0075822. 10.1128/msystems.00758-22. [DOI] [PMC free article] [PubMed] [Google Scholar]
  27. Hernández, M. , Ancona S., Hereira‐Pacheco S., De La Díaz Vega‐Pérez A. H., and Navarro‐Noya Y. E.. 2023. “Comparative Analysis of Two Nonlethal Methods for the Study of the Gut Bacterial Communities in Wild Lizards.” Integrative Zoology 18, no. 6: 1056–1071. [DOI] [PubMed] [Google Scholar]
  28. Hill, M. O. 1973. “Diversity and Evenness: A Unifying Notation and Its Consequences.” Ecology 54, no. 2: 427–432. [Google Scholar]
  29. Idaghdour, Y. , Broderick D., and Korrida A.. 2003. “Faeces as a Source of DNA for Molecular Studies in a Threatened Population of Great Bustards.” Conservation Genetics 4: 789–792. 10.1023/B:COGE.0000006110.03529.95. [DOI] [Google Scholar]
  30. James, D. A. , Venables W. N., and Ripley B. D.. 1996. “Modern Applied Statistics With S‐PLUS.” Technometrics: A Journal of Statistics for the Physical, Chemical, and Engineering Sciences 38, no. 1: 77. [Google Scholar]
  31. Jansson, L. , Aili Fagerholm S., Börkén E., et al. 2024. “Assessment of DNA Quality for Whole Genome Library Preparation.” Analytical Biochemistry 695: 115636. [DOI] [PubMed] [Google Scholar]
  32. Jost, L. 2007. “Partitioning Diversity Into Independent Alpha and Beta Components.” Ecology 88, no. 10: 2427–2439. [DOI] [PubMed] [Google Scholar]
  33. Kang, D. D. , Li F., Kirton E., et al. 2019. “MetaBAT 2: An Adaptive Binning Algorithm for Robust and Efficient Genome Reconstruction From Metagenome Assemblies.” PeerJ 7: e7359. [DOI] [PMC free article] [PubMed] [Google Scholar]
  34. Koziol, A. , Odriozola I., Leonard A., et al. 2023. “Mammals Show Distinct Functional Gut Microbiome Dynamics to Identical Series of Environmental Stressors.” MBio 14, no. 5: e0160623. [DOI] [PMC free article] [PubMed] [Google Scholar]
  35. Kuznetsova, A. , Brockhoff P., and Christensen R.. 2017. “LmerTest Package: Tests in Linear Mixed Effects Models.” Journal of Statistical Software 82, no. 13: 1–26. [Google Scholar]
  36. Langmead, B. , and Salzberg S. L.. 2012. “Fast Gapped‐Read Alignment With Bowtie 2.” Nature Methods 9, no. 4: 357–359. [DOI] [PMC free article] [PubMed] [Google Scholar]
  37. Leonard, A. , Earth Hologenome Initiative Consortium , and Alberdi A.. 2024. “A Global Initiative for Ecological and Evolutionary Hologenomics.” Trends in Ecology & Evolution 39, no. 7: 616–620. [DOI] [PubMed] [Google Scholar]
  38. Ley, R. E. , Hamady M., Lozupone C., et al. 2008. “Evolution of Mammals and Their Gut Microbes.” Science 320, no. 5883: 1647–1651. [DOI] [PMC free article] [PubMed] [Google Scholar]
  39. Li, D. , Liu C.‐M., Luo R., Sadakane K., and Lam T.‐W.. 2015. “MEGAHIT: An Ultra‐Fast Single‐Node Solution for Large and Complex Metagenomics Assembly via Succinct de Bruijn Graph.” Bioinformatics 31, no. 10: 1674–1676. [DOI] [PubMed] [Google Scholar]
  40. Mak, S. S. T. , Gopalakrishnan S., Carøe C., et al. 2017. “Comparative Performance of the BGISEQ‐500 vs Illumina HiSeq2500 Sequencing Platforms for Palaeogenomic Sequencing.” GigaScience 6, no. 8: 1–13. [DOI] [PMC free article] [PubMed] [Google Scholar]
  41. Marcos, S. , Parejo M., Estonba A., and Alberdi A.. 2022. “Recovering High‐Quality Host Genomes From Gut Metagenomic Data Through Genotype Imputation.” Advances in Genetics 3, no. 3: 2100065. [DOI] [PMC free article] [PubMed] [Google Scholar]
  42. McKee, A. M. , Spear S. F., and Pierson T. W.. 2015. “The Effect of Dilution and the Use of a Post‐Extraction Nucleic Acid Purification Column on the Accuracy, Precision, and Inhibition of Environmental DNA Samples.” Biological Conservation 183: 70–76. [Google Scholar]
  43. Monteiro, L. , Bonnemaison D., Vekris A., et al. 1997. “Complex Polysaccharides as PCR Inhibitors in Feces: Helicobacter pylori Model.” Journal of Clinical Microbiology 35, no. 4: 995–998. [DOI] [PMC free article] [PubMed] [Google Scholar]
  44. Murray, D. C. , Coghlan M. L., and Bunce M.. 2015. “From Benchtop to Desktop: Important Considerations When Designing Amplicon Sequencing Workflows.” PLoS One 10, no. 4: e0124671. [DOI] [PMC free article] [PubMed] [Google Scholar]
  45. Nyholm, L. , Koziol A., Marcos S., et al. 2020. “Holo‐Omics: Integrated Host‐Microbiota Multi‐Omics for Basic and Applied Biological Research.” IScience 23, no. 8: 101414. [DOI] [PMC free article] [PubMed] [Google Scholar]
  46. Oksanen, J. , Blanchet F. G., Kindt R., et al. 2013. “Package “vegan.” Community Ecology Package, Version, 2(9).” http://cran.ism.ac.jp/web/packages/vegan/vegan.pdf.
  47. Olm, M. R. , Brown C. T., Brooks B., and Banfield J. F.. 2017. “dRep: A Tool for Fast and Accurate Genomic Comparisons That Enables Improved Genome Recovery From Metagenomes Through de‐Replication.” ISME Journal 11, no. 12: 2864–2868. [DOI] [PMC free article] [PubMed] [Google Scholar]
  48. Opel, K. L. , Chung D., and McCord B. R.. 2010. “A Study of PCR Inhibition Mechanisms Using Real Time PCR.” Journal of Forensic Sciences 55, no. 1: 25–33. [DOI] [PubMed] [Google Scholar]
  49. Papke, L. E. , and Wooldridge J. M.. 1996. “Econometric Methods for Fractional Response Variables With an Application to 401(k) Plan Participation Rates.” Journal of Applied Economics 11, no. 6: 619–632. [Google Scholar]
  50. Picard Toolkit . 2019. “In Broad Institute, GitHub repository. Broad Institute.” https://broadinstitute.github.io/picard/.
  51. Pietroni, C. , Gaun N., Leonard A., et al. 2025. “Hologenomic Data Generation and Analysis in Wild Vertebrates.” Methods in Ecology and Evolution 16, no. 1: 97–107. [Google Scholar]
  52. Rodgers, T. W. , and Janečka J. E.. 2013. “Applications and Techniques for Non‐Invasive Faecal Genetics Research in Felid Conservation.” European Journal of Wildlife Research 59, no. 1: 1–16. [Google Scholar]
  53. Rodriguez, R. L. M. , Gunturu S., Tiedje J. M., Cole J. R., and Konstantinidis K. T.. 2018. “Nonpareil 3: Fast Estimation of Metagenomic Coverage and Sequence Diversity.” mSystems 3, no. 3: e00039‐18. 10.1128/mSystems.00039-18. [DOI] [PMC free article] [PubMed] [Google Scholar]
  54. Rohland, N. , and Reich D.. 2012. “Cost‐Effective, High‐Throughput DNA Sequencing Libraries for Multiplexed Target Capture.” Genome Research 22, no. 5: 939–946. [DOI] [PMC free article] [PubMed] [Google Scholar]
  55. Russell, A. C. , Kenna M. A., Van Huynh A., and Rice A. M.. 2024. “Microbial DNA Extraction Method for Avian Feces and Preen Oil From Diverse Species.” Ecology and Evolution 14, no. 9: e70220. [DOI] [PMC free article] [PubMed] [Google Scholar]
  56. Schiebelhut, L. M. , Abboud S. S., Gómez Daglio L. E., Swift H. F., and Dawson M. N.. 2017. “A Comparison of DNA Extraction Methods for High‐Throughput DNA Analyses.” Molecular Ecology Resources 17, no. 4: 721–729. [DOI] [PubMed] [Google Scholar]
  57. Schrader, C. , Schielke A., Ellerbroek L., and Johne R.. 2012. “PCR Inhibitors—Occurrence, Properties and Removal.” Journal of Applied Microbiology 113, no. 5: 1014–1026. [DOI] [PubMed] [Google Scholar]
  58. Shaffer, M. , Borton M. A., McGivern B. B., et al. 2020. “DRAM for Distilling Microbial Metabolism to Automate the Curation of Microbiome Function.” Nucleic Acids Research 48, no. 16: 8883–8900. [DOI] [PMC free article] [PubMed] [Google Scholar]
  59. Song, S. J. , Sanders J. G., Delsuc F., et al. 2020. “Comparative Analyses of Vertebrate Gut Microbiomes Reveal Convergence Between Birds and Bats.” MBio 11, no. 1: e02901‐19. 10.1128/mBio.02901-19. [DOI] [PMC free article] [PubMed] [Google Scholar]
  60. Srivathsan, A. , Ang A., Vogler A. P., and Meier R.. 2016. “Fecal Metagenomics for the Simultaneous Assessment of Diet, Parasites, and Population Genetics of an Understudied Primate.” Frontiers in Zoology 13: 17. [DOI] [PMC free article] [PubMed] [Google Scholar]
  61. Tang, Q. , Jin G., Wang G., et al. 2020. “Current Sampling Methods for Gut Microbiota: A Call for More Precise Devices.” Frontiers in Cellular and Infection Microbiology 10: 151. [DOI] [PMC free article] [PubMed] [Google Scholar]
  62. Theis, K. R. 2018. “Hologenomics: Systems‐Level Host Biology.” mSystems 3, no. 2: e00164‐17. 10.1128/mSystems.00164-17. [DOI] [PMC free article] [PubMed] [Google Scholar]
  63. Thomas, T. , Gilbert J., and Meyer F.. 2012. “Metagenomics—A Guide From Sampling to Data Analysis.” Microbial Informatics and Experimentation 2, no. 1: 3. [DOI] [PMC free article] [PubMed] [Google Scholar]
  64. Thomas, V. , Clark J., and Doré J.. 2015. “Fecal Microbiota Analysis: An Overview of Sample Collection Methods and Sequencing Strategies.” Future Microbiology 10, no. 9: 1485–1504. [DOI] [PubMed] [Google Scholar]
  65. Uritskiy, G. V. , DiRuggiero J., and Taylor J.. 2018. “MetaWRAP—A Flexible Pipeline for Genome‐Resolved Metagenomic Data Analysis.” Microbiome 6, no. 1: 1–13. [DOI] [PMC free article] [PubMed] [Google Scholar]
  66. Weidner, L. , Laner‐Plamberger S., Horner D., et al. 2022. “Sample Buffer Containing Guanidine‐Hydrochloride Combines Biological Safety and RNA Preservation for SARS‐CoV‐2 Molecular Diagnostics.” Diagnostics 12, no. 5: 1186. 10.3390/diagnostics12051186. [DOI] [PMC free article] [PubMed] [Google Scholar]
  67. Woodcroft, B. J. , Aroney S. T. N., Zhao R., et al. 2024. “SingleM and Sandpiper: Robust Microbial Taxonomic Profiles From Metagenomic Data.” bioRxiv. 2024.01.30.578060. 10.1101/2024.01.30.578060. [DOI]
  68. Wu, Y.‐W. , Simmons B. A., and Singer S. W.. 2016. “MaxBin 2.0: An Automated Binning Algorithm to Recover Genomes From Multiple Metagenomic Datasets.” Bioinformatics 32, no. 4: 605–607. [DOI] [PubMed] [Google Scholar]
  69. Zemanova, M. A. 2020. “Towards More Compassionate Wildlife Research Through the 3Rs Principles: Moving From Invasive to Non‐Invasive Methods.” Wildlife Biology 2020, no. 1: 1–17. 10.2981/wlb.00607. [DOI] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Data S1: men70042‐sup‐0001‐DataS1.pdf.

MEN-25-e70042-s003.pdf (32.3KB, pdf)

Data S2: men70042‐sup‐0002‐DataS2.pdf.

MEN-25-e70042-s002.pdf (31.5KB, pdf)

Figure S1: Integrity of the DNA extracted using the three methods in four randomly selected samples. The first peak corresponds with the marker, while the second peak indicates the proportion of DNA fragments with distinct sizes.

Figure S2: Laboratory and sequencing performance metrics of the three extraction protocols. (A) 260/280 nm absorbance ratio indicating the purity of the DNA extraction. (B) 260/230 nm absorbance ratio indicating the purity of the DNA extraction. (C) Sequencing depth in gigabases (thousands of millions of nucleic acid bases). (D) Kmer‐based estimation of the recovered metagenomic complexity. (E) Proportion of the estimated microbial community recovered in each sample.

MEN-25-e70042-s001.docx (937.8KB, docx)

Data Availability Statement

The raw data tables containing the quantitative information, data accession codes, and bioinformatic code scripts used for the analysis are available in a dedicated GitHub repository (https://github.com/earthhologenome/EHI_extraction_test), both rendered into a HTML webbook (https://www.earthhologenome.org/EHI_extraction_test) and frozen in Zenodo with doi:10.5281/zenodo.16757704. Raw sequencing data and microbial genome sequences, along with their respective sampling, export, ABS and import permits, were published as part of the 1st EHI Data Release (Gaun et al. 2025), under Bioproject PRJEB76898. The samples used for the validation of the extraction methodologies were procured following the standardised administrative procedures of the Earth Hologenome Initiative, which has strict rules regarding local and international Access and Benefit Sharing regulations, including the Nagoya Protocol. The relevant authorisations for this work include permits 168E/2022 and 254E/2023 by the regional Government of Navarra; permit 2023FAUNA0000010501 by the regional Government of Gipuzkoa; permit 04–23‐550/16 ZM by the Federal Ministry of Environment and Tourism of Serbia; permit MAAE‐ARSFC‐2022‐2083 by the Ministry of Environment, Water and Ecological Transition of Ecuador; permit 23‐20 by the Agricultural Department of Sweden; permit DGMNB/SEN/avp_21_187 by the regional Government of Castilla‐La Mancha; permit DD‐1015 by the Region of Lombardia; permit DBCA‐AEC 2021‐01A by the Australian Wildlife Conservancy; and permit 648/2023/CAPT by the Institute of Nature Conservation of Portugal. The benefits of this research include sharing our code, data, and results on public databases as outlined above.


Articles from Molecular Ecology Resources are provided here courtesy of Wiley

RESOURCES