Skip to main content
Wiley Open Access Collection logoLink to Wiley Open Access Collection
. 2025 Nov 8;26(1):e70077. doi: 10.1111/1755-0998.70077

Comparison of Metabarcoding and Shotgun Sequencing Confirms the Relevance of Chloroplastic rRNA Genes to Assess Community Structure of Lake Phytoplankton

Clarisse Lemonnier 1,, Benjamin Alric 1,2, Isabelle Domaizon 1,2, Frédéric Rimet 1,2,
PMCID: PMC12627912  PMID: 41204870

ABSTRACT

Studying the taxonomic composition of phytoplankton has been revolutionised by the emergence of metabarcoding approaches which theoretically provides access to all phytoplankton diversity. However, metabarcoding has its limitations, including biases related to primers efficiency in covering all phytoplanktonic taxonomic groups, biases introduced during amplification and due to databases completion. To assess the importance of these biases, we compared the compositions of phytoplankton in two large peri‐alpine lakes, using metabarcoding with primers targeting hypervariable regions of chloroplastic 16S and 23S rRNA, designed to cover the full taxonomic diversity of phytoplankton, and shotgun sequencing. To be able to directly compare the two methods, we extracted reads coming from the full sequences of these rRNA genes in shotgun sequencing data and used the same reference database for taxonomic assignation. The results show that the relative abundances of dominant groups of phytoplankton, including Cyanobacteria, Cryptophyta, Haptophyta, and Bacillariophyta, are consistent between the two approaches, validating the primers used for metabarcoding analysis. However, two phyla, Chlorophyta and non‐diatom Ochrophyta showed greater divergence in their relative abundance, due to under‐amplification or lack of amplification of certain taxonomic groups in metabarcoding. This is likely due to the high diversity of these groups, not covered yet by the reference databases, as well as a possible presence of introns in their choloroplastic ribosomal genes. These limitations are expected to be overcome with increasing reference database completion and the use of long‐read metabarcoding. Overall, our study confirms the relevance of using chloroplastic primers for assessing the phytoplankton composition of lakes.

Keywords: lakes, metabarcoding, PCR biases, phytoplankton, shotgun, taxonomic composition

1. Introduction

Lakes play a crucial role in ecological functions like the carbon cycle and climate regulation (Cole et al. 2007) while providing essential ecosystem services (Gleick 1993; Sterner et al. 2020). Phytoplankton (photosynthetic micro‐eukaryotes and cyanobacteria) is a key biological component of these ecosystems. Investigating the mechanisms affecting phytoplankton diversity, the processes shaping community assembly is of primary interest to understand lake ecosystem stability and functioning (e.g., Reynolds 1987; Interlandi and Kilham 2001; Salmaso 2010; Pomati et al. 2011). Its large diversity and fast responses to environmental variations make it also a key bioindicator used by the Water Framework Directive (WFD) to assess lake ecological quality (European Commission 2000). Accurately characterising phytoplankton diversity and abundance is thus critical to understand lakes ecological functioning and their responses to anthropogenic pressures and climate change that affect lakes worldwide (Borics et al. 2021 ).

The taxonomic composition of phytoplankton is classically determined via light microscopy (e.g., Utermöhl 1958), currently, the reference method for WFD biomonitoring. However, this approach has limitations, struggling to capture the full spectrum of picoplankton diversity, cryptic diversity and very rare taxa (Bergkemper and Weisse 2018; Smayda 2011). Variability in the taxonomic identification between analysts further affects water quality assessments based on micro‐algae (Prygiel et al. 2002; Kahlert et al. 2009; Gao et al. 2018). To overcome these limitations, an alternative methodology based on high‐throughput sequencing of genetic barcodes (i.e., metabarcoding) has been tested with success (Canino et al. 2021; Nicolosi Gelis et al. 2024). One of the advantages of metabarcoding is that it can theoretically target the entire phytoplankton diversity, regardless their size. It is also seen as a more reliable, cheaper and faster method than microscopy (Abad et al. 2016). These strengths led metabarcoding to be rapidly applied to the study of microalgal ecology, increasing the temporal or spatial resolution at an unprecedented level (Sommeria‐Klein et al. 2021) and revealing a diversity never seen before (De Vargas et al. 2015). It has also been successfully implemented in biomonitoring (Vasselon et al. 2019; Kelly et al. 2020; Andersson et al. 2023).

However, metabarcoding relies on some critical steps potentially affecting both the number of taxonomic groups detected and their relative abundances (Santoferrara 2019). First, the barcode selection has consequences on the detected diversity. Identifying a single barcode common to all targeted taxa is challenging, especially in a highly diverse and polyphyletic group like phytoplankton. For instance, small hypervariable regions of the 18S rRNA genes are by far the most commonly used barcodes for phytoplankton surveys (Stoeck et al. 2010), but it misses all Cyanobacteria which then needs be amplified aside, using for instance the prokaryotic 16S rRNA gene (Nübel et al. 1997). Also, it is common to get a high number of 18S barcodes from unknown Dinoflagellates that cannot be classified as either photosynthetic or non‐photosynthetic. An alternative is to use chloroplastic rRNA genes that target both prokaryotic and eukaryotic photosynthetic microorganisms. However, some taxa such as Dinoflagellates or Chomprodellids are poorly represented with these barcodes as they present highly divergent plastidics rRNA genes or fragmented chloroplastic DNA (Howe et al. 2008; Barbrook et al. 2019). In addition, the barcode selection can influence the relative abundances observed. The number of copies of the rRNA gene varies depending on the species, which leads to an artifactual over‐estimation of some taxa (Pierella Karlusich et al. 2023; Santi et al. 2021; Martin et al. 2022). For instance 18S rRNA genes can variate from one to more than 100,000 per cells in some dinoflagellates (Martin et al. 2022). If these variations are less pronounced for chloroplastic rRNA genes, they can also be due to changes in gene copy number per plastome, plastomes per plastid and plastids per cell, and in addition, the number of chloroplasts can change depending on organismal life's stage (Decelle et al. 2019; Vasselon et al. 2018). Second, PCR is another critical step, which can distort the diversity and proportions of phytoplankton taxa detected in the samples (Acinas et al. 2005; Santoferrara 2019). Primers mismatches on targeted genes of certain taxa can lead to erroneous abundances of entire taxonomic groups (from species to phylum) that are poorly or not amplified at all (Bradley et al. 2016). The PCR step itself can over or under amplify certain taxa and will tend to better amplify shorter and structurally simpler DNA fragments (Taberlet et al. 2012; Geisen et al. 2015). This is also at this step and during the sequencing that some errors are added in the amplified sequences, which can lead to misclassification and over‐estimation of some phytoplanktonic taxa (Behnke et al. 2011; Bachy et al. 2013).

Usually, metabarcoding biases impact on diversity and taxonomic composition of phytoplankton communities are tested using either in silico methods (Latz et al. 2022; Canino et al. 2023) or mock samples with a known diversity (Nübel et al. 1997; Marinchel et al. 2023), or also by comparing species composition obtained in an environmental sample with microscopic counts (Huo et al. 2020; Brown et al. 2022; Nicolosi Gelis et al. 2024; Mikhailov et al. 2025). However, conclusions of such approaches are partial, insofar as (i) mock samples represent simplified samples compared to natural community diversity and (ii) comparison with microscopy is limited due to different technical aspects which were clearly identified in Nicolosi Gelis et al. (2024). The four main technical aspects that emerge from the study of Nicolosi Gelis et al. (2024) are listed and summarised below. First, with metabarcoding a filtration step is necessary to recover cells and DNA which is not the case for microscopy, as a sedimentation is carried out to visualise cell under inverted microscopy. Second, certain groups, particularly picocyanobacteria and picoeukaryodes, were poorly detected under microscope because too small, whereas metabarcoding enabled to detect a significant proportion of reads and diversity of these groups. Third, reference barcoding libraries are incomplete for some groups which are on the contrary well detected in microscopy (e.g., non‐diatom Ochrophyta are poorly represented with 23S rRNA gene compared to the known diversity). Fourth, some taxonomic groups, although large enough to be detected under a microscope, were mistaken with others because their morphological characteristics are difficult or even impossible to observe with routine identifications on material fixed with Lugol's solution (e.g., Chlorella like identified in microscopy were in fact Eustigmotophyceae). For all these reasons, the real impact of metabarcoding biases on the composition of phytoplankton taxonomic inventories remains to be determined, as the resources needed to estimate phytoplankton metabarcoding reliability for environmental samples are still lacking.

Shotgun sequencing is an alternative DNA‐based approach that can be used to assess microbial composition, but unlike metabarcoding, the entire genetic information is sequenced, without prior amplification. It thus represents an interesting reference method for studying the impact of PCR biases on taxonomic composition retrieved in environmental samples (Tedersoo et al. 2015; van der Loos and Nijland 2021). Several studies compared the taxonomic compositions of microbial communties obtained with shotgun and metabarcoding data (e.g., 16S rRNA gene in Tessler et al. 2017; Stothart et al. 2023). Overall, these studies showed that classical diversity metrics (e.g., alpha and beta‐diversity) are relatively comparable between the two methods (Rieder et al. 2023; Stothart et al. 2023). However, they did not specifically focus on the impact of PCR biases. On this specific point, their conclusions are limited as they often do not compare the same information: taxonomic diversity obtained from the whole genome for shotgun vs. taxonomic diversity obtained from one small barcode for metabarcoding. In addition, the capacity to recover different taxa in shotgun data is also strongly influenced by the reference databases used (Lind and Pollard 2021; Stothart et al. 2023) and retrieving a precise taxonomical classification from shotgun reads remains challenging (Breitwieser et al. 2019). For instance, some tools used for shotgun diversity analysis can result in a large number of false positive classifications as well as erroneous abundances (Milanese et al. 2019; Chakoory et al. 2022). To finely study the potential impact of PCR biases on the taxonomic composition of phytoplankton, one should compare the community composition obtained using the same barcode and the same reference database (e.g., 16S rRNA genes in Logares et al. 2014; ITS in Tedersoo et al. 2015; 18S‐V4 in Pernice et al. 2016). For phytoplankton, very few studies did so on 18S hyper variable regions barcodes for assessing marine communities' composition (Obiol et al. 2020; Pierella Karlusich et al. 2023). These studies were able to disentangle PCR biases, showing that some taxa were under‐represented in metabarcoding, such as Prymnesiophyceae for 18S‐V4 or Haptophyta for 18S‐V9.

Our aim in this study is to use shotgun sequencing to infer the importance of PCR biases present in metabarcoding, in assessing the community structure of lake phytoplankton. To this end, we compared the relative abundances of different taxa obtained from the V3‐V4 and UPA hypervariable regions of the 16S and 23S chloroplastic rRNA genes amplified through metabarcoding, with reads of these full rRNA genes extracted in shotgun sequencing. Taxonomy was assessed using the same reference barcoding library (Phytool; Canino et al. 2021). The purposes of this comparison are (i) to check the reliability of the relative abundances obtained by metabarcoding for the dominant phyla and (ii) to assess the impact of PCR biases on these relative abundances and their consequences on the community structure. Prior to this, we validated an open‐access bioinformatic pipeline for the analysis of chloroplastic rRNA genes from shotgun data.

2. Material and Methods

2.1. Study Sites and Sampling

Two large peri‐alpine lakes in France were sampled monthly or bimonthly throughout year 2021 as part of the Observatory on LAkes (OLA; Rimet et al. 2020): lake Geneva (18 samples) and lake Annecy (14 samples). An integrated water sample (using IWS Hydro‐bios) was collected in the first 20 m depth above the deepest point of the lakes. Each sample was split into two subsamples, one for routine microscopic identification and counting (100 mL), and the second (250 mL) was used for molecular analysis. Microscopic and metabarcoding data were compared in Nicolosi Gelis et al. (2024). The subsamples dedicated to molecular analyses, were filtered in the laboratory on open filters (0.45 μm MF‐Millipore, cellulose nitrate membrane filter) and immediately frozen at −20°C in a 2 mL Eppendorf microtube until DNA extraction.

2.2. DNA Extraction, Metabarcoding and Shotgun Libraries Preparation

DNA was extracted using the NucleoSpin soil kit (Macherey‐Nagel), following the recommendations of the manufacturer. The same DNA extract was used for metabarcoding and shotgun libraries preparation.

For the metabarcoding approach, taxonomic composition of phytoplankton was assessed using both 16S and 23S chloroplastic barcodes as described in Nicolosi Gelis et al. (2024). Briefly, the V3–V4 region of the 16S rRNA (a barcode of 380 bp) was amplified using the pair of primers CYA359F (5′GGGGAATYTTCCGCAATGG‐3′) and CYA781Rd (5′‐GACTACWGGGGTATCTAATCCCWTT‐3′) (Nübel et al. 1997), and the UPA region of the 23S rRNA gene (a barcode of 358 bp) was amplified using the pair of primers ECLA23S_F1 (5′‐ACAGWAAGACCCTATGAAGCTT‐3′) ECLA23S_R1 (5′CCTGTTATCCCTAGAGTAACTT‐3′) (Canino et al. 2023). PCR products were sent for sequencing to the PGTB platform (Plateforme de Genomique et Transcriptomique, Bordeaux, France), using Illumina MiSeq technology and the v3 reagent kit (2 × 250 bp), with an expected sequencing depth of around 50,000 reads per sample.

Shotgun libraries were constructed in GeTplaGe plateform at Toulouse following the TrueSeq Nano kit from Illumina. Sequencing of the 32 libraries was performed on a NovaSeq6000 2 × 150 bp using a S4 flowcell with a sequencing depth of around 100 million reads per sample.

2.3. Bioinformatic Analysis

2.3.1. Metabarcoding Bioinformatic Analysis

Composition of phytoplankton was assessed by defining Amplicons Sequence Variants (ASVs) using DADA2 pipeline (Figure 1A), implemented in the R‐package dada2 (Callahan et al. 2016). First, primers were checked and removed from demultiplexed FASTQ reads using Cutadapt v3.5 (Martin 2011). Then, reads were filtered using the function filterAndTrim() with parameters set as default. After inference of sequencing errors with dada(), forward and reverse reads were merged with mergePairs() function, and chimaera were detected de novo and removed using the function removeBimeraDenovo() with the consensus method. Sequences were size‐filtered to keep those of length between 375–385 bp for 16S barcode and 345–370 bp for 23S barcode, to avoid non‐target sequences. Taxonomic assignment of each ASV was done with the RDP (Ribosomal Database Project) naive Bayesian classifier (Wang et al. 2007) implemented in the R‐package dada2, using Phytool v2 (Canino et al. 2021) and SILVA 138.1 (Quast et al. 2012) as reference databases. Phytool is a reference barcoding library dedicated to eukaryotic algae and cyanobacteria with an up to date and curated taxonomy, while SILVA is a generalist database for rRNA genes. The taxonomic assignment of a sequence with Phytool was kept if it had a bootstrap value greater than 80% and if it was assigned as Chloroplast or Cyanobacteria with SILVA database. This last step was necessary to remove non‐phytoplanktonic ASVs, which could still be assigned with Phytool (e.g., Bacteria) due to the low number of sequences in the reference database. Finally, ASV affiliated to Trachaeophyta were removed. Abundance of the different ASVs or taxa was expressed in relative abundance of reads in a sample, to account for sequencing depth variability.

FIGURE 1.

FIGURE 1

Schematic resuming the different analyses done in this study, using 23S as an example. (A) Metabarcoding steps with DNA amplification, MiSeq sequencing, DADA2 pipeline and taxonomic assignment using assign. taxonomy() command and Phytool reference databases. (B) The metagenomic steps to recover taxonomic composition of phytoplankton with DNA fragmentation, NovaSeq sequencing, SortMeRNA to recover rRNA reads and taxonomic assignment with blast of the extracted rRNA reads against Phytool. (C) A fine comparison of the ASV sequence and abundance using mapping of the shotgun reads. A sequence with a colour means that it was assigned taxonomically.

2.3.2. Shotgun Bioinformatic Analysis

Raw reads were quality‐filtered as in the first step of anvio metagenomic workflow (Eren et al. 2015), following the recommendations of Minoche et al. (2011).

To have the best comparability between metabarcoding and shotgun methods to assess taxonomic composition of phytoplankton, we targeted the same genes (16S and 23S) with the shotgun approach. There are multiple tools developed to assess the microbial diversity from shotgun data based on rRNA genes, but most of them target only the small subunit of rRNA (e.g., Ribotaxa and Phyloflash) and none of them gives a robust taxonomy for chloroplastics rRNAs (e.g., Metaxa2 is based on SILVA taxonomy). In this study, we thus optimised and validated a two‐steps pipeline to assess freshwater phytoplankton composition based on shotgun data (Figure 1B). The first step involved extracting reads that corresponds to parts of all rRNA genes using SortMeRNA v4.3.6 (Kopylova et al. 2012), and then blasting them into the Phytool reference database to obtain their taxonomic identity. The only difference with metabarcoding approach is that results obtained are based on the full 16S or 23S genes and not on a specific barcode. We explain in details this pipeline here below.

SortMeRNA is a dedicated tool to extract ribosomal RNAs present in a dataset. To do this, it uses a reference database and builds statistical models of the rRNA sequences, making it able to extract rRNA reads even if they were not present in the reference database. We chose SILVA as reference database, as it is the most generalist one for rRNA genes of microorganisms. In order to optimise the read extraction (i.e., avoid non‐ribosomal reads to be extracted), SILVA databases were curated as suggested in Phyloflash pipeline (Gruber‐Vodicka et al. 2020). Briefly, SILVA sequences were transformed from RNA to DNA alphabet and ambiguous bases were randomly replaced using a python script so that they can be used by SortMeRNA. SILVA sequences corresponding to the wrong ribosomal subunit were detected using the nhmmer tool (Wheeler and Eddy 2013) and removed using the Hidden Markov Model (HMM) profiles of SSU or LSU of barrnap (https://github.com/tseemann/barrnap), adapted in Phyloflash. This ensures that only reads from the desire subunit are extracted. Sequences with low complexity were masked using bbmask program of the BBTools suite of bioinformatics tools (https://sourceforge.et/projects/bbmap/), and sequences identified as vector were detected using bbduk program form BBTools against NCBI vector database and then deleted. Once this curation was complete, reads belonging to SSU (18S and 16S) and LSU (28S and 23S) ribosomal genes were extracted separately within each sample using SortMeRNA setting an e‐value cutoff of 1e‐9.

The SSU and LSU SILVA reference database used with SortMeRNA differs in their completion (SILVA NR99 SSU has 509,885 sequences after curation, while SILVA NR99 LSU has only 94,811 sequences). To assess to which point the extraction of rRNA reads could be influenced by the reference database, we compared the number of reads extracted for the SSU with the number extracted for the LSU. As the LSU genes are longer than SSU genes, they logically extract a higher number of reads. For the comparison, we thus normalised the number of read extract by the mean gene size between prokaryotic and eukaryotic rRNA LSU (2906 bp for 23S and 5070 bp for 28S) and SSU (1542 bp for 16S and 1869 bp for 18S).

Taxonomy of reads coming from freshwater phytoplankton was obtained by blasting the extracted rRNA reads against Phytool 16S and 23S reference databases, using nblast (Camacho et al. 2009). An assignment was validated by setting high selection criteria: (i) a hit (i.e., match with a sequence in the database) was retained if it presented more than 95% identity over a minimum length of 140 bp and (ii) with less than one mismatch or open gap. Then, an assignment was retained if all hits of a read matched the same phyla in Phytool reference database and if no better blast results was obtained against SILVA. This last step was necessary, particularly to avoid false taxonomic assignment for reads that match a conserved region of the rRNA gene and could come from another phylum or even from heterotrophic bacteria.

Similarly to SILVA reference database, Phytool 16S and Phytool 23S differ in their completion (8479 and 1997 sequences, respectively). To assess to which point this affects the number of reads assigned to a taxon, we compared the relative abundance of each phyla obtained for the 16S and 23S at the end of the pipeline.

For a more complete analysis we also compared the taxonomic composition of chloroplastic rRNA genes based on shotgun reads with microscopy data and the nuclear 18S rRNA gene. For the later, this was done using the exact same protocol as the 16S and 23S but using Phytool 18S database. Results are given in Data S2.

In addition to rRNA extraction and comparison of the taxonomic composition of 23S obtained by shotgun sequencing and metabarcoding, we tested another gene, psbO, which could be promising for obtaining unbiased taxonomic communities' compositions. Indeed, this nuclear gene is present in a single copy per cell (Pierella Karlusich et al. 2023). This analysis is presented in Data S1, but unfortunately was not successful (see reasons in Data S1).

2.3.3. Evaluation of Metabarcoding Biases

The first question regarding metabarcoding biases concerned the ability of primers to cover all the main phytoplanktonic groups and provide reliable relative abundances. After the two‐step pipeline for shotgun data, we obtained the abundance of the main phytoplankton groups. The abundance of a given taxon corresponds to the number of reads assigned to this taxon in a sample. As for metabarcoding data, abundance of the different taxa in a sample was expressed in relative abundance (number of reads assigned to a taxon divided by the total number of reads in a sample), to account for sequencing depth variability.

Another major source of metabarcoding biases is during the PCR, during which the amplification efficiency can be variable and errors can be introduced into the sequences (as during sequencing). To investigate the importance of these biases we directly mapped (i.e., align with 100% identity) the short‐reads obtained with shotgun sequencing onto the ASV reconstructed with the metabarcoding data.

Phytoplankton 16S rRNA (or 23S) reads extracted with SortMeRNA were mapped against ASV of 16S (or 23S), using bowtie2 v2.5.1 (Langmead and Salzberg 2012). This mapping was done twice. A first mapping was done to distinguish « true » ASVs from putative remaining ASVs with sequencing errors left by DADA2 pipeline. DADA2 aims to infer the real sequence, corrected of their sequencing errors using error models, but some errors can remain (Prodan et al. 2020). We defined an ASV as « true » if 80% or more of its length was covered by reads mapping with 100% identity (given with bowtie option score‐min L, 0, −0.0), meaning that this exact ASV sequence was recovered with both metabarcoding and shotgun reads. Then, a second mapping at 100% identity was done on these ASVs to obtain their abundance consisting in the number of short‐reads mapped in an ASV within each sample. As previously, these abundances were expressed as relative abundance to account for sequencing depth variability.

Using this approach, we could assess to which point ASVs really represent the diversity and composition present in a sample as found with the shotgun approach. Indeed, if a low proportion of shotgun reads affiliated to a phylum mapped onto ASVs of this phylum, we hypothesise that this is because some taxa from this phylum were less recovered in metabarcoding. We then assessed the over‐ or under‐estimated ASVs between metabarcoding and shotgun approach. For each ASV, we calculated the number of times it was found more abundant or less abundant in metabarcoding data in all the 32 samples. This was done on clr‐transformed data to account for compositional bias by using decostand() function of the R‐package vegan (Oksanen et al. 2013). Prior to this, 0 values for ASV in a sample were replaced by the 2/3 of the smallest non‐null value of an ASV in the sample. Following the same analysis, we also looked for over‐ or under‐estimated taxa at the family level.

Finally, to see if the amplification step affects the structure of phytoplankton community, we estimated the congruence between ASVs tables obtained from metabarcoding approach and shotgun mapping. This was assessed in a multivariate way using a co‐inertia analysis, based on two normalised PCAs, performed with the coinertia() function of the R‐package ade4 (Dray et al. 2007). The significance of the association was tested using the Monte‐Carlo test by comparing the observed RV coefficient to an expected distribution obtained by 9999 permutations. Prior to this analysis, relative abundance table from metabarcoding and shotgun‐based ASV abundance were clr‐transformed, which converts the data so it can be used in the Euclidian space of the PCA (Aitchison 1986).

All statistics and graphics were performed using R v. 4.4.1.

3. Results

3.1. Validation of the Shotgun Pipeline to Assess Phytoplankton Community Structure

At the end of the first stage of our bioinformatics pipeline (Figure 1B), rRNA reads extracted with SortMeRNA represented between 0.10% and 0.19% of the total number of reads depending of the sample considered for the SSU and between 0.17% and 0.33% percent of reads for the LSU. Despite a very different number of sequences in the reference databases used for their extraction (509,885 for SILVA NR99 SSU and 94,811 for SILVA NR99 LSU, after curation), the number of reads extracted from the two subunits was comparable, as illustrated by a very good fit of the linear model (adjusted R 2 = 0.97) and a slope of the model equal to one (p‐value < 0.001; Figure 2A).

FIGURE 2.

FIGURE 2

(A) Comparison of the number of reads extracted with SortMeRNA for the SSU (18S and 16S) and LSU (28S and 23S) rRNA genes. As LSU genes are longer than SSU genes (i.e., composed of more reads), the abundance of the reads was normalised by the mean length of the two genes. (B) Comparison of the relative abundance obtained for the main phytoplankton phyla based on 16S and 23S shotgun rRNA reads. Adjusted R‐squared and equation of the linear model are indicated for each group.

After the second stage of our bioinformatics pipeline, corresponding to a blast against Phytool reference database, depending on the sample, from 3.8% to 32.1% of the total rRNA reads extracted for the 16S and from 2.6% to 32.3% for the 23S matched to micro‐algae. For the main phyla, the associated linear models highlighted a significant relationship between the relative abundances obtained by shotgun 23S and 16S (with adjusted R 2 greater than or equal to 0.95, and a slope between 0.8 and 1.3; Figure 2B). These results suggest that taxonomic assignment at phylum level was not biassed by the reference databases of the two marker genes, although there is a large difference in their completion (i.e., 8479 sequences in Phytool 16S versus 1997 for Phytool 23S). Bacillariophyta was slightly under‐represented with Phytool 23S, while Haptophyta and non‐diatom Ochrophyta were slightly under‐represented with Phytool 16S. Overall, these results showed that the bioinformatics pipeline for assessing phytoplankton community structure using shotgun data yields robust and comparative relative abundances between 16S and 23S marker genes, supporting the hypothesis that these data can be used as a reference to verify and validate the relative abundances obtained with the metabarcoding approach.

3.2. Are the Relative Abundances of the Main Phytoplankton Groups Similar Between the Two Methods?

Comparison of the relative abundances of the main phyla obtained by metabarcoding and shotgun showed comparable results (Figure 3). However, there were some disparities depending on the phylum considered. Indeed, Cryptophyta, Cyanobacteria, Bacillariophyta and Haptophyta showed the strongest correlations between the two methods (adjusted R 2 from 0.70 to 0.84), whereas Chlorophyta and non‐diatom Ochrophyta showed weaker correlations (adjusted R 2 of 0.54 and 0.67 for Chlorophyta and 0.66 and 0.79 for non‐diatom Ochrophyta). This latter phylum presented lower relative abundances for 16S metabarcoding (with a slope of 0.2).

FIGURE 3.

FIGURE 3

Comparison of the relative abundances obtained for the main phytoplankton phyla based on metabarcoding and shotgun sequencing. As the relative abundances obtained in shotgun did not differ depending on the barcode, here, values of the 23S were taken as a reference for shotgun values. R‐squared and the equation of the linear model are indicated for each group and each barcode.

The phytoplankton community composition obtained using the chloroplastic rRNA genes contrast with the one obtained using the nuclear 18S rRNA gene (Data S2B). As expected for this gene, Dinoflagellates (Miozoa phylum) were highly abundant, representing up to 72% of the phytoplankton community whereas they were poorly represented in both microscopy and chloroplastic rRNA results (Data S2A). But overall, molecular analysis shows more comparable results as compared to microscopy (Data S2A).

3.3. What Could Explain the Discrepancies Between Shotgun and Metabarcoding?

After the two mapping steps of reads, a large majority of the dominant ASVs were kept: 149 over 1157 ASVs for 23S, which represents 91% of the reads in total. For 16S, 147 ASVs over 426 were kept, which represents 97% of the reads in total. This validated that almost all abundant ASVs were likely « true » ASVs, as they were recovered at 80% minimum by shotgun mapping with 100% identity (Data S3). Within each taxonomic group, the number of ASVs detected with shotgun was similar between the two marker genes (Data S4), with the exception of Chlorophyta which presented only 6 ASVs that were validated in 16S against 18 in 23S, and Cryptophyta which presented only 19 ASVs that were detected in 23S against 34 in 16S.

Phyla with strong correlations in relative abundance between shotgun and metabarcoding (e.g., Cyanobacteria, Cryptophyta or Haptophyta in Figure 3) were those presenting the highest percentages of mapping in Figure 4. Chlorophyta and non‐diatom Ochrophyta had the lowest number of shotgun‐reads mapping to them for both barcodes, suggesting that the ASV obtained with metabarcoding may not adequately reflect their diversity and taxonomic composition in lakes Annecy and Geneva. These phyla were the only ones presenting orders detected in microscopy that were not detected at all with metabarcoding, for example, Prasiolales for the Chlorophyta, Goniochloridales, Hibberdiales, Pedinellales, Phaeothamniales and Tribonematales for the non‐diatom Ochrophyta (Data S5).

FIGURE 4.

FIGURE 4

Percentage of shotgun reads that could be mapped to an ASV within each phylum, as compared to the total number of shotgun reads that were assigned to this phylum, for 16S (A) and 23S (B). The dot line corresponds to the percentage of reads expected: As shotgun reads possibly originate from the entire ribosomal gene, we could not expect that 100% of reads mapped onto the ASVs, which represent only a part of the ribosomal genes. Instead, we could expect a maximum percentage of mapping reads of 24.02% and 12.3% for the 16S and 23S, respectively, which corresponds to the fraction each barcode represents of the full ribosomal gene length (16S: 380 bp/1542 bp = 24.02 and 23S: 358 bp/2906 bp = 12.3).

By comparing the clr‐transformed relative abundance of each ASV, we found that rare ASVs were repeatedly (in more than 50% of the samples) under‐estimated with metabarcoding, while most of the dominant ASVs were generally over‐estimated (Data S6). We then investigated whether these differences in relative abundances had repercussions on the estimation of family‐level relative abundances (Data S5). For the 23S, 5 families were regularly under‐estimated in metabarcoding among Chlorophyta (4 families: Ulotrichaceae, Chlorellaceae, Monomastigaceae and Volvocaceae) and Cryptophyta (1 family: Cryptomonadaceae). For the 16S barcode, 7 families were recurrently under‐estimated in non‐diatom Ochrophyta (4 families: Dinobryaceae, Chromulceae, Monodospidaceae and Mallomonadaceae) and Chlorophyta (3 families: Coccomyxaceae, Chlamydomonadaceae and Mychonastaceae).

3.4. Is There a Congruence in the Structure of Phytoplankton Communities Obtained With the Two Methods?

For both barcodes, the co‐structure between the two approaches were significantly similar, with RV of 0.67 (p < 0.001) and 0.77 (p < 0.001) for 16S and 23S, respectively (Figure 5). Regardless on the barcode, samples from the two markers clearly separates based on the lake origin on the first axis and depending on the season on the second axis, indicating a spatio‐temporal structure of the phytoplankton communities. Despites these trends in community structure, the length of the arrows, that represents the level of congruence between shotgun and metabarcoding is variable between samples: some have a good congruence (short arrows) while others have a poorer congruence (long arrows), indicating discrepancies in the ASV composition.

FIGURE 5.

FIGURE 5

Co‐inertia analysis describing the co‐structure between phytoplankton communities characterised by metabarcoding and shotgun approaches, for each marker gene (16S and 23S). The dot in the graph corresponds to metabarcoding samples and the end of the arrows to shotgun‐mapping samples.

4. Discussion

While metabarcoding is increasingly used to characterise the diversity and taxonomic composition of complex biological communities and is seen as a promising tool for assessing the ecological status of ecosystems (Pawlowski et al. 2018; Blancher et al. 2022), it is important to know the extent of its technical biases and its consequences for biodiversity assessment. By sequencing all the DNA present in a sample without any prior amplification, shotgun sequencing highlights PCR biases. Our study, which targeted rRNA marker genes in metagenomic samples, required a high sequencing depth per sample (100 million reads per sample) to ensure a robust detection. Given the current state of sequencing technology, this remains costly in terms of both finances and computing resources, which limits the number of samples than can be analysed (in our case, 32). To our knowledge, this is the first study comparing shotgun and metabarcoding data of lake phytoplankton using the same chloroplastic marker genes. In this study, we were able to extract and detect chloroplastic rRNA genes of the phytoplankton communities. This approach, likely not impacted by reference databases, allows us to draw robust conclusions about the reliability of metabarcoding in characterising the composition of lake phytoplankton communities.

4.1. Metabarcoding and Shotgun Reveal Similar Community Structure and Phyla Relative Abundances

Our analysis shows that PCRs performed during the metabarcoding workflow did not significantly bias the community structures and the taxa proportion of the phytoplankton communities in the two studied lake.

First, we observed a significant congruence in community structures obtained with both molecular approaches. Comparison of the co‐structure of microalgal communities characterised by metabarcoding and microscopy were carried out in former studies and showed also a good congruence particularly when using a taxonomy‐free approach (Brown et al. 2022; Nicolosi Gelis et al. 2024). Such comparisons are important, because they demonstrate that both methods reveal similar phytoplankton community structures. Given that microscopy data capture changes in phytoplankton composition induced by environmental conditions, such as seasons or nutrient levels, which are key factors shaping phytoplankton communities of these lakes (e.g., Anneville et al. 2018), these findings show that metabarcoding is also able of capturing them. We showed that some known sources of bias had limited impact on the global structure of the phytoplankton community. While the primers used introduced an expected PCR bias, namely over‐estimation of abundant ASVs and under‐estimation of low‐abundance ASVs (Gonzalez et al. 2012), the community structure assessed from metabarcoding remained consistent with that obtained by shotgun sequencing. Similarly, regarding the ASV sequences, we could observe that the algorithms currently used for metabarcoding (i.e., DADA2 pipeline) are performant to infer the true sequences and correct them from their PCR and sequencing errors, giving similar number of ASVs with the two barcodes (16S and 23S) for most of the phytoplankton groups.

Second, the relative abundances of main phytoplankton groups (i.e., Cyanobacteria, Cryptophyta, Bacillariophyta and Haptophyta) obtained with shotgun and metabarcoding were correlated, confirming the ability of the 16S and 23S chloroplastic primers to represent similarly the relative abundances of both prokaryotic and eukaryotic lake phytoplanktonic taxa. The two molecular approaches consistently highlighted the overwhelming dominance of Cyanobacteria and Cryptophyta in the oligotrophic and mesotrophic studied lakes (Callieri and Pinolini, 1995; Personnic et al. 2009). This result supports the use of such chloroplastic barcodes to monitor the community structure and taxonomic composition of freshwater phytoplankton, as they represent the only ones capable of assessing all the phytoplankton groups, including eukaryotic algae and Cyanobacteria (in opposition to 18S rRNA metabarcoding), from all sizes (pico‐ to microorganisms in opposition to light microscopy).

These results are of essential interest for water managers, and particularly in terms of the confidence that we can place in the metabarcoding approach to develop biotic indices based on phytoplankton composition for assessing the ecological status of lakes.

4.2. Several Factors Limit the Detection of Certain Phyla

Although the relative abundances of most phyla obtained with metabarcoding were similar to those obtained with shotgun, notable differences were observed for Chlorophyta and non‐diatom Ochrophyta, for both barcodes. For these two phyla, part of their taxonomic diversity could not be amplified using metabarcoding. This was further confirmed by several families observed in microscopy that were not detected in metabarcoding or under‐amplified as compared to shotgun data.

A first explanation for these discrepancies observed for Chlorophyta and non‐diatom Ochrophyta is that numerous of their taxa are absent in the reference databases. Several families of these two phyla have fewer than 10 representative sequences in Phytool 16S and 23S (e.g., within Chlorophyta, the family Selenastraceae has only reference sequences for 5 species in Phytool 23S—3 species in Phytool 16S—while having 115 described species in AlgaeBase). Incomplete reference databases are a known bottleneck to design robust and efficient primers as it increases the number of putative mismatches with sequences of the missing diversity (Taberlet et al. 2012). This can result in a fail to amplify some taxa or their under‐amplification, skewing their relative abundance (Liu et al. 2023). These limits will likely be reduced in the future with increasing number of representative sequences available in reference databases. As an example, Cyanobacteria is a very diverse group but well represented in the reference databases (above 100 sequences for most families in Phytool) which could explain the efficiency of both 16S and 23S to recover their diversity in our study.

A previous study comparing phytoplankton diversity in a hypersaline system also observed that Chlorophyta were less detected with 23S chloroplastic primers than with light microscopy (Brown et al. 2022). Another explanation for Chlorophyta could lie on the complexity of their chloroplastic genome which exhibits a highly variable structure, gene diversity and intron content between species (Turmel et al. 2015). Chloroplast genome in general presents a quadripartite structure, with two inverted repeat (IR) regions that harbour the rRNA genes operon (Green 2011). In Chlorophyta, these IR regions have been remodelled extensively during their evolution, resulting in the presence of large introns that can be inserted in different genes, including the 16S and 23S rRNA genes (Lemieux et al. 2014; Turmel et al. 2017). These introns size are from few hundred nucleotides to more than 2000 (Brouard et al. 2008). Thus, the presence of an intron in the targeted barcode can lead to a strict elimination of the amplicons during the sequencing step, as we used Illumina short‐reads sequencing that limits barcode size to 500 bp. Limits due to barcode length can be overcome using long‐read amplicon sequencing like Oxford Nanopore or PacBio technologies (Geisen et al. 2019; Egeter et al. 2022). These sequencing technologies are increasingly tested for metabarcoding studies of other micro‐eukaryotes and phytoplankton with success (Latz et al. 2022; Chang et al. 2024; Gaonkar and Campbell 2024). However, it would be important to test if these chloroplastic rRNA genes with introns are as successfully amplified as genes without introns.

4.3. Other PCR Biases

However, other steps of metabarcoding can also influence abundance and diversity estimates, which could not be addressed in this study. The variation in the number of copies of the targeted marker gene from one cell to another, can also lead to over‐ or under‐estimate the abundances of certain taxa. Chloroplastic genes have less drastic variations in copy number compared to nuclear rRNA genes (e.g., up to 6 chloroplastic genes copies in Pedinomonas minor, Decelle et al. 2015). This was confirmed by a previous study that compared 16S, 18S and microscopy‐based cell counts, and showed that the plastid marker gene reflected best the cell number as compared to the 18S for phytoplankton (Xu et al. 2023). However, the number of chloroplasts itself can be highly variable from a taxon to another, ranging from only one in Zygnema sp. to 100 in some centric diatoms (Decelle et al. 2015). A recent study suggested using psbO, a single copy nuclear gene coding for a PSII protein, to reliably assess phytoplankton cell abundances (Pierella Karlusich et al. 2023). We attempted to use this gene to estimate the importance of copy number bias in chloroplastic genes, but our dataset did not allow for sufficient gene reconstruction (n = 37), probably due to the taxonomic complexity of our samples, where eukaryotic reads were outnumbered by prokaryotic reads in the shotgun dataset. When mapping the shotgun reads against psbO gene reference database, we observed a bias for some eukaryotes.

5. Conclusion

In this study, we investigated the impact of different marker genes and the PCR on the assessment of community structure of phytoplankton obtained by metabarcoding in two alpine lakes. Despite their great advantage (i.e., enabling cost‐effective assessment of both prokaryotic and eukaryotic phytoplankton diversity), chloroplastic genes remain underutilised compared to 18S rRNA for phytoplankton diversity. By comparing the relative abundance of taxa obtained with shotgun data, we demonstrated that the 16S and 23S primers reliably capture phytoplankton community structure in the two alpine lakes. However, our study also highlight the low representation of freshwater eukaryotic phytoplankton (e.g., non‐diatom Ochrophyta and Chlorophyta) in molecular databases (chloroplastic rRNA gene databases and psbO gene reference database).

Author Contributions

Frédéric Rimet and Clarisse Lemonnier designed the project. Clarisse Lemonnier conducted the study and performed the primary data analysis and visualisation. Benjamin Alric contributed on analytical tools. Clarisse Lemonnier wrote the manuscript with substantial input from all the authors.

Disclosure

Benefit‐Sharing Statement: Benefit from this research accrue from the sharing of our data and scripts.

Conflicts of Interest

The authors declare no conflicts of interest.

Supporting information

Data S1–S6. men70064‐sup‐0001‐DataS1–S6.docx.

MEN-26-e70077-s001.docx (1.4MB, docx)

Acknowledgements

This study was carried out in the framework of the project PhytoDOM, funded by OFB and INRAE, and developed at the Pôle R&D ECLA (ECosystèmes LAcustres). Clarisse Lemonnier was financed in the framework of the European Union H2020 BIOLAWEB project. We are grateful to the genotoul bioinformatics platform Toulouse Occitanie (Bioinfo Genotoul, https://doi.org/10.15454/1.5572369328961167E12) for providing computing and storage resources.

Lemonnier, C. , Alric B., Domaizon I., and Rimet F.. 2026. “Comparison of Metabarcoding and Shotgun Sequencing Confirms the Relevance of Chloroplastic rRNA Genes to Assess Community Structure of Lake Phytoplankton.” Molecular Ecology Resources 26, no. 1: e70077. 10.1111/1755-0998.70077.

Handling Editor: Brent Emerson

Funding: This study was in part carried out in the framework of the project PhytoDOM, funded by OFB and INRAE.

Contributor Information

Clarisse Lemonnier, Email: clarisse.lemonnier@inrae.fr.

Frédéric Rimet, Email: frederic.rimet@inrae.fr.

Data Availability Statement

Data used or generated in this study were deposited in ENA under different project names: PRJEB88020 for phytoplankton 16S metabarcoding; PRJEB88073 for phytoplankton 23S metabarcoding; and PRJEB88169 for shotgun data. Scripts used for the analysis are available here: https://github.com/clarilemon/Comparison‐shotgun‐metabarcoding.

References

  1. Abad, D. , Albaina A., Aguirre M., et al. 2016. “Is Metabarcoding Suitable for Estuarine Plankton Monitoring? A Comparative Study With Microscopy.” Marine Biology 163: 1–13. [Google Scholar]
  2. Acinas, S. G. , Sarma‐Rupavtarm R., Klepac‐Ceraj V., and Polz M. F.. 2005. “PCR‐Induced Sequence Artifacts and Bias: Insights From Comparison of Two 16S rRNA Clone Libraries Constructed From the Same Sample.” Applied and Environmental Microbiology 71, no. 12: 8966–8969. [DOI] [PMC free article] [PubMed] [Google Scholar]
  3. Aitchison, J. 1986. The Statistical Analysis of Compositional Data. Monographs on Statistics and Applied Probability. Chapman & Hall Ltd. [Google Scholar]
  4. Andersson, A. , Zhao L., Brugel S., Figueroa D., and Huseby S.. 2023. “Metabarcoding vs Microscopy: Comparison of Methods to Monitor Phytoplankton Communities.” ACS ES&T Water 3, no. 8: 2671–2680. [Google Scholar]
  5. Anneville, O. , Dur G., Rimet F., and Souissi S.. 2018. “Plasticity in Phytoplankton Annual Periodicity: An Adaptation to Long‐Term Environmental Changes.” Hydrobiologia 824, no. 1: 121–141. [Google Scholar]
  6. Bachy, C. , Dolan J. R., López‐García P., Deschamps P., and Moreira D.. 2013. “Accuracy of Protist Diversity Assessments: Morphology Compared With Cloning and Direct Pyrosequencing of 18S rRNA Genes and ITS Regions Using the Conspicuous Tintinnid Ciliates as a Case Study.” ISME Journal 7, no. 2: 244–255. [DOI] [PMC free article] [PubMed] [Google Scholar]
  7. Barbrook, A. C. , Howe C. J., and Nisbet R. E. R.. 2019. “Breaking Up Is Hard to Do: The Complexity of the Dinoflagellate Chloroplast Genome.” Perspectives in Phycology 6: 31–37. [Google Scholar]
  8. Behnke, A. , Engel M., Christen R., Nebel M., Klein R. R., and Stoeck T.. 2011. “Depicting More Accurate Pictures of Protistan Community Complexity Using Pyrosequencing of Hypervariable SSU rRNA Gene Regions.” Environmental Microbiology 13, no. 2: 340–349. [DOI] [PubMed] [Google Scholar]
  9. Bergkemper, V. , and Weisse T.. 2018. “Do Current European Lake Monitoring Programmes Reliably Estimate Phytoplankton Community Changes?” Hydrobiologia 824, no. 1: 143–162. [Google Scholar]
  10. Blancher, P. , Lefrançois E., Rimet F., et al. 2022. “A Strategy for Successful Integration of DNA‐Based Methods in Aquatic Monitoring.” Metabarcoding and Metagenomics 6: 215–226. [Google Scholar]
  11. Borics, G. , Abonyi A., Salmaso N., and Ptacnik R.. 2021. “Freshwater Phytoplankton Diversity: Models, Drivers and Implications for Ecosystem Properties.” Hydrobiologia 848: 53–75. [DOI] [PMC free article] [PubMed] [Google Scholar]
  12. Bradley, I. M. , Pinto A. J., and Guest J. S.. 2016. “Design and Evaluation of Illumina MiSeq‐Compatible, 18S rRNA Gene‐Specific Primers for Improved Characterization of Mixed Phototrophic Communities.” Applied and Environmental Microbiology 82, no. 19: 5878–5891. [DOI] [PMC free article] [PubMed] [Google Scholar]
  13. Breitwieser, F. P. , Lu J., and Salzberg S. L.. 2019. “A Review of Methods and Databases for Metagenomic Classification and Assembly.” Briefings in Bioinformatics 20, no. 4: 1125–1136. [DOI] [PMC free article] [PubMed] [Google Scholar]
  14. Brouard, J.‐S. , Otis C., Lemieux C., and Turmel M.. 2008. “Chloroplast DNA Sequence of the Green Alga Oedogonium cardiacum (Chlorophyceae): Unique Genome Architecture, Derived Characters Shared With the Chaetophorales and Novel Genes Acquired Through Horizontal Transfer.” BMC Genomics 9: 1–16. [DOI] [PMC free article] [PubMed] [Google Scholar]
  15. Brown, P. D. , Craine J. M., Richards D., Chapman A., and Marden B.. 2022. “DNA Metabarcoding of the Phytoplankton of Great Salt Lake's Gilbert Bay: Spatiotemporal Assemblage Changes and Comparisons to Microscopy.” Journal of Great Lakes Research 48, no. 1: 110–124. [Google Scholar]
  16. Callahan, B. J. , McMurdie P. J., Rosen M. J., Han A. W., Johnson A. J. A., and Holmes S. P.. 2016. “DADA2: High‐Resolution Sample Inference From Illumina Amplicon Data.” Nature Methods 13, no. 7: 581–583. [DOI] [PMC free article] [PubMed] [Google Scholar]
  17. Callieri, C. , Bertoni R., Amicucci E., and Pinolini M. L.. 1995. “Picoplankton Composition and Size Frequency Distribution in Lago Maggiore, N. Italy.” Memorie‐Istituto Italiano Di Idrobiologia Dott Marco De Marchi 53: 177–190. [Google Scholar]
  18. Camacho, C. , Coulouris G., Avagyan V., et al. 2009. “BLAST+: Architecture and Applications.” BMC Bioinformatics 10: 1–9. [DOI] [PMC free article] [PubMed] [Google Scholar]
  19. Canino, A. , Bouchez A., Laplace‐Treyture C., Domaizon I., and Rimet F.. 2021. “Phytool, a ShinyApp to Homogenise Taxonomy of Freshwater Microalgae From DNA Barcodes and Microscopic Observations.” Metabarcoding and Metagenomics 5: e74096. [Google Scholar]
  20. Canino, A. , Lemonnier C., Alric B., et al. 2023. “Which Barcode to Decipher Freshwater Microalgal Assemblages? Tests on Mock Communities.” International Journal of Limnology 59: 8. [Google Scholar]
  21. Chakoory, O. , Comtet‐Marre S., and Peyret P.. 2022. “RiboTaxa: Combined Approaches for rRNA Genes Taxonomic Resolution Down to the Species Level From Metagenomics Data Revealing Novelties.” NAR Genomics and Bioinformatics 4, no. 3: lqac070. [DOI] [PMC free article] [PubMed] [Google Scholar]
  22. Chang, J. J. M. , Ip Y. C. A., Neo W. L., Mowe M. A. D., Jaafar Z., and Huang D.. 2024. “Primed and Ready: Nanopore Metabarcoding Can Now Recover Highly Accurate Consensus Barcodes That Are Generally Indel‐Free.” BMC Genomics 25, no. 1: 842. [DOI] [PMC free article] [PubMed] [Google Scholar]
  23. Cole, J. J. , Prairie Y. T., Caraco N. F., et al. 2007. “Plumbing the Global Carbon Cycle: Integrating Inland Waters Into the Terrestrial Carbon Budget.” Ecosystems 10: 172–185. [Google Scholar]
  24. De Vargas, C. , Audic S., Henry N., et al. 2015. “Eukaryotic Plankton Diversity in the Sunlit Ocean.” Science 348, no. 6237: 1261605. [DOI] [PubMed] [Google Scholar]
  25. Decelle, J. , Romac S., Stern R. F., et al. 2015. “Phyto REF: A Reference Database of the Plastidial 16S rRNA Gene of Photosynthetic Eukaryotes With Curated Taxonomy.” Molecular Ecology Resources 15, no. 6: 1435–1445. [DOI] [PubMed] [Google Scholar]
  26. Decelle, J. , Stryhanyuk H., Gallet B., et al. 2019. “Algal Remodeling in a Ubiquitous Planktonic Photosymbiosis.” Current Biology 29, no. 6: 968–978. 10.1016/j.cub.2019.01.073. [DOI] [PubMed] [Google Scholar]
  27. Dray, S. , Dufour A. B., and Chessel D.. 2007. “The ade4 Package‐II: Two‐Table and K‐Table Methods.” R News 7, no. 2: 47–52. [Google Scholar]
  28. Egeter, B. , Verissimo J., Lopes‐Lima M., et al. 2022. “Speeding Up the Detection of Invasive Bivalve Species Using Environmental DNA: A Nanopore and Illumina Sequencing Comparison.” Molecular Ecology Resources 22, no. 6: 2232–2247. [DOI] [PubMed] [Google Scholar]
  29. Eren, A. M. , Esen Ö. C., Quince C., et al. 2015. “Anvi'o: An Advanced Analysis and Visualization Platform for ‘Omics Data’.” PeerJ 3: e1319. [DOI] [PMC free article] [PubMed] [Google Scholar]
  30. European Commission . 2000. “Directive 2000/60/EC of the European Parliament and of the Council of 23rd October 2000 Establishing a Framework for Community Action in the Field of Water Policy.” Official Journal of the European Communities 327: 1–72. [Google Scholar]
  31. Gao, Y. , Sun L., Wu C., et al. 2018. “Inter‐Annual and Seasonal Variations of Phytoplankton Community and Its Relation to Water Pollution in Futian Mangrove of Shenzhen, China.” Continental Shelf Research 166: 138–147. [Google Scholar]
  32. Gaonkar, C. C. , and Campbell L.. 2024. “A Full‐Length 18S Ribosomal DNA Metabarcoding Approach for Determining Protist Community Diversity Using Nanopore Sequencing.” Ecology and Evolution 14, no. 4: e11232. [DOI] [PMC free article] [PubMed] [Google Scholar]
  33. Geisen, S. , Laros I., Vizcaíno A., Bonkowski M., and De Groot G. A.. 2015. “Not All Are Free‐Living: High‐Throughput DNA Metabarcoding Reveals a Diverse Community of Protists Parasitizing Soil Metazoa.” Molecular Ecology 24, no. 17: 4556–4569. [DOI] [PubMed] [Google Scholar]
  34. Geisen, S. , Vaulot D., Mahé F., Lara E., de Vargas C., and Bass D.. 2019. “A User Guide to Environmental Protistology: Primers, Metabarcoding, Sequencing, and Analyses.” bioRxiv. 850610.
  35. Gleick, P. H. 1993. Water in Crisis. Vol. 82. Oxford University Press. [Google Scholar]
  36. Gonzalez, J. M. , Portillo M. C., Belda‐Ferre P., and Mira A.. 2012. “Amplification by PCR Artificially Reduces the Proportion of the Rare Biosphere in Microbial Communities.” PLoS One 7, no. 1: e29973. [DOI] [PMC free article] [PubMed] [Google Scholar]
  37. Green, B. R. 2011. “Chloroplast Genomes of Photosynthetic Eukaryotes.” Plant Journal 66, no. 1: 34–44. [DOI] [PubMed] [Google Scholar]
  38. Gruber‐Vodicka, H. R. , Seah B. K. B., and Pruesse E.. 2020. “phyloFlash: Rapid Small‐Subunit rRNA Profiling and Targeted Assembly From Metagenomes.” mSystems 5, no. 5: 10–128. [DOI] [PMC free article] [PubMed] [Google Scholar]
  39. Howe, C. J. , Nisbet R. E. R., and Barbrook A. C.. 2008. “The Remarkable Chloroplast Genome of Dinoflagellates.” Journal of Experimental Botany 59, no. 5: 1035–1045. [DOI] [PubMed] [Google Scholar]
  40. Huo, S. , Li X., Xi B., Zhang H., Ma C., and He Z.. 2020. “Combining Morphological and Metabarcoding Approaches Reveals the Freshwater Eukaryotic Phytoplankton Community.” Environmental Sciences Europe 32: 1–14. [Google Scholar]
  41. Interlandi, S. J. , and Kilham S. S.. 2001. “Limiting Resources and the Regulation of Diversity in Phytoplankton Communities.” Ecology 82, no. 5: 1270–1282. [Google Scholar]
  42. Kahlert, M. , Albert R.‐L., Anttila E.‐L., et al. 2009. “Harmonization Is More Important Than Experience—Results of the First Nordic–Baltic Diatom Intercalibration Exercise 2007 (Stream Monitoring).” Journal of Applied Phycology 21: 471–482. [Google Scholar]
  43. Kelly, M. G. , Juggins S., Mann D. G., et al. 2020. “Development of a Novel Metric for Evaluating Diatom Assemblages in Rivers Using DNA Metabarcoding.” Ecological Indicators 118: 106725. [Google Scholar]
  44. Kopylova, E. , Noé L., and Touzet H.. 2012. “SortMeRNA: Fast and Accurate Filtering of Ribosomal RNAs in Metatranscriptomic Data.” Bioinformatics 28, no. 24: 3211–3217. [DOI] [PubMed] [Google Scholar]
  45. Langmead, B. , and Salzberg S. L.. 2012. “Fast Gapped‐Read Alignment With Bowtie 2.” Nature Methods 9, no. 4: 357–359. [DOI] [PMC free article] [PubMed] [Google Scholar]
  46. Latz, M. A. C. , Grujcic V., Brugel S., et al. 2022. “Short‐ and Long‐Read Metabarcoding of the Eukaryotic rRNA Operon: Evaluation of Primers and Comparison to Shotgun Metagenomics Sequencing.” Molecular Ecology Resources 22, no. 6: 2304–2318. [DOI] [PubMed] [Google Scholar]
  47. Lemieux, C. , Otis C., and Turmel M.. 2014. “Six Newly Sequenced Chloroplast Genomes From Prasinophyte Green Algae Provide Insights Into the Relationships Among Prasinophyte Lineages and the Diversity of Streamlined Genome Architecture in Picoplanktonic Species.” BMC Genomics 15: 1–20. [DOI] [PMC free article] [PubMed] [Google Scholar]
  48. Lind, A. L. , and Pollard K. S.. 2021. “Accurate and Sensitive Detection of Microbial Eukaryotes From Whole Metagenome Shotgun Sequencing.” Microbiome 9: 1–18. [DOI] [PMC free article] [PubMed] [Google Scholar]
  49. Liu, M. , Burridge C. P., Clarke L. J., Baker S. C., and Jordan G. J.. 2023. “Does Phylogeny Explain Bias in Quantitative DNA Metabarcoding?” Metabarcoding and Metagenomics 7: e101266. [Google Scholar]
  50. Logares, R. , Sunagawa S., Salazar G., et al. 2014. “Metagenomic 16S rDNA I Llumina Tags Are a Powerful Alternative to Amplicon Sequencing to Explore Diversity and Structure of Microbial Communities.” Environmental Microbiology 16, no. 9: 2659–2671. [DOI] [PubMed] [Google Scholar]
  51. Marinchel, N. , Marchesini A., Nardi D., et al. 2023. “Mock Community Experiments Can Inform on the Reliability of eDNA Metabarcoding Data: A Case Study on Marine Phytoplankton.” Scientific Reports 13, no. 1: 20164. [DOI] [PMC free article] [PubMed] [Google Scholar]
  52. Martin, J. L. , Santi I., Pitta P., John U., and Gypens N.. 2022. “Towards Quantitative Metabarcoding of Eukaryotic Plankton: An Approach to Improve 18S rRNA Gene Copy Number Bias.” Metabarcoding and Metagenomics 6: e85794. [Google Scholar]
  53. Martin, M. 2011. “Cutadapt Removes Adapter Sequences From High‐Throughput Sequencing Reads.” EMBnet.Journal 17, no. 1: 10–12. [Google Scholar]
  54. Mikhailov, I. S. , Yurij S. B., Firsova A. D., Petrova D. P., and Likhoshway Y. V.. 2025. “Comparison of Relative and Absolute Abundance and Biomass of Freshwater Phytoplankton Taxa Using Metabarcoding and Microscopy.” Ecology and Evolution 15, no. 3: e70856. [DOI] [PMC free article] [PubMed] [Google Scholar]
  55. Milanese, A. , Mende D. R., Paoli L., et al. 2019. “Microbial Abundance, Activity and Population Genomic Profiling With mOTUs2.” Nature Communications 10, no. 1: 1014. [DOI] [PMC free article] [PubMed] [Google Scholar]
  56. Minoche, A. E. , Dohm J. C., and Himmelbauer H.. 2011. “Evaluation of Genomic High‐Throughput Sequencing Data Generated on Illumina HiSeq and Genome Analyzer Systems.” Genome Biology 12: 1–15. [DOI] [PMC free article] [PubMed] [Google Scholar]
  57. Nicolosi Gelis, M. M. , Canino A., Bouchez A., et al. 2024. “Assessing the Relevance of DNA Metabarcoding Compared to Morphological Identification for Lake Phytoplankton Monitoring.” Science of the Total Environment 914: 169774. [DOI] [PubMed] [Google Scholar]
  58. Nübel, U. , Garcia‐Pichel F., and Muyzer G.. 1997. “PCR Primers to Amplify 16S rRNA Genes From Cyanobacteria.” Applied and Environmental Microbiology 63, no. 8: 3327–3332. [DOI] [PMC free article] [PubMed] [Google Scholar]
  59. Obiol, A. , Giner C. R., Sánchez P., Duarte C. M., Acinas S. G., and Massana R.. 2020. “A Metagenomic Assessment of Microbial Eukaryotic Diversity in the Global Ocean.” Molecular Ecology Resources 20, no. 3: 718–731. [DOI] [PubMed] [Google Scholar]
  60. Oksanen, J. , Blanchet F. G., Kindt R., et al. 2013. “Package ‘Vegan’.” Community Ecology Package, Version 2, 9: 1–295.
  61. Pawlowski, J. , Kelly‐Quinn M., Altermatt F., et al. 2018. “The Future of Biotic Indices in the Ecogenomic Era: Integrating (e) DNA Metabarcoding in Biological Assessment of Aquatic Ecosystems.” Science of the Total Environment 637: 1295–1310. [DOI] [PubMed] [Google Scholar]
  62. Pernice, M. C. , Giner C. R., Logares R., et al. 2016. “Large Variability of Bathypelagic Microbial Eukaryotic Communities Across the World's Oceans.” ISME Journal 10, no. 4: 945–958. [DOI] [PMC free article] [PubMed] [Google Scholar]
  63. Personnic, S. , Domaizon I., Dorigo U., Berdjeb L., and Jacquet S.. 2009. “Seasonal and Spatial Variability of Virio‐, Bacterio‐, and Picophytoplanktonic Abundances in Three Peri‐Alpine Lakes.” Hydrobiologia 627, no. 1: 99–116. [Google Scholar]
  64. Pierella Karlusich, J. J. , Pelletier E., Zinger L., et al. 2023. “A Robust Approach to Estimate Relative Phytoplankton Cell Abundances From Metagenomes.” Molecular Ecology Resources 23, no. 1: 16–40. [DOI] [PMC free article] [PubMed] [Google Scholar]
  65. Pomati, F. , Jokela J., Simona M., Veronesi M., and Ibelings B. W.. 2011. “An Automated Platform for Phytoplankton Ecology and Aquatic Ecosystem Monitoring.” Environmental Science and Technology 45, no. 22: 9658–9665. [DOI] [PubMed] [Google Scholar]
  66. Prodan, A. , Tremaroli V., Brolin H., Zwinderman A. H., Nieuwdorp M., and Levin E.. 2020. “Comparing Bioinformatic Pipelines for Microbial 16S rRNA Amplicon Sequencing.” PLoS One 15, no. 1: e0227434. [DOI] [PMC free article] [PubMed] [Google Scholar]
  67. Prygiel, J. , Carpentier P., Almeida S., et al. 2002. “Determination of the Biological Diatom Index (IBD NF T 90–354): Results of an Intercomparison Exercise.” Journal of Applied Phycology 14: 27–39. [Google Scholar]
  68. Quast, C. , Pruesse E., Yilmaz P., et al. 2012. “The SILVA Ribosomal RNA Gene Database Project: Improved Data Processing and Web‐Based Tools.” Nucleic Acids Research 41, no. D1: D590–D596. [DOI] [PMC free article] [PubMed] [Google Scholar]
  69. Reynolds, C. S. 1987. “The Response of Phytoplankton Communities to Changing Lake Environments.” Swiss Journal of Hydrology 49: 220–236. [Google Scholar]
  70. Rieder, J. , Kapopoulou A., Bank C., and Adrian‐Kalchhauser I.. 2023. “Metagenomics and Metabarcoding: Experimental Choices and Their Impact on Microbial Community Characterization in Freshwater Recirculating Aquaculture Systems.” Environmental Microbiomes 18, no. 1: 8. [DOI] [PMC free article] [PubMed] [Google Scholar]
  71. Rimet, F. , Anneville O., Barbet D., et al. 2020. “The Observatory on LAkes (OLA) Database: Sixty Years of Environmental Data Accessible to the Public.” Journal of Limnology 79, no. 2: 164–178. [Google Scholar]
  72. Salmaso, N. 2010. “Long‐Term Phytoplankton Community Changes in a Deep Subalpine Lake: Responses to Nutrient Availability and Climatic Fluctuations.” Freshwater Biology 55, no. 4: 825–846. [Google Scholar]
  73. Santi, I. , Kasapidis P., Karakassis I., and Pitta P.. 2021. “A Comparison of DNA Metabarcoding and Microscopy Methodologies for the Study of Aquatic Microbial Eukaryotes.” Diversity 13, no. 5: 180. [Google Scholar]
  74. Santoferrara, L. F. 2019. “Current Practice in Plankton Metabarcoding: Optimization and Error Management.” Journal of Plankton Research 41, no. 5: 571–582. [Google Scholar]
  75. Smayda, T. J. 2011. “Cryptic Planktonic Diatom Challenges Phytoplankton Ecologists.” Proceedings of the National Academy of Sciences 108, no. 11: 4269–4270. [DOI] [PMC free article] [PubMed] [Google Scholar]
  76. Sommeria‐Klein, G. , Watteaux R., Ibarbalz F. M., et al. 2021. “Global Drivers of Eukaryotic Plankton Biogeography in the Sunlit Ocean.” Science 374, no. 6567: 594–599. [DOI] [PubMed] [Google Scholar]
  77. Sterner, R. W. , Keeler B., Polasky S., Poudel R., Rhude K., and Rogers M.. 2020. “Ecosystem Services of Earth's Largest Freshwater Lakes.” Ecosystem Services 41: 101046. [Google Scholar]
  78. Stoeck, T. , Bass D., Nebel M., et al. 2010. “Multiple Marker Parallel Tag Environmental DNA Sequencing Reveals a Highly Complex Eukaryotic Community in Marine Anoxic Water.” Molecular Ecology 19: 21–31. [DOI] [PubMed] [Google Scholar]
  79. Stothart, M. R. , McLoughlin P. D., and Poissant J.. 2023. “Shallow Shotgun Sequencing of the Microbiome Recapitulates 16S Amplicon Results and Provides Functional Insights.” Molecular Ecology Resources 23, no. 3: 549–564. [DOI] [PubMed] [Google Scholar]
  80. Taberlet, P. , Coissac E., Pompanon F., Brochmann C., and Willerslev E.. 2012. “Towards Next‐Generation Biodiversity Assessment Using DNA Metabarcoding.” Molecular Ecology 21, no. 8: 2045–2050. [DOI] [PubMed] [Google Scholar]
  81. Tedersoo, L. , Anslan S., Bahram M., et al. 2015. “Shotgun Metagenomes and Multiple Primer Pair‐Barcode Combinations of Amplicons Reveal Biases in Metabarcoding Analyses of Fungi.” MycoKeys 10: 1–43. [Google Scholar]
  82. Tessler, M. , Neumann J. S., Afshinnekoo E., et al. 2017. “Large‐Scale Differences in Microbial Biodiversity Discovery Between 16S Amplicon and Shotgun Sequencing.” Scientific Reports 7, no. 1: 6589. 10.1038/s41598-017-06665-3. [DOI] [PMC free article] [PubMed] [Google Scholar]
  83. Turmel, M. , Otis C., and Lemieux C.. 2015. “Dynamic Evolution of the Chloroplast Genome in the Green Algal Classes Pedinophyceae and Trebouxiophyceae.” Genome Biology and Evolution 7, no. 7: 2062–2082. [DOI] [PMC free article] [PubMed] [Google Scholar]
  84. Turmel, M. , Otis C., and Lemieux C.. 2017. “Divergent Copies of the Large Inverted Repeat in the Chloroplast Genomes of Ulvophycean Green Algae.” Scientific Reports 7, no. 1: 994. [DOI] [PMC free article] [PubMed] [Google Scholar]
  85. Utermöhl, H. 1958. “Zur Vervollkommnung der Quantitativen Phytoplankton‐Methodik: Mit 1 Tabelle Und 15 Abbildungen im Text Und Auf 1 Tafel.” SIL Communications, 1953–1996 9, no. 1: 1–38. [Google Scholar]
  86. van der Loos, L. M. , and Nijland R.. 2021. “Biases in Bulk: DNA Metabarcoding of Marine Communities and the Methodology Involved.” Molecular Ecology 30, no. 13: 3270–3288. [DOI] [PMC free article] [PubMed] [Google Scholar]
  87. Vasselon, V. , Bouchez A., Rimet F., et al. 2018. “Avoiding Quantification Bias in Metabarcoding: Application of a Cell Biovolume Correction Factor in Diatom Molecular Biomonitoring.” Methods in Ecology and Evolution 9, no. 4: 1060–1069. 10.1111/2041-210X.12960. [DOI] [Google Scholar]
  88. Vasselon, V. , Rimet F., Domaizon I., Monnier O., Reyjol Y., and Bouchez A.. 2019. “Assessing Pollution of Aquatic Environments With Diatoms' DNA Metabarcoding: Experience and Developments From France Water Framework Directive Networks.” Metabarcoding and Metagenomics 3: 39646. [Google Scholar]
  89. Wang, Q. , Garrity G. M., Tiedje J. M., and Cole J. R.. 2007. “Naive Bayesian Classifier for Rapid Assignment of rRNA Sequences Into the New Bacterial Taxonomy.” Applied and Environmental Microbiology 73, no. 16: 5261–5267. [DOI] [PMC free article] [PubMed] [Google Scholar]
  90. Wheeler, T. J. , and Eddy S. R.. 2013. “Nhmmer: DNA Homology Search With Profile HMMs.” Bioinformatics 29, no. 19: 2487–2489. [DOI] [PMC free article] [PubMed] [Google Scholar]
  91. Xu, S. , Li G., He C., et al. 2023. “Diversity, Community Structure, and Quantity of Eukaryotic Phytoplankton Revealed Using 18S rRNA and Plastid 16S rRNA Genes and Pigment Markers: A Case Study of the Pearl River Estuary.” Marine Life Science & Technology 5, no. 3: 415–430. [DOI] [PMC free article] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Data S1–S6. men70064‐sup‐0001‐DataS1–S6.docx.

MEN-26-e70077-s001.docx (1.4MB, docx)

Data Availability Statement

Data used or generated in this study were deposited in ENA under different project names: PRJEB88020 for phytoplankton 16S metabarcoding; PRJEB88073 for phytoplankton 23S metabarcoding; and PRJEB88169 for shotgun data. Scripts used for the analysis are available here: https://github.com/clarilemon/Comparison‐shotgun‐metabarcoding.


Articles from Molecular Ecology Resources are provided here courtesy of Wiley

RESOURCES