ABSTRACT
A detailed characterization of the microbial ecosystem involved in the production processes of fermented foods is essential. Although fermented foods are an important part of the human diet and have seen an increasing interest nowadays, some challenges still need to be solved. Specifically, yeast identification through culture-independent methodologies is still limited to the genus level. Unlike bacterial species identifications, long-read sequencing technologies have barely been used for yeast species identification, and, to the best of the authors’ knowledge, it has not been validated with mock communities reflecting food fermentation processes yet. Therefore, in the current study, we present an amplicon-based metabarcoding approach targeting the full-length internal transcribed spacer (ITS) region comprising ITS1, the 5.8S rRNA gene, and ITS2, using the PacBio HiFi sequencing platform. This method was validated using DNA-based mock communities composed of yeast species involved in sourdough, lambic beer, and cocoa fermentation processes. Accurate species-level identification was achieved for most of the species. However, special attention should be given to Saccharomyces-rich niches, as accurate species-level identification for this genus is still challenging. Furthermore, underestimation of the relative abundance of species with short ITS regions, such as Pichia and Brettanomyces, occurred. In addition, the method was successfully applied to describe the yeast diversity present in two sourdough and two lambic beer samples. Overall, the current method provides an unprecedented way of determining the species-level yeast composition of complex ecosystems present in fermented food products.
IMPORTANCE
To date, species-level identification of common yeasts present in food fermentation ecosystems has been difficult, if not impossible, when using short-read sequencing methods. However, species-level identification is essential when evaluating and describing the characteristics of fermented food microbiomes. The current study reports on the development and validation of an amplicon-based metabarcoding approach combined with long-read PacBio HiFi sequencing targeting the full internal transcribed spacer (ITS) region, comprising the ITS1 and ITS2 regions, as well as the 5.8S rRNA gene. The described methodology enables species-level identification of the most common yeasts present in food fermentation ecosystems. This new methodology provides an important tool not only for the investigation of fermented foods but also for other fields engaged in complex microbial community analysis.
KEYWORDS: metagenetics, yeasts, food fermentation, PacBio, ITS
INTRODUCTION
Fermented foods and beverages have been an important part of the human diet for centuries and have seen an increasing interest nowadays (1). As a result of the great importance of fermented foods and beverages, studying the microbial ecosystem composition involved in their production processes in detail is essential. Fermented foods are produced through the metabolic activity of yeasts, lactic acid bacteria (LAB), coagulase-negative cocci, acetic acid bacteria, and/or filamentous fungi (1–3). Among the microorganisms involved, yeasts are key microorganisms during the production of sourdough, beer, and wine, among others, as well as during cocoa fermentation (2, 4).
Yeast diversity in food fermentation processes has traditionally been studied by culture-dependent approaches in the same way as is done for bacterial diversity (5). However, besides the benefit of being able to construct a collection of microbial strains for further investigation, those approaches are time-consuming and are biased by the selective growth media and conditions used (5, 6). The advances in high-throughput sequencing technologies allow the study of microbial communities using a culture-independent approach (5, 7). Specifically, amplicon-based metabarcoding, also known as metagenetics, has been the gold standard in this field, as it has a high throughput with relatively low costs (7–11). This technique consists of amplifying a specific genomic region expected to be present in all genomes of the microorganisms present in the ecosystem, followed by sequencing the amplicons obtained (7). Therefore, selecting an appropriate genomic region that enables a good taxonomic resolution is of crucial importance. Furthermore, selecting a proper sequencing technology is also of importance, as the technology chosen can also impose certain constraints. For example, Illumina MiSeq and Illumina NovaSeq are the most commonly used technologies, but they only allow for sequencing of short reads, despite their high accuracy. This short length results in a taxonomical resolution that seldom reaches the species level (10, 12). Long-read sequencing technologies, such as Pacific Biosciences’ (PacBio) HiFi long-read sequencing or Oxford Nanopore Technologies’ nanopore sequencing, are available on the market. The former one has emerged as a suitable technology for amplicon-based metabarcoding as it combines long reads with high accuracy (10, 13).
In the case of bacterial communities, the 16S rRNA gene is commonly considered to be the most suitable gene for identification (14, 15). Due to the length restriction imposed for Illumina sequencing, specific variable regions or combinations thereof, such as V1-V3, V4, and V3-V5, have been targeted (15–17). However, this only allowed for genus-level resolution (17, 18). Recently, PacBio’s HiFi sequencing allows the analysis of the full-length 16S rRNA gene (10, 15, 19, 20) and, therefore, a species-level resolution for most of the LAB commonly found during food fermentation. Specifically, it has been used to reveal the bacterial diversity of beer (21, 22), cheese (23, 24), and sourdough (25).
In the case of yeasts, there is no such gold standard. Variable regions of the 18S rRNA gene, the internal transcribed spacer 1 (ITS1) region, and the ITS2 region are some examples of the genomic regions targeted (9, 26–28). In all cases, the short length imposed by Illumina sequencing also leads to less accurate genus-level identification compared to long-read amplicon sequencing (29). In addition, comparisons of taxonomic assignment using different regions are available for specific samples, such as beer (27), wine (27), soil (9), human saliva (9), grape must (9), and bioaerosols (30). However, different microbial compositions are retrieved for different targeted regions, and the preferred region is environment-dependent. No consensus has been reached regarding which genomic region to target yet. Nevertheless, all studies pointed out the ITS regions as the primary fungal barcode because of a higher number of sequence variations (12, 31). Recently, PacBio HiFi sequencing of the genomic region comprising the ITS1, 5.8S rRNA gene, and ITS2 has been performed for metabarcoding of eukaryote-containing samples. Specifically, it has been used to describe the fungal diversity of coffee plant samples (32), wheat, maize, barley, and cover crop (clover, hairy vetch, oilseed radish, and fallow) samples (33), seawater samples (34), and soil samples (31, 35), with most studies focusing on filamentous fungi rather than yeasts. To the best of the authors’ knowledge, this method has barely been applied to the analysis of fermented food samples (21), nor has it been validated using mock communities reflecting food fermentation processes yet.
Hence, this study aimed to assess whether amplicon-based metabarcoding targeting the full ITS region, comprising the ITS1 and ITS2 regions as well as the 5.8S rRNA gene, enables species-level identification of the most common yeasts present in food fermentation ecosystems. After determining the appropriate primer sets and amplification conditions, the technique was tested using DNA-based mock communities (DbMC) representing the yeast composition of different fermentation stages of relevant fermented foods and beverages, namely sourdough, lambic beer, and cocoa. Finally, the fungal diversity of two sourdough samples and two lambic beer samples was assessed using this technique for validation.
RESULTS
Optimization of the PCR conditions: BITS-ITS4 vs ITS1-ITS4
PCR amplification using a general DbMC as DNA template and the BITS-ITS4 primer set (the BITS primer was extended with barcodes 04, 05, and 06; the ITS4 primer was extended with barcodes 19 and 20) resulted in no visible fragments on gel electrophoresis (data not shown). Hence, the primer combination BITS-ITS4 was deemed unsuccessful in amplifying the targeted microorganisms. On the contrary, the use of the ITS1-ITS4 primer set (the ITS1 primer was extended with barcodes 07, 08, and 09; the ITS4 primer was extended with barcodes 19 and 20) and the same DbMC as DNA template resulted in visually successful amplification. Specifically, the clearest bands were obtained when an annealing temperature of 52°C was used. The purification of the amplicons resulted in DNA concentrations between 4.9 and 49.0 ng/µL, with no relationship observed between the concentrations and primer-barcode combination or sample types. However, PCR amplifications involving the ITS1 primer extended with barcode 05 resulted in no amplification, most likely due to secondary structures formed between the primer and barcode sequences. Indeed, when checking the extended primer in OligoEvaluator, four possible moderate secondary structures were present involving the 3′ end of the primer, possibly preventing primer annealing.
Sequencing data
An average of 118,998 reads per sample was obtained with a minimum of 9,718 reads and a maximum of 348,099 reads (Table S3). Those reads had an average length of 763 nucleotides, with 90% of the reads being between 600 and 900 nucleotides, while 10% of the reads were shorter than 600 nucleotides. Primer removal, quality filtering, and denoising resulted in the removal of 11% of the reads, yielding an average of 105,470 reads per sample, with a minimum of 8,007 reads (Lambic beer A) and a maximum of 310,563 reads (DbMC lambic beer equal 1). Most of the removed reads (8% of the total reads) were eliminated during the quality-filtering step. After data processing, the average length of the reads was 722 nucleotides, with a minimum length of 405 and a maximum length of 842 nucleotides. Only 15% of the reads were shorter than 600 bp.
Rarefaction analysis showed that all DbMCs and fermented food samples were sequenced with a sufficient depth to reliably assess the full fungal diversity (Fig. S1). The two DbMCs (cocoa equal 1 and cocoa late fermentation stages 3) with the highest number of species retrieved (16 species) reached a plateau (within 85,121 and 42,981 reads, respectively; Fig. S1A). In the case of the Lambic beer A and B samples, the plateau was reached with 4,741 and 5,301 reads, respectively, representing a maximum of 17 and 6 species, respectively (Fig. S1B). In the case of the Sourdough A and B samples, the plateau was reached with 43,521 and 8,041 reads, respectively, representing a maximum of 30 and 33 species, respectively.
Taxonomic classification
Unidentifiable fungal genus reads
Taxonomical assignment at the species or genus level could not be achieved for all amplicon sequence variants (ASVs) obtained. The proportion of unidentifiable fungal genus reads (UFGRs) was small for the different DbMCs, with a maximal percentage of 0.2% of the total reads in the DbMC cocoa equal 3. The Lambic beer A and B samples displayed similarly low percentages of UFGR. On the contrary, the sourdough samples had large amounts of UFGR, with 16.7% and 60.8% in Sourdough A and B, respectively. Blastn searches of these ASVs revealed that 99.4% and 99.8% of the UFGR, respectively, aligned to the ITS region of Simplicia felix (87%–91% sequence identity and 85.6%–88.3% coverage). Furthermore, blastn searches of these ASVs against the genomes of the type material of Triticum aestivum (soft wheat) and Secale cereale (rye) resulted in alignments of the ASVs with the ITS regions of these two plant species, with 100% sequence identity and coverage. Therefore, the UFGRs were removed before calculating the relative abundances of fungal species in each sample.
DNA-based mock communities
A total of seven different DbMCs, representing common yeast species present in the microbial communities of various food fermentation processes, namely those of sourdough, lambic beer, and cocoa, were constructed (Table 1) and analyzed in triplicate. The results obtained for the three replicates of each DbMC were similar, showing good reproducibility of the PCR amplification and sequencing methodology employed (Fig. 1).
TABLE 1.
Yeast strains and composition of the DNA-based mock communities used in the studya
| Strain | Isolation source | DNA (%) | ||
|---|---|---|---|---|
| Sourdough equal | ||||
| Maudiozyma humilis VUB-H061-Y3 | Household sourdough (IMDO-VUB) | 14 | ||
| Maudiozyma saulgeensis VUB-H029-Y1 | Household sourdough (IMDO-VUB) | 14 | ||
| Monosporozyma unispora MUCL 51234 | Bakery sourdough (MUCL) | 14 | ||
| Naumovozyma castellii VUB-H022-Y10 | Household sourdough (IMDO-VUB) | 14 | ||
| Saccharomyces cerevisiae MUCL 51207 | Bakery sourdough (MUCL) | 14 | ||
| Pichia fermentans WM_6w_t18w_c_Y7 | Laboratory sourdough (25) | 14 | ||
| Wickerhamomyces anomalus MUCL 51207 | Bakery sourdough (MUCL) | 14 | ||
MUCL, Agro-food and Environmental Fungal Collection at the Belgian Coordinated Collections of Microorganisms; IMDO-VUB, Research Group of Industrial Microbiology and Food Biotechnology; and NA, not available.
Fig 1.
Microbial composition of the sourdough (A), lambic beer (B), and cocoa (C) DNA-based mock communities, depicting both the expected and obtained relative abundances. Sample numbers 1, 2, and 3 represent the different replicates.
Sourdough DbMC
All species included in the DbMC were retrieved, corresponding to a precision, recall, and F1 score of 1 (Fig. 1A and Table 2). The divergence between the expected composition and the obtained one was 0.27, as some species were found in relative abundances different from the expected ones. Specifically, the obtained relative abundance of Saccharomyces cerevisiae was very similar to the expected one (on average, 0.6 ± 0.7 percentage points [%pt] more). In the case of Maudiozyma humilis, Maudiozyma saulgeensis, and Naumovozyma castellii, their relative abundances were slightly below the expected ones (on average, 4.6 ± 0.2 %pt less). In contrast, Monosporozyma unispora and Wickerhamomyces anomalus were found at a higher relative abundance than the expected ones (on average 13.3 ± 0.8 %pt more). Finally, the relative abundance of Pichia fermentans was on average 0.8% ± 0.1%, almost 20 times lower than the expected one.
TABLE 2.
Metrics describing the deviations obtained from the expected composition of the DNA-based mock communitiesa
| DNA-based mock community | Divergence | False negatives | False positives | True positives | Precision | Recall | F1 score |
|---|---|---|---|---|---|---|---|
| Sourdough equal | 0.2743 ± 0.0051 | 0 | 0 | 7 | 1 | 1 | 1 |
| Lambic beer equal | 0.6915 ± 0.0006 | 4 | 1 | 5 | 0.83 | 0.56 | 0.67 |
| Lambic beer alcoholic fermentation | 0.5527 ± 0.0016 | 4 | 1 | 5 | 0.83 | 0.56 | 0.67 |
| Lambic beer maturation | 0.5073 ± 0.3595 | 4 | 1 | 5 | 0.83 | 0.56 | 0.67 |
| Cocoa equal | 0.5321 ± 0.0063 | 1 | 1 | 2 | 0.67 | 0.67 | 0.67 |
| Cocoa early fermentation stages | 0.8318 ± 0.0228 | 1 | 1 | 2 | 0.67 | 0.67 | 0.67 |
| Cocoa late fermentation stages | 0.3391 ± 0.0081 | 1 | 1 | 2 | 0.67 | 0.67 | 0.67 |
The divergence was calculated as the Bray-Curtis distances between the expected and the observed relative abundance. FN, false negatives, the number of expected taxa that were not recovered; FP, false positives, the number of recovered taxa that were not expected; TP, true positives, the number of recovered taxa that were expected; precision, TP/(TP + FP); recall, TP/(TP + FN); and F1 score, 2TP/(2TP + 2FP + FN).
Lambic beer DbMCs
The composition of the lambic beer DbMCs was determined with a precision of 0.83, a recall of 0.56, and an F1 score of 0.67 in all cases, and a divergence of, on average, 0.69, 0.55, and 0.51 in the case of the lambic beer equal, lambic beer alcoholic fermentation, and lambic beer maturation DbMCs, respectively (Table 2). A precision, recall, and F1 score lower than 1 could be explained by the fact that the species Saccharomyces eubayanus, Saccharomyces kudriavzevii, Saccharomyces pasteurianus, and Saccharomyces uvarum were not found in any of the lambic beer DbMCs (Fig. 1B). Remarkably, 10.0% ± 0.0%, 7.7% ± 0.1%, and 6.3% ± 0.2% of the reads of the lambic beer equal, lambic beer alcoholic fermentation, and lambic beer maturation DbMCs, respectively, were assigned to Saccharomyces sp. Blastn searches using the ASVs assigned to this taxonomic unit as a query resulted in the alignment of seven out of eight ASVs to S. kudriavzevii, with at least 98.4% sequence identity and 100% coverage. In contrast, the eighth ASV aligned to S. kudriavzevii and Saccharomyces paradoxus with a sequence identity of 95.8% and a coverage of 100%. As was the case for the sourdough DbMC, S. cerevisiae was found at a similar relative abundance to the expected ones in the lambic beer equal (on average 11.4% ± 0.1% vs 11.1%) and the lambic beer alcoholic fermentation DbMCs (on average 28.7% ± 1.4% vs 30.0%). However, this species was found at a relative abundance of 6.5% ± 0.4%, compared to the expected 3.3% in the lambic beer maturation DbMCs. Saccharomyces bayanus was found at a higher relative abundance than those expected in the lambic beer equal (70.1% ± 0.1% vs 11.1%), lambic beer alcoholic fermentation (57.4% ± 0.2% vs 10.0%), and lambic beer maturation (44.2% ± 3.6% vs 3.3%) DbMCs. Furthermore, the Brettanomyces species were all underrepresented in the three lambic beer DbMCs, with the case of Brettanomyces bruxellensis being the most remarkable one (relative abundances lower than 2.4% in all cases).
Cocoa DbMCs
The cocoa DbMCs were described with a precision of 0.67, a recall of 0.67, and an F1 score of 0.67 in all cases, with divergences of 0.53, 0.83, and 0.34 in the case of the cocoa equal, cocoa early fermentation stages, and cocoa late fermentation stages DbMCs, respectively (Table 2). These values were the result of incorrect identification of the Hanseniaspora species, leading to higher divergences for the DbMCs with higher expected relative abundances for Hanseniaspora. Specifically, reads from Hanseniaspora opuntiae were identified as Hanseniaspora uvarum. Nevertheless, its relative abundances were very similar to the expected ones in the cocoa equal (35.5% ± 1.7% vs 33.3%), cocoa early fermentation stage (75.5% ± 2.8 vs 80%), and cocoa late fermentation stage (10.6% ± 1.4% vs 10%) DbMCs. S. cerevisiae was found with a relative abundance 1.5 times higher than the expected one in all cocoa DbMCs (Fig. 1C). In contrast, Pichia kudriavzevii was found with a relative abundance that was 2.5, 1.4, and 2.1 times lower than the expected ones for the equal cocoa, cocoa early fermentation stage, and cocoa late fermentation stage DbMCs, respectively.
Fermented food samples
When applying the developed method to the fermented food samples, a high diversity of yeast species was retrieved (Fig. S1B). However, most of the species were present with a relative abundance lower than 0.5% (combined in the “Minorities < 0.5%” category; Fig. 2). The lambic beer samples, Lambic beer A and B, were both mostly inhabited by S. cerevisiae (relative abundance of 76.8% and 50.8%, respectively), S. bayanus (10.3% and 19.8%, respectively), and Brettanomyces anomalus (3.0% and 29.1%, respectively). The sourdough samples displayed different yeast compositions, with Sourdough A containing Maudiozyma pseudohumilis (90.1%), Debaryomyces prosopidis (3.0%), and S. cerevisiae (1.9%) and Sourdough B containing Ma. humilis (47.9%), D. prosopidis (27.7%), and S. cerevisiae (11.2%). In the sourdough samples, two filamentous fungi were retrieved at low relative abundances, namely, Alternaria sp. and Cladosporium herbarum.
Fig 2.
Fungal composition of the fermented food samples. Lambic beer A was taken after 5 days of barrel fermentation, and Lambic beer B was taken after 3 months. Sourdough A was a 3-year-old bakery rye sourdough, and Sourdough B was a 9-year-old bakery wheat sourdough. “Minorities < 0.5%” included species with a relative abundance < 0.5%.
DISCUSSION
In the present study, the potential of an amplicon-based metabarcoding method targeting the full ITS region to reveal the yeast diversity of fermented food samples was investigated. Up to now, culture-independent approaches revealing the yeast diversity in food fermentations have been performed using short-read sequencing, targeting a small region of the fungal rRNA operon, resulting in an identification that was limited to the genus level (22, 25, 27, 36–38). However, the contribution to food fermentation processes of different species within a genus might differ. Therefore, targeting a larger region that would allow species-level identification was needed in the field.
Thanks to PacBio’s long-read HiFi sequencing technology, in combination with PacBio’s Kinnex library construction strategy, longer amplicons can be sequenced in a cost-efficient manner, providing a higher level of taxonomic resolution compared to previous methods. While the PacBio-based method has already been used to sequence the full ITS region to investigate fungal diversity, those studies focused, in most cases, on filamentous fungi and not on yeasts (21, 31, 33–35). Whether this method would be applicable in the field of food fermentation, in which yeasts are key microorganisms (1–3), remained unclear.
As a first step, seven DNA-based mock communities were composed, representing different stages of sourdough, lambic beer, and cocoa fermentation processes. These DbMCs were composed of different species of the same genus, including B. anomalus, B. bruxellensis, Brettanomyces custersianus, Ma. humilis, Ma. saulgeensis, Mo. unispora, N. castelli, S. cerevisiae, P. fermentans, P. kudriavzevii, and W. anomalus. Overall, species-level identification was successfully obtained, with two exceptions, which align with current knowledge in the field. Hanseniaspora opuntiae was misidentified as H. uvarum due to their known close genetic relationship (39). Also, a limited resolution was observed within the Saccharomyces genus, also due to the close genetic relationship among species in this genus (40). Additionally, Saccharomyces species are also known to produce natural hybrids, which are frequently found in fermented foods such as beers (41). Therefore, when analyzing samples from Saccharomyces-rich niches, the results should be interpreted with caution, as the Saccharomyces complexity might not be fully captured. Moreover, further research could tackle the search for other genomic regions that allow for species-level identification of Saccharomyces and Hanseniaspora species and ideally for species of other genera relevant to fermented foods.
The lack of species-level resolution or species misidentifications, as for the specific taxa discussed above, could be solved by targeting longer regions, such as the full rRNA region, spanning the 18S rRNA gene, ITS1, the 5.8S rRNA gene, ITS2, and the 28S rRNA gene (42, 43). However, such an approach would imply other challenges. Although the length of the expected amplicon, 6 kb, fits perfectly within PacBio’s HiFi and Kinnex sequencing strategies, well-curated, publicly available databases containing the full operon are currently not available. This could be overcome by the construction of custom databases (42, 43). However, the need to construct such custom databases as well as perform regular updates could prevent this strategy from becoming a routine strategy in microbiology laboratories unspecialized in advanced bioinformatics. The launch of the Ribosomal Operon Database (44) might change this scenario, although a regular update of such a recent database cannot be taken for granted, which, given the regular reclassifications of yeasts in recent years, is crucial for accurate yeast identification (45–49).
In addition to accurate species-level identification, the quantitative estimation of the fungal composition, expressed as relative abundance, should be as accurate as possible to obtain a good representation of the real community. However, this is not always possible, as different biases play a role that are difficult (if not impossible) to avoid. In particular, the lower-than-expected relative abundances of Pichia species in the sourdough and cocoa DbMCs are in line with previous studies (22, 25). Due to the short ITS regions for both Pichia and Brettanomyces (approximately 400–500 bp compared to 700–800 bp for Saccharomyces), the underestimation of Brettanomyces species in the lambic beer DbMCs might have been caused by the same effects. However, further investigation is needed, as length biases do not seem to have a great impact on the relative abundances when targeting either the ITS1 or ITS2 region (12). In addition, the different copy numbers of the ITS region in different yeasts might have introduced a bias during the PCR amplification. As information on yeast rRNA region copy numbers remains scarce and seems to be strain-dependent (44), fungal community profiling should be considered as semi-quantitative, rather than quantitative. This highlights the need for long-read whole-genome sequences, curated and deposited in a public database, to enable accurate copy number determination at the species or strain level, moving toward a more quantitative profiling of fungal communities, irrespective of their habitat.
Given the promising results of the DbMCs, the use of the newly developed method on sourdough and lambic beer samples resulted in a confident description of the fungal diversity at the species level using a culture-independent approach. The species corresponded with those generally described for these matrices using a culture-dependent approach (50, 51). Special attention should be paid to S. bayanus for lambic beer samples, as other Saccharomyces species might also be present. Furthermore, the high level of UFGR in the sourdough samples could be explained by the presence of the ITS region in plants, which, in some cases, might be similar enough to allow primer hybridization, as described before (12, 52). Therefore, when using this method in cereal-rich matrices, a high sequencing depth should be pursued to have a sufficient number of sequence reads to be able to reliably describe the full yeast diversity after the removal of the UFGR, as was the case in the present study. Moreover, optimization of the centrifugation steps before DNA extractions to ensure that most of the plant cells are removed could be considered.
Finally, biases introduced during DNA extraction were not assessed in the current study, as the same DNA extraction methods were applied to make the DNA-based mock communities that were previously successfully applied for different fermented food matrices (24, 25, 36, 37, 53, 54). Further optimization of DNA extraction methods has the potential to improve the community composition analysis.
In conclusion, amplicon-based metabarcoding targeting the full ITS region using the PacBio HiFi sequencing technology is a promising method to unravel yeast diversity in fermented food samples, achieving a better taxonomic resolution compared to the commonly used Illumina-based methodologies. However, further research is needed to find ways to improve the identification at the species level in environments rich in Saccharomyces and Hanseniaspora species.
MATERIALS AND METHODS
DNA-based mock communities
To critically assess the methodology developed during the present study that aimed to investigate yeast diversity through amplicon-based metabarcoding of the ITS1-2 region, different DbMCs were constructed representing strains of species that play an important role in food fermentation processes.
Strains, growth conditions, and yeast genomic DNA extraction
All yeast strains used in the present study (Table 1) were stored at −80°C in cryovials containing yeast extract-peptone-glucose (YPG; Oxoid, Basingstoke, Hampshire, UK) cultures (55), supplemented with 25% (vol/vol) glycerol (Sigma-Aldrich, Saint-Louis, MO, USA) as cryoprotectant, as part of the laboratory collection of the Research Group of Industrial Microbiology and Food Biotechnology (IMDO-VUB). Each yeast strain was streaked on YPG agar medium supplemented with 200 ppm of chloramphenicol (Sigma-Aldrich) and incubated at 30°C for 48 h. A single colony of each strain was transferred to 10 mL of liquid YPG medium and incubated at the same temperature for 24 h. Yeast cell pellets were obtained by centrifugation of 2 mL of the latter cultures at 21,100 × g for 5 min. Yeast genomic DNA was extracted from these cell pellets using 600 µL of enzymatic lysis buffer with lyticase (200 U; Sigma-Aldrich), zymolyase (15 U; G-Biosciences, Saint Louis, MO, USA), and proteinase K (60 mAnsonU; Merck, Darmstadt, Germany) as described before (24). DNA purification was performed using a DNeasy Blood and Tissue kit (Qiagen, Hilden, Germany). The DNA concentrations were measured by fluorimetry (Qubit; Thermo Fisher Scientific, Waltham, MA, USA).
Composition of the DbMCs
A general DbMC, already available in the IMDO-VUB research group, composed of 17 strains of fermented food-related yeast species (Table S1), was used to evaluate the success of the PCR amplification (Temperature gradient section).
A total of seven different DbMCs, including the most abundant yeast species present in the microbial communities of various food fermentation processes, namely those of sourdough, lambic beer, and cocoa, were constructed (Table 1). In the case of sourdough, one DbMC containing equal concentrations of DNA of seven common sourdough yeast species (51, 56) was constructed. In the case of lambic beer, three DbMCs were considered, namely one DbMC with equal concentrations of DNA of three Brettanomyces and six Saccharomyces species, one DbMC representing the alcoholic fermentation phase of lambic beer that is rich in Saccharomyces species, and one DbMC representing the maturation phase of lambic beer that is rich in Brettanomyces species (50). In the case of cocoa, three DbMCs were considered, namely, one DbMC with equal concentrations of DNA of H. opuntiae, S. cerevisiae, and P. kudriavzevii, one DbMC representing the early stage of the cocoa fermentation process that is rich in Hanseniaspora species (4), and one DbMC representing the late stage of the cocoa fermentation process that is rich in S. cerevisiae and Pichia species (4). Each DbMC was constructed by pooling the strains’ DNA (Table 1) to get a final total DNA concentration of 1.5 ng/µL.
Fermented food samples
Sampling
To assess the method’s validity in fermented food samples, two sourdough samples and two lambic beer samples were analyzed. However, no recently collected cocoa samples were available at the time of the study. The sourdough samples were both spontaneous Type 1 sourdoughs originating from two artisan bakeries, with Sourdough A being a 3-year-old rye sourdough and Sourdough B being a 9-year-old wheat sourdough. Cell pellets were collected by diluting the sourdoughs 1:10 with sterile peptone saline (0.1% bacteriological peptone [Oxoid] and 0.85% NaCl [Merck]) as described before (25). After mixing to get a homogenous sample, a two-step centrifugation was performed. First, the homogenous samples were centrifuged at 1,000 × g for 5 min. Second, the flour debris-free supernatants were centrifuged at 5,000 × g for 20 min to collect the cell pellet. The lambic beer samples consisted of samples taken during the production of lambic beer, with Lambic beer A taken after 5 days of barrel fermentation and Lambic beer B taken after 3 months of barrel fermentation. The cell pellets were collected by a centrifugation step at 5,000 × g for 20 min (21). In both cases, the cell pellets were stored at −20°C until further DNA extraction.
DNA extraction from fermented food samples
Extraction of total DNA from the cell pellets was performed as described for sourdough (25) and lambic beer samples (22). Briefly, in both cases, a combination of both a bacterial lysis solution (300 µL of bacterial lysis buffer with 100 U mutanolysin [Sigma-Aldrich] and 320 kU of lysozyme [Merck]) and a yeast lysis solution (600 µL of yeast lysis buffer with 15 U zymolyase [G-Biosciences] and 200 U lyticase [Sigma-Aldrich]) was used to lyse the bacterial and yeast cells. Furthermore, mechanical lysis was performed using UV-C-irradiated glass beads (Sigma-Aldrich). Then, protein digestion was performed with 40.0 μL of a 20.0% (m/m) sodium dodecyl sulfate (Sigma-Aldrich) solution and 50.0 μL of proteinase K solution (2.0 mg/mL of proteinase K; Merck) per sample. In the case of the sourdough samples, they were further incubated in a 5.0 M NaCl (Merck) solution and a 10.0% (m/m) cetyltrimethylammonium bromide (Merck) solution. Finally, DNA was purified using chloroform-phenol-isoamyl alcohol solution and a DNeasy Blood and Tissue kit (Qiagen). The DNA concentrations were measured by Qubit fluorimetry (Thermo Fisher Scientific).
PCR
Selection of the primer sets
From all primers described to target the yeast rRNA operon (27, 57, 58) and, in particular, from those that allow for amplifying the full internal transcribed region, two pairs of primers were selected. Specifically, primers BITS (5′-ACCTGCGGARGGATCA-3′) and ITS1 (5′-TCCGTAGGTGAACCTGCGG-3′) were selected as forward primers, and primer ITS4 (5′-TCCTCCGCTTATTGATATGC-3′) was selected as reverse primer. The BITS primer, together with B58S3 as a reverse primer, has previously been used to target the fungal ITS1 region to describe the fungal diversity of food fermentation processes using Illumina amplicon-based metabarcoding (22, 25, 27, 36–38). The primers ITS1 and ITS4 are generally used to amplify the ITS1-2 region to identify yeast isolates in a culture-dependent approach (22, 25, 37, 38, 59, 60). The primers were tagged with Kinnex adaptors and 5′ sample-specific barcodes (Integrated DNA Technologies, Leuven, Belgium; Table S2) to allow for multiplexed sequencing, according to the manufacturer’s instructions (Pacific Biosciences, Menlo Park, CA, USA).
Temperature gradient
PCR assays were performed using KAPA HiFi DNA Polymerase (Hot Start and Ready Mix; Roche, Basel, Switzerland) based on the PCR assay described for 16S rRNA gene amplification for PacBio HiFi sequencing (24), but with a different annealing temperature. These PCR amplifications were done using the general DbMC as a DNA template. Specifically, per sample, 1.5 µL of DNase and RNase-free water (VWR International, Radnor, PA, USA), 12.5 µL of 2× KAPA HiFi HotStart ReadyMix (Roche), 3 µL of the forward primer, 3 µL of the reverse primers (2.5 µM each), and 5 µL of template DNA (1.5 ng/µL) were mixed. To assess the most optimal annealing temperature, PCR assays were performed applying a temperature gradient using a TProfessional Basic Gradient thermocycler (BioMetra, Göttingen, Germany) with values 52.0°C, 53.0°C, 54.2°C, 55.4°C, 56.0°C, and 57.0°C as annealing temperatures, in combination with an initial denaturation at 95°C for 3 min, followed by 27 cycles of denaturation at 95°C for 30 s, annealing for 30 s, and extension at 72°C for 60 s. The final PCR assays consisted of 3 min of denaturation at 95°C, followed by 27 cycles of denaturation at 95°C for 30 s, annealing at 52°C for 30 s, and extension at 72°C for 60 s. The PCR amplicons obtained were visualized through gel electrophoresis using 1.5% (m/v) agarose gels and performed at 100 V for 1 h. PCR amplicon purification was performed using a Wizard Plus SV Mini-preps DNA purification system (Promega, Madison, WI, USA). PCR assays for all DbMCs and food samples were performed in triplicate with sample-specific barcodes (Table S3).
PacBio HiFi long-read sequencing
After PCR amplification with the sample-specific barcoded primers, sequencing library preparation was done following PacBio’s workflow (“Procedure & checklist—preparing Kinnex libraries from 16S rRNA amplicons,” PacBio Document 103-238-800). Briefly, amplicons were pooled, and a Kinnex PCR was performed, during which Kinnex terminal adaptors were added to the amplicons to enable concatenation. Next, amplicons were concatenated using the Kinnex enzyme and ligase. Finally, SMRTbell adapters were ligated to generate circular DNA templates suitable for circular consensus sequencing. The obtained SMRTbell library was sequenced on a Revio platform (Pacific Biosciences) at the VIB Nucleomics Core Facility (Leuven, Belgium).
Data processing
The reads obtained from the Revio platform were deconcatenated using Skera (v1.3.0, PacBio), and individual amplicons were demultiplexed using Lima (v2.12.0, Pacific Biosciences), with a minimal length cutoff set at 50 bp to avoid removing any short ITS sequences. Demultiplexed reads were further processed using RStudio (version 4.2.1 [61]), and amplicon sequence variants were obtained using the DADA2 package (version 1.26.0 [62]). The filtering parameters used were minQ = 2, minLen = 300, maxLen = 1,000, maxN = 0, and maxEE = 5. Taxonomy was assigned using the fungal-specific version of the UNITE database (v10.0 [63]) using a minBoot = 80. Reads not identified at the genus level (“unidentifiable fungal genus reads”) were filtered out. ASVs were grouped per species, and relative abundances of each species per sample (number of reads assigned to one species in one sample/total number of reads of that sample × 100) were calculated, including those species with a relative abundance below 0.5% under the category “Minorities < 0.5%.” Furthermore, the UFGRs were used as queries for blastn searches using the core_nt database of the National Center for Biotechnology Information (July 2025).
Rarefaction curve analysis was performed using the R package vegan (v2.6-4 [64]). The divergence rate was calculated as the Bray-Curtis distance between the expected and the observed relative abundance using the same R package. The expected relative abundance was calculated based on the percentage of DNA of each species included in each DbMC, whereas the observed relative abundance was the one obtained after processing the sequence data. The number of false negatives (FNs) was defined as the number of expected taxa that were not recovered, and the false positives (FPs) were defined as the number of recovered taxa that were not expected. Furthermore, the number of true positives (TPs) was defined as the number of recovered taxa that were expected. Finally, the precision of the method was calculated as TP/(TP + FP), the recall as TP/(TP + FN), and the F1 score as 2TP/(2TP + 2FP + FN) (12).
ACKNOWLEDGMENTS
The authors thank Stéphane Plaisance, Lim De Swert, and the VIB Nucleomics Core Facility team for the sequencing and technical guidance provided.
This work was supported by the Research Council of the Vrije Universiteit Brussel (SRP71). Part of the computational resources used were provided by the Flemish Supercomputer Centre (VSC), funded by the Research Foundation-Flanders (FWO-Vlaanderen) and the Flemish Government. T.G. is the recipient of a PhD fellowship from the Vrije Universiteit Brussel as part of the European research project HealthFerm, which is co-funded by the European Union under the Horizon Europe grant agreement No. 101060247 and the Swiss State Secretariat for Education, Research and Innovation (SERI) under contract No. 22.00210. Views and opinions expressed are, however, those of the authors only and do not necessarily reflect those of the European Union or European Research Executive Agency (REA). Neither the European Union nor REA can be held responsible for them.
Contributor Information
Stefan Weckx, Email: Stefan.Weckx@vub.be.
Kalliopi Rantsiou, University of Turin, Gruglisco, Turin, Italy.
DATA AVAILABILITY
The sequence reads are available at the European Nucleotide Archive of the European Bioinformatics Institute (ENA/EBI) under the BioProject accession number PRJEB101075.
SUPPLEMENTAL MATERIAL
The following material is available online at https://doi.org/10.1128/spectrum.00060-26.
Tables S1 to S4 and Figure S1.
ASM does not own the copyrights to Supplemental Material that may be linked to, or accessed through, an article. The authors have granted ASM a non-exclusive, world-wide license to publish the Supplemental Material files. Please contact the corresponding author directly for reuse.
REFERENCES
- 1. Marco ML, Sanders ME, Gänzle M, Arrieta MC, Cotter PD, De Vuyst L, Hill C, Holzapfel W, Lebeer S, Merenstein D, et al. 2021. The international scientific association for probiotics and prebiotics (ISAPP) consensus statement on fermented foods. Nat Rev Gastroenterol Hepatol 18:196–208. doi: 10.1038/s41575-020-00390-5 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 2. Gänzle M. 2022. The periodic table of fermented foods: limitations and opportunities. Appl Microbiol Biotechnol 106:2815–2826. doi: 10.1007/s00253-022-11909-y [DOI] [PubMed] [Google Scholar]
- 3. Valentino V, Magliulo R, Farsi D, Cotter PD, O’Sullivan O, Ercolini D, De Filippis F. 2024. Fermented foods, their microbiome and its potential in boosting human health. Microb Biotechnol 17:e14428. doi: 10.1111/1751-7915.14428 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 4. Díaz-Muñoz C, De Vuyst L. 2022. Functional yeast starter cultures for cocoa fermentation. J Appl Microbiol 133:39–66. doi: 10.1111/jam.15312 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 5. Zhao Z, Gänzle MG. 2025. Sequence based characterization of microbial communities in food: the panacea for smart detection of food microbes or dirty deeds done dirt cheap? Trends Food Sci Technol 162:105113. doi: 10.1016/j.tifs.2025.105113 [DOI] [Google Scholar]
- 6. Yap M, Ercolini D, Álvarez-Ordóñez A, O’Toole PW, O’Sullivan O, Cotter PD. 2022. Next-generation food research: use of meta-omic approaches for characterizing microbial communities along the food chain. Annu Rev Food Sci Technol 13:361–384. doi: 10.1146/annurev-food-052720-010751 [DOI] [PubMed] [Google Scholar]
- 7. Weckx S, Van Kerrebroeck S, De Vuyst L. 2019. Omics approaches to understand sourdough fermentation processes. Int J Food Microbiol 302:90–102. doi: 10.1016/j.ijfoodmicro.2018.05.029 [DOI] [PubMed] [Google Scholar]
- 8. Callahan BJ, McMurdie PJ, Rosen MJ, Han AW, Johnson AJA, Holmes SP. 2016. DADA2: high-resolution sample inference from illumina amplicon data. Nat Methods 13:581–583. doi: 10.1038/nmeth.3869 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 9. De Filippis F, Laiola M, Blaiotta G, Ercolini D. 2017. Different amplicon targets for sequencing-based studies of fungal diversity. Appl Environ Microbiol 83:e00905-17. doi: 10.1128/AEM.00905-17 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 10. Walsh AM, Crispie F, Claesson MJ, Cotter PD. 2017. Translating omics to food microbiology. Annu Rev Food Sci Technol 8:113–134. doi: 10.1146/annurev-food-030216-025729 [DOI] [PubMed] [Google Scholar]
- 11. Billington C, Kingsbury JM, Rivas L. 2022. Metagenomics approaches for improving food safety: a review. J Food Prot 85:448–464. doi: 10.4315/JFP-21-301 [DOI] [PubMed] [Google Scholar]
- 12. Rué O, Coton M, Dugat-Bony E, Howell K, Irlinger F, Legras J-L, Loux V, Michel E, Mounier J, Neuvéglise C, Sicard D. 2023. Comparison of metabarcoding taxonomic markers to describe fungal communities in fermented foods. Peer Community Journal 3:e97. doi: 10.24072/pcjournal.321 [DOI] [Google Scholar]
- 13. Tedersoo L, Albertsen M, Anslan S, Callahan B. 2021. Perspectives and benefits of high-throughput long-read sequencing in microbial ecology. Appl Environ Microbiol 87:e0062621. doi: 10.1128/AEM.00626-21 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 14. Yarza P, Yilmaz P, Pruesse E, Glöckner FO, Ludwig W, Schleifer K-H, Whitman WB, Euzéby J, Amann R, Rosselló-Móra R. 2014. Uniting the classification of cultured and uncultured bacteria and archaea using 16S rRNA gene sequences. Nat Rev Microbiol 12:635–645. doi: 10.1038/nrmicro3330 [DOI] [PubMed] [Google Scholar]
- 15. Johnson JS, Spakowicz DJ, Hong B-Y, Petersen LM, Demkowicz P, Chen L, Leopold SR, Hanson BM, Agresta HO, Gerstein M, et al. 2019. Evaluation of 16S rRNA gene sequencing for species and strain-level microbiome analysis. Nat Commun 10:5029. doi: 10.1038/s41467-019-13036-1 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 16. Sun D-L, Jiang X, Wu QL, Zhou N-Y. 2013. Intragenomic heterogeneity of 16S rRNA genes causes overestimation of prokaryotic diversity. Appl Environ Microbiol 79:5962–5969. doi: 10.1128/AEM.01282-13 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 17. Wagner J, Coupland P, Browne HP, Lawley TD, Francis SC, Parkhill J. 2016. Evaluation of PacBio sequencing for full-length bacterial 16S rRNA gene classification. BMC Microbiol 16:274. doi: 10.1186/s12866-016-0891-4 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 18. Schloss PD, Girard RA, Martin T, Edwards J, Thrash JC. 2016. Status of the archaeal and bacterial census: an update. mBio 7:e00201-16. doi: 10.1128/mBio.00201-16 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 19. Buetas E, Jordán-López M, López-Roldán A, D’Auria G, Martínez-Priego L, De Marco G, Carda-Diéguez M, Mira A. 2024. Full-length 16S rRNA gene sequencing by PacBio improves taxonomic resolution in human microbiome samples. BMC Genomics 25:310. doi: 10.1186/s12864-024-10213-5 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 20. Callahan BJ, Wong J, Heiner C, Oh S, Theriot CM, Gulati AS, McGill SK, Dougherty MK. 2019. High-throughput amplicon sequencing of the full-length 16S rRNA gene with single-nucleotide resolution. Nucleic Acids Res 47:e103. doi: 10.1093/nar/gkz569 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 21. Hlangwani E, Abrahams A, Masenya K, Adebo OA. 2023. Analysis of the bacterial and fungal populations in South African sorghum beer (umqombothi) using full-length 16S rRNA amplicon sequencing. World J Microbiol Biotechnol 39:350. doi: 10.1007/s11274-023-03764-4 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 22. Bongaerts D, Bouchez A, De Roos J, Cnockaert M, Wieme AD, Vandamme P, Weckx S, De Vuyst L. 2024. Refermentation and maturation of lambic beer in bottles: a necessary step for gueuze production. Appl Environ Microbiol 90:e0186923. doi: 10.1128/aem.01869-23 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 23. Jin H, Mo L, Pan L, Hou Q, Li C, Darima I, Yu J. 2018. Using PacBio sequencing to investigate the bacterial microbiota of traditional buryatian cottage cheese and comparison with Italian and Kazakhstan artisanal cheeses. J Dairy Sci 101:6885–6896. doi: 10.3168/jds.2018-14403 [DOI] [PubMed] [Google Scholar]
- 24. Decadt H, Weckx S, De Vuyst L. 2023. The rotation of primary starter culture mixtures results in batch-to-batch variations during gouda cheese production. Front Microbiol 14:1128394. doi: 10.3389/fmicb.2023.1128394 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 25. Pradal I, González-Alonso V, Wardhana YR, Cnockaert M, Wieme AD, Vandamme P, De Vuyst L. 2024. Various cold storage-backslopping cycles show the robustness of Limosilactobacillus fermentum IMDO 130101 as starter culture for type 3 sourdough production. Int J Food Microbiol 411:110522. doi: 10.1016/j.ijfoodmicro.2023.110522 [DOI] [PubMed] [Google Scholar]
- 26. Schoch CL, Seifert KA, Huhndorf S, Robert V, Spouge JL, Levesque CA, Chen W, Fungal Barcoding Consortium . 2012. Nuclear ribosomal internal transcribed spacer (ITS) region as a universal DNA barcode marker for Fungi. Proc Natl Acad Sci USA 109:6241–6246. doi: 10.1073/pnas.1117018109 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 27. Bokulich NA, Mills DA. 2013. Improved selection of internal transcribed spacer-specific primers enables quantitative, ultra-high-throughput profiling of fungal communities. Appl Environ Microbiol 79:2519–2526. doi: 10.1128/AEM.03870-12 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 28. Hoggard M, Vesty A, Wong G, Montgomery JM, Fourie C, Douglas RG, Biswas K, Taylor MW. 2018. Characterizing the human mycobiota: a comparison of small subunit rRNA, ITS1, ITS2, and large subunit rRNA genomic targets. Front Microbiol 9:2208. doi: 10.3389/fmicb.2018.02208 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 29. Hu Y, Irinyi L, Hoang MTV, Eenjes T, Graetz A, Stone EA, Meyer W, Schwessinger B, Rathjen JP. 2022. Inferring species compositions of complex fungal communities from long- and short-read sequence data. mBio 13:e0244421. doi: 10.1128/mbio.02444-21 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 30. Mbareche H, Veillette M, Bilodeau G, Duchaine C. 2020. Comparison of the performance of ITS1 and ITS2 as barcodes in amplicon-based sequencing of bioaerosols. PeerJ 8:e8523. doi: 10.7717/peerj.8523 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 31. Tedersoo L, Anslan S, Bahram M, Kõljalg U, Abarenkov K. 2020. Identifying the “unidentified” fungi: a global-scale long-read third-generation sequencing approach. Fungal Divers 103:273–293. doi: 10.1007/s13225-020-00456-4 [DOI] [Google Scholar]
- 32. James TY, Marino JA, Perfecto I, Vandermeer J. 2016. Identification of putative coffee rust mycoparasites via single-molecule DNA sequencing of infected pustules. Appl Environ Microbiol 82:631–639. doi: 10.1128/AEM.02639-15 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 33. Walder F, Schlaeppi K, Wittwer R, Held AY, Vogelgsang S, van der Heijden MGA. 2017. Community profiling of Fusarium in combination with other plant-associated fungi in different crop species using SMRT sequencing. Front Plant Sci 8:2019. doi: 10.3389/fpls.2017.02019 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 34. Latz MAC, Grujcic V, Brugel S, Lycken J, John U, Karlson B, Andersson A, Andersson AF. 2022. Short- and long-read metabarcoding of the eukaryotic rRNA operon: evaluation of primers and comparison to shotgun metagenomics sequencing. Mol Ecol Resour 22:2304–2318. doi: 10.1111/1755-0998.13623 [DOI] [PubMed] [Google Scholar]
- 35. Tedersoo L, Tooming-Klunderud A, Anslan S. 2018. PacBio metabarcoding of fungi and other eukaryotes: errors, biases and perspectives. New Phytol 217:1370–1385. doi: 10.1111/nph.14776 [DOI] [PubMed] [Google Scholar]
- 36. Díaz-Muñoz C, Van de Voorde D, Comasio A, Verce M, Hernandez CE, Weckx S, De Vuyst L. 2020. Curing of cocoa beans: fine-scale monitoring of the starter cultures applied and metabolomics of the fermentation and drying steps. Front Microbiol 11:616875. doi: 10.3389/fmicb.2020.616875 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 37. González-Alonso V, Pradal I, Wardhana YR, Cnockaert M, Wieme AD, Vandamme P, De Vuyst L. 2024. Microbial ecology and metabolite dynamics of backslopped triticale sourdough productions and the impact of scale. Int J Food Microbiol 408:110445. doi: 10.1016/j.ijfoodmicro.2023.110445 [DOI] [Google Scholar]
- 38. Pradal I, González-Alonso V, Wardhana YR, De Vuyst L. 2025. Companilactobacillus crustorum LMG 23699 and Wickerhamomyces anomalus IMDO 010110 form a candidate, stable, mixed-strain starter culture for sourdough production. Int J Food Microbiol 440:111278. doi: 10.1016/j.ijfoodmicro.2025.111278 [DOI] [PubMed] [Google Scholar]
- 39. Čadež N, Poot GA, Raspor P, Smith MT. 2003. Hanseniaspora meyeri sp. nov., Hanseniaspora clermontiae sp. nov., Hanseniaspora lachancei sp. nov. and Hanseniaspora opuntiae sp. nov., novel apiculate yeast species. Int J Syst Evol Microbiol 53:1671–1680. doi: 10.1099/ijs.0.02618-0 [DOI] [PubMed] [Google Scholar]
- 40. Wang Q-M, Li J, Wang S-A, Bai F-Y. 2008. Rapid differentiation of phenotypically similar yeast species by single-strand conformation polymorphism analysis of ribosomal DNA. Appl Environ Microbiol 74:2604–2611. doi: 10.1128/AEM.02223-07 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 41. Bendixsen DP, Frazão JG, Stelkens R. 2022. Saccharomyces yeast hybrids on the rise. Yeast 39:40–54. doi: 10.1002/yea.3684 [DOI] [PubMed] [Google Scholar]
- 42. Heeger F, Bourne EC, Baschien C, Yurkov A, Bunk B, Spröer C, Overmann J, Mazzoni CJ, Monaghan MT. 2018. Long-read DNA metabarcoding of ribosomal RNA in the analysis of fungi from aquatic environments. Mol Ecol Resour 18:1500–1514. doi: 10.1111/1755-0998.12937 [DOI] [PubMed] [Google Scholar]
- 43. D’Andreano S, Cuscó A, Francino O. 2020. Rapid and real-time identification of fungi up to species level with long amplicon nanopore sequencing from clinical samples. Biol Methods Protoc 6:bpaa026. doi: 10.1093/biomethods/bpaa026 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 44. Krabberød AK, Stokke E, Thoen E, Skrede I, Kauserud H. 2025. The ribosomal operon database: a full-length rDNA operon database derived from genome assemblies. Mol Ecol Resour 25:e14031. doi: 10.1111/1755-0998.14031 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 45. Kurtzman CP. 2003. Phylogenetic circumscription of Saccharomyces, Kluyveromyces and other members of the Saccharomycetaceae, and the proposal of the new genera Lachancea, Nakaseomyces, Naumovia, Vanderwaltozyma and Zygotorulaspora. FEMS Yeast Res 4:233–245. doi: 10.1016/S1567-1356(03)00175-2 [DOI] [PubMed] [Google Scholar]
- 46. Kurtzman CP, Robnett CJ, Basehoar-Powers E. 2008. Phylogenetic relationships among species of Pichia, Issatchenkia and Williopsis determined from multigene sequence analysis, and the proposal of Barnettozyma gen. nov., Lindnera gen. nov. and Wickerhamomyces gen. nov. FEMS Yeast Res 8:939–954. doi: 10.1111/j.1567-1364.2008.00419.x [DOI] [PubMed] [Google Scholar]
- 47. Wang L, Groenewald M, Wang QM, Boekhout T. 2015. Reclassification of Saccharomycodes sinensis, proposal of Yueomyces sinensis gen. nov., comb. nov. within Saccharomycetaceae (Saccharomycetales, Saccharomycotina). PLoS One 10:e0136987. doi: 10.1371/journal.pone.0136987 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 48. Groenewald M, Hittinger CT, Bensch K, Opulente DA, Shen X-X, Li Y, Liu C, LaBella AL, Zhou X, Limtong S, et al. 2023. A genome-informed higher rank classification of the biotechnologically important fungal subphylum Saccharomycotina. Stud Mycol 105:1–22. doi: 10.3114/sim.2023.105.01 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 49. Liu F, Hu Z-D, Yurkov A, Chen X-H, Bao W-J, Ma Q, Zhao WN, Pan S, Zhao X-M, Liu J-H, et al. 2024. Saccharomycetaceae: delineation of fungal genera based on phylogenomic analyses, genomic relatedness indices and genomics-based synapomorphies. Persoonia 52:1–21. doi: 10.3767/persoonia.2024.52.01 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 50. Bongaerts D, De Roos J, De Vuyst L. 2021. Technological and environmental features determine the uniqueness of the lambic beer microbiota and production process. Appl Environ Microbiol 87:e0061221. doi: 10.1128/AEM.00612-21 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 51. von Gastrow L, Gianotti A, Vernocchi P, Serrazanetti DI, Sicard D. 2023. Chapter 7, Taxonomy, biodiversity, and physiology of sourdough yeasts, p 161–212. In M Gobbetti, M Gänzle (ed), Handbook on sourdough biotechnology. Springer Nature, Switzerland. [Google Scholar]
- 52. Urien C, Legrand J, Montalent P, Casaregola S, Sicard D. 2019. Fungal species diversity in French bread sourdoughs made of organic wheat flour. Front Microbiol 10:201. doi: 10.3389/fmicb.2019.00201 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 53. Vermote L, Verce M, De Vuyst L, Weckx S. 2018. Amplicon and shotgun metagenomic sequencing indicates that microbial ecosystems present in cheese brines reflect environmental inoculation during the cheese production process. Int Dairy J 87:44–53. doi: 10.1016/j.idairyj.2018.07.010 [DOI] [Google Scholar]
- 54. Van der Veken D, Poortmans M, Dewulf L, Fraeye I, Michiels C, Leroy F. 2023. Challenge tests reveal limited outgrowth of proteolytic Clostridium botulinum during the production of nitrate- and nitrite-free fermented sausages. Meat Sci 200:109158. doi: 10.1016/j.meatsci.2023.109158 [DOI] [PubMed] [Google Scholar]
- 55. Comasio A, Harth H, Weckx S, De Vuyst L. 2019. The addition of citrate stimulates the production of acetoin and diacetyl by a citrate-positive Lactobacillus crustorum strain during wheat sourdough fermentation. Int J Food Microbiol 289:88–105. doi: 10.1016/j.ijfoodmicro.2018.08.030 [DOI] [PubMed] [Google Scholar]
- 56. Arora K, Ameur H, Polo A, Di Cagno R, Rizzello CG, Gobbetti M. 2021. Thirty years of knowledge on sourdough fermentation: a systematic review. Trends Food Sci Technol 108:71–83. doi: 10.1016/j.tifs.2020.12.008 [DOI] [Google Scholar]
- 57. Tedersoo L, Lindahl B. 2016. Fungal identification biases in microbiome projects. Environ Microbiol Rep 8:774–779. doi: 10.1111/1758-2229.12438 [DOI] [PubMed] [Google Scholar]
- 58. Tedersoo L, Anslan S. 2019. Towards PacBio-based pan-eukaryote metabarcoding using full-length ITS sequences. Environ Microbiol Rep 11:659–668. doi: 10.1111/1758-2229.12776 [DOI] [PubMed] [Google Scholar]
- 59. Bazalová O, Cihlář JZ, Dlouhá Z, Bár L, Dráb V, Kavková M. 2022. Rapid sourdough yeast identification using panfungal PCR combined with high resolution melting analysis. J Microbiol Methods 199:106522. doi: 10.1016/j.mimet.2022.106522 [DOI] [PubMed] [Google Scholar]
- 60. Sánchez-Adriá IE, Sanmartín G, Prieto JA, Estruch F, Fortis E, Randez-Gil F. 2023. Technological and acid stress performance of yeast isolates from industrial sourdough. LWT 184:114957. doi: 10.1016/j.lwt.2023.114957 [DOI] [Google Scholar]
- 61. Posit team . 2025. Posit Software, PBC. Boston, MA. http://www.posit.co. [Google Scholar]
- 62. Callahan BJ, McMurdie PJ, Holmes SP. 2017. Exact sequence variants should replace operational taxonomic units in marker-gene data analysis. ISME J 11:2639–2643. doi: 10.1038/ismej.2017.119 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 63. Abarenkov K, Zirk A, Piirmann T, Pöhönen R, Ivanov F, Nilsson RH, Kõljalg U. 2025. UNITE general FASTA release for fungi. UNITE Community 10:15156. doi: 10.1093/nar/gkad1039 [Google Scholar]
- 64. Oksanen J, Simpson G, Blanchet F, Kindt R, Legendre P, Minchin P, O’Hara R, Solymos P, Stevens M, Szoecs E, et al. 2022. Vegan: Community Ecology Package. Available from: https://CRAN.R-project.org/package=vegan
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Tables S1 to S4 and Figure S1.
Data Availability Statement
The sequence reads are available at the European Nucleotide Archive of the European Bioinformatics Institute (ENA/EBI) under the BioProject accession number PRJEB101075.


