Skip to main content
Scientific Reports logoLink to Scientific Reports
. 2025 Nov 4;15:38587. doi: 10.1038/s41598-025-22523-z

Manually weighted taxonomy classifiers improve species-specific rumen microbiome analysis compared to unweighted or average weighted taxonomy classifiers

Ryukseok Kang 1, Zhongtang Yu 2, Hanbeen Kim 3,4, Jakyeom Seo 4, Minseok Kim 5, Tansol Park 1,✉
PMCID: PMC12586449  PMID: 41188334

Abstract

Previous research has demonstrated that applying taxonomic weights to shotgun metagenomic data can improve species identification in 16S rRNA gene-based microbiome analysis. However, such an approach does not allow for accurate analysis of samples collected from less studied habitats, such as rumen. In the present study, we developed a method to incorporate taxonomic weights based on relative abundance of species identified from shotgun sequencing and amplicon sequencing data derived from rumen. Using this weighting method, we evaluated latest versions of five prominent databases—SILVA, Greengenes2 (GG2), RDP, NCBI RefSeq, and GTDB—against the BLAST 16S rRNA database, assessing classification counts, fully classified ratios (proportion of ASVs classified to a known genus and species), and error rates. Our results indicated that providing taxonomic weights partially increased classification counts and fully classified ratios, although the extent of improvement varied across databases. A reduction in error rates was also observed compared to the unweighted taxonomy classifier (P < 0.05). While GG2 and SILVA struggled with accurate classification at the species level owing to their inherent database characteristics, GTDB consistently improved all metrics using the manually weighted taxonomy classifier, achieving up to an 8% error rate reduction at the species level. NCBI RefSeq and RDP also exhibited remarkable improvement in the classification counts and fully classified ratios, along with error rate reductions by up to 47% at the species level. These findings demonstrate that amplicon sequencing datasets can enhance rumen microbiome analyses through effective weighting methods. While SILVA is commonly used in metataxonomic analyses of the rumen microbiome, we recommend NCBI RefSeq for species-level classification due to its superior accuracy and minimal ambiguous classification (e.g., “uncultured” or “sp.“) in future metataxonomic studies.

Supplementary Information

The online version contains supplementary material available at 10.1038/s41598-025-22523-z.

Keywords: Taxonomy classifier, Ruminal microbiome, Manually weighted taxonomy classifier, 16S microbiome analysis, High-resolution analysis, NCBI RefSeq

Subject terms: Metagenomics, Microbiome, Bioinformatics, Zoology, Microbial ecology

Introduction

Ruminants host a diverse and highly specialized microbiome in their rumen, which enables them to digest a wide range of feedstuffs, particularly fibrous materials that would otherwise be indigestible1,2. Accurately identifying the key microbes and understanding their contributions to various rumen fermentation processes are crucial for improving animal nutrition and productivity and mitigating the environmental impact associated with this unique animal group. The rumen microbiome has been intensively studied using culture-based and culture-independent nucleic acid-based techniques, with recent efforts predominantly leveraging the latter that utilize sequencing approaches like metataxonomics, metagenomics, and metatranscriptomics3. Advances in sequencing technologies and the expansion of reference databases have significantly improved the accuracy of microbiome characterization, including taxonomic classification. However, most gut microbiome databases are disproportionally dominated by human microbiome data4. Such database bias may reduce the accuracy of analyses for non-human microbiomes, including the rumen microbiome.

To address this challenge, several initiatives have been proposed, including the Global Rumen Census project, the “Hungate1000,” “Holoruminant,” and meta-omic approaches aimed at enhancing comprehensive characterization of the rumen microbiome5–8. Despite these efforts, the accuracy of 16S rRNA gene-based metataxonomic analysis of the rumen microbiome remains uncertain. Recent advances, such as weighted taxonomy classifiers, have shown promise in improving taxonomic resolution. For example, Kaehler et al. (2019) utilized q2-clawback and the Earth Microbiome Project Ontology (EMPO) 3 datasets to apply weighting to taxonomy classifiers. Specifically, this approach assigns varying weights to individual taxa: higher weights to taxa prevalent in specific environments and lower weights to rare or absent taxa9. This approach improved species-level identification and abundance estimation accuracy. It was also suggested that weighted taxonomy classifiers could further improve taxonomic resolution when applied to metagenomic analyses. However, since the ruminal microbiome is not included in the EMPO habitat types, the proposed classifiers probably cannot be applied to analyses of rumen samples.

The distinct rumen microbiome profiles observed across ruminant species10,11 suggest that generalized taxonomic weights may compromise classification accuracy when applied to a specific ruminant species. Therefore, creating and implementing ruminant species-specific taxonomy classifiers would be essential for achieving greater classification accuracy. Additionally, while shotgun sequencing-based metagenomics has been increasingly used for improved resolution, 16S rRNA gene amplicon-based metataxonomics remains widely used for rumen microbiome profiling due to its cost-effectiveness. To achieve greater accuracy in metataxonomic analyses of the rumen microbiome, we utilized multiple datasets of partial and full-length 16S rRNA gene sequences along with shotgun sequences obtained from native Korean Hanwoo cattle in developing weighted taxonomic classifiers. Furthermore, unlike the earlier studies by Kaehler et al. (2019) that used an outdated version of Greengenes for validating weighted methods, we used the most recent Greengenes database (GG2). We also included the latest versions of other commonly used databases—SILVA, NCBI reference sequence database (RefSeq), GTDB, and RDP—to explore how the choice of databases affects classification accuracy. We developed an enhanced taxonomy classification method that integrates weight assignment, additional preprocessing steps, and q2-clawback by utilizing data from native Korean Hanwoo cattle. This study indicates that applying taxonomic weights can significantly improve rumen microbiome analyses.

Thus, the aims of this study were to (i) evaluate whether incorporating amplicon sequences, in addition to shotgun sequences, into taxonomic classifier weighting could improve classification resolution, (ii) assess the applicability of the manually weighted classifier in the rumen microbiome using latest databases, and (iii) determine which database is most suitable for species-level classification of the rumen microbiome.

Results

The taxonomy classifier evaluation method was validated using microbiota datasets derived from both in vitro experiments and in vivo samples collected from the rumen of Hanwoo cattle. The in vitro fermentation experiment involved the collection of ruminal fluid and incubation with a mixed microbial substrate under controlled conditions, enabling the study of fermentation dynamics and microbial composition. This approach is widely used in ruminant research12. A recent study validated this approach by demonstrating that in vitro batch cultures of rumen fluid effectively maintained the ruminal microbiome for up to 48 h13. Based on this evidence, both in vivo and in vitro datasets were selected to validate the taxonomy classifiers.

The unprocessed taxonomy classifier was referred to as the unweighted taxonomy classifier (UWTC). The classifier provided by Quantitative Insights Into Microbial Ecology 2 (QIIME2) resource (https://library.qiime2.org/data-resources), which applies average weights derived from a large collection of EMPO datasets, was designated as the average weighted taxonomy classifier (AWTC). The classifier manually weighted using Hanwoo-specific microbiome data generated in this study was termed the manually weighted taxonomy classifier (MWTC). We compared UWTC, AWTC, and MWTC to evaluate their differences and to identify the most suitable database for species-level classification in rumen microbiome analysis.

To comprehensively evaluate the performance of taxonomic classifiers and databases, three core metrics: classification counts, fully classified ratios, and error rates, along with alpha and beta diversity were assessed. These evaluation results are summarized in Table 2, and detailed outcomes for both in vivo and in vitro datasets are described in the following sections.

Table 2.

Summary of key metrics including relative abundance, classification counts, fully classified ratios, and error rates comparing weighted classifiers (AWTC and MWTC) and UWTC across multiple databases in rumen microbiome analysis. Comparison results are presented in the order of in vivo full-length, in vivo V3–V4, in vitro full-length, and in vitro V3–V4.

Database Classifier Top 20 relative abundance Classification
counts
Fully classified
ratio
Error rate
Genus Species
GG2 2022.10 AWTC ○, ↓, ↓, ↓ 1) ↓, ↓, ↓, ↓ ○, ○, ↓, ↓ ↓, ○, ↓, ↓ N.A.
MWTC ○, ↑, ○, ○ ↑, ↑, ○, ↑ ↑, ↑, ○, ○ ○, ↑, ○, ○ N.A.
GG2 2024.09 AWTC ○, ↓, ↓, ↓ ↓, ↑, ↓, ↓ ○, ↓, ↓, ↓ ↑, ↓, ↓, ↓ ↑, ↑, ○, ↑
MWTC ○, ↑, ○, ○ ↑, ↑, ↓, ↑ ↑, ↑, ○, ○ ↑, ↑, ○, ↓ ↑, ↑, ○, ↑
GTDB 220.0 AWTC ↑, ↑, ↑, ○ ↑, ↑, ↑, ↓ ↑, ↑, ↑, ○ ↑, ○, ↑, ○ ↑, ↓, ○, ○
MWTC ↑, ↑, ↑, ↑ ↑, ↑, ↑, ↑ ↑, ↑, ↑, ↑ ↑, ↑, ↑, ↑ ↓, ↓, ○, ↓
NCBI 2024.10 AWTC ↑, ↑, ↑, ↑ ↑, ↑, ↑, ↑ ↑, ↑, ↑, ↑ ↑, ↑, ↑, ↑ ○, ↓, ↓, ↓
MWTC ↑, ↑, ↑, ↑ ↑, ↑, ↑, ↑ ↑, ↑, ↑, ↑ ↑, ↑, ↑, ↑ ↓, ↓, ↓, ↓
RDP 2.14 AWTC ↑, ↑, ↑, ↑ ↑, ↑, ↑, ↑ ↑, ↑, ↑, ↑ ↑, ↑, ↑, ↑ ↓, ↓, ○, ↓
MWTC ↑, ↑, ↑, ↑ ↑, ↑, ↑, ↑ ↑, ↑, ↑, ↑ ↑, ↑, ↑, ↑ ↓, ↓, ↓, ↓
SILVA 138.1 AWTC ○, ○, ○, ↑ ↑, ↓, ↑, ↑ ↓, ↑, ○, ○ ↓, ○, ↓, ○ N.A.
MWTC ○, ↑, ○, ↑ ↑, ↑, ↑, ↑ ↑, ↑, ↑, ↑ ↑, ↑, ○, ↑ N.A.
SILVA 138.2 AWTC ↓, ↓, ○, ↓ ↑, ↑, ↓, ↑ ↑, ↑, ○, ○ ○, ○, ↓, ○ ○, ↑, ○, ↑
MWTC ○, ↑, ○, ↑ ↑, ↑, ↑, ↑ ↑, ↑, ↑, ↑ ↑, ↑, ○, ↑ ○, ↑, ○, ○

1) Increase (↑), decrease (↓), no change or inconsistent (○) compared to UWTC.

2) N.A., Not analyzed.

Classifier evaluation using in vivo datasets

The classifiers were comparatively evaluated for diversity analyses and taxonomic identification of rumen microbes. To assess within-sample diversity and compare the performance of different classifiers, alpha diversity indices were measured at the genus and species level (Tables S1 and S2). Observed features were statistically different (P < 0.05) across all classifiers and taxonomy levels, except for GTDB with full-length amplicon sequences at the species level (P = 0.0526). The MWTC exhibited the highest values of observed features compared to the UWTC and AWTC with NCBI RefSeq and RDP (P < 0.01) but either lower or comparable diversity metric values with GG2, SILVA, and GTDB (P < 0.01). With the full-length amplicon sequences, MWTC showed the lowest evenness, Shannon index, and Simpson index compared to UWTC with NCBI RefSeq and RDP (P < 0.01), but with the V3-V4 amplicon sequences, these alpha diversity metrics were higher or comparable. When GTDB was used, AWTC and MWTC yielded higher Shannon, Simpson, and Inverse Simpson indices at the genus level compared to UWTC (P < 0.01) but lower indices at the species level (P < 0.01).

To assess differences in microbiota identified by the three classification methods (i.e., UWTC, AWTC, and MWTC), we performed beta diversity analysis using PERMANOVA. Significant differences (P = 0.001) were noted in microbiota composition depicted by the full-length amplicon sequences at the genus level with NCBI RefSeq and RDP, while no significant differences were noted with the other databases (Fig. 1A). At the species level, however, the microbiota compositions were significantly affected (P = 0.001) by the classifier used irrespective of databases (Fig. 1A). With respect to the microbiota composition revealed by the V3-V4 amplicon sequences, significant differences were identified at both the genus and the species levels across all databases (P = 0.001) (Fig. 1B). Pairwise comparisons revealed that with the GG2, microbiota profiles classified using UWTC and MWTC were similar, whereas AWTC produced distinct microbiota profiles. Significant differences in microbiota profiles were observed among UWTC, AWTC, and MWTC with all other databases (Table S3).

Fig. 1.

Fig. 1

PCA plots showing differences in rumen microbiomes among taxonomy classifier types within each database at genus and species levels for full-length amplicon sequences (A) and for V3-V4 amplicon sequences (B) in an in vivo study. Significant differences were observed for all classifiers except for NCBI RefSeq and RDP at the genus level.

The classifiers were evaluated for taxonomic classification of key rumen microbial taxa and the proportion of successfully annotated taxa, focusing on the top 20 most dominant genera and species. Across all databases, MWTC yielded higher proportion of the top 20 dominant taxa compared to UWTC except at the genus level in full-length sequence with GG2 and SILVA, while maintaining the same taxonomic composition (Fig. 2, S1). With NCBI RefSeq and RDP, AWTC assigned some of the full-length and V3-V4 amplicon sequences to Segatella copri, whereas MWTC and UWTC assigned these sequences to Xylanibacter ruminicola (Fig. 2A, S1B). With other databases, UWTC and MWTC identified similar dominant taxa. With GG2 and GTDB, over 70% of the top 20 taxa were classified as “sp.,” whereas with SILVA, more than 80% of the top 20 taxa were annotated as “uncultured”. Notably, with SILVA at the species level, AWTC predominantly assigned uncultured microorganisms as “bacterium,” whereas MWTC annotated them as “rumen”.

Fig. 2.

Fig. 2

Taxonomic barplot showing the top 20 taxa for each taxonomy classifier in an in vivo study. Taxonomy classification was performed at the species level using full-length amplicon sequences (A) and at the genus level using V3-V4 amplicon sequences (B).

In terms of classification counts and fully classified ratios, MWTC outperformed UWTC and AWTC with both the full-length amplicon sequences and V3-V4 amplicon sequences across all databases (Figs. 3 and 4). Specifically, MWTC exhibited similar or higher classification counts at the genus level although it showed lower or no difference in classification counts at the phylum and family levels compared to UWTC and AWTC with the GG2 and SILVA databases. At the species level, MWTC increased classification counts by over 5% and 40% with the GG2 and the SILVA databases, respectively, compared with UWTC (P < 0.001). Similarly, with NCBI RefSeq and RDP, MWTC demonstrated a > 20% increase in higher classification counts and fully classified ratios over UWTC and even greater increases over AWTC at both the genus and species levels (P < 0.001) (Tables S4 and S5). Notably, with SILVA 138.1 version, AWTC performed worse than UWTC with the full-length amplicon sequences. With the V3-V4 amplicon sequences, AWTC also performed worse than UWTC at the genus level but slightly better than MWTC at the species level (Tables S4 and S5).

Fig. 3.

Fig. 3

Counts of classified ASVs at the phylum, family, genus, and species levels for each database (A) and the fully classified ratios (the proportion of completely classified features relative to the total ASVs) for each sample (B), identified using the full-length amplicon sequences in an in vivo study. Classifiers were categorized as follows: UWTC, a classifier without any adjustments; AWTC, a classifier based on average data from the EMPO3 dataset; and MWTC, a classifier manually curated using metagenomic and amplicon sequencing data.

Fig. 4.

Fig. 4

Counts of classified ASVs at the phylum, family, genus, and species levels for each database (A) and the fully classified ratios (the proportion of completely classified features relative to the total ASVs) for each sample (B), identified using the V3-V4 amplicon sequences in an in vivo study. Classifiers were categorized as follows: UWTC, a classifier without any adjustments; AWTC, a classifier based on average data from the EMPO3 dataset; and MWTC, a classifier manually curated using metagenomic and amplicon sequencing data.

Classifiers evaluation using in vitro experimental samples

In vitro experiments are often used to simulate the rumen environment when conducting preliminary comparisons or when large-scale animal trials are not feasible. Therefore, since 16S rRNA amplicon sequences derived from in vitro experiments also require improved resolution through MWTC, we validated MWTC using in vitro experimental samples.

The alpha diversity indices demonstrated notable differences across classifiers (Tables S6 and S7). Observed features were statistically significant across all classifiers (P < 0.05), except for GTDB with full-length amplicon sequences. Consistent with the in vivo datasets, NCBI RefSeq and RDP demonstrated higher observed features with MWTC compared to UWTC (P < 0.05), whereas GG2 and SILVA exhibited lower or similar values. Shannon, Simpson, and inverse Simpson indices were significant in both full-length and V3-V4 amplicon sequences, except for GG2 and SILVA at the genus level in full-length sequences (P < 0.05). When used with NCBI RefSeq and RDP, MWTC yielded higher values in these indices compared to UWTC and AWTC, irrespective of sequencing types or taxonomic levels. In contrast, coupled with GG2 and SILVA, MWTC exhibited lower or similar values than UWTC. AWTC displayed inconsistent trends, with the values fluctuating, relative to UWTC.

The beta diversity analysis revealed no significant differences as determined with the full-length amplicon sequences at the genus level with the two SILVA database releases (Fig. S2A). In contrast, other genus-level data, except for the SILVA databases, and all species-level data were statistically different in beta diversity (P < 0.05) (Fig. S2A). Both genus- and species-level data determined by the V3-V4 amplicon sequences differed statistically significantly across all databases (Fig. S2B), consistent with the in vivo results (Table S8).

Similar to the in vivo dataset, MWTC achieved the highest proportion accounted for by the top 20 dominant taxa compared to UWTC (Fig. S3, S4). MWTC showed an overall higher microbial abundance than UWTC. Across all the classifiers, at the species level, with GG2 and GTDB, 75% of the top 20 taxa were classified as “sp.,” whereas with SILVA, 90% of the top 20 taxa were annotated as “uncultured” or “unidentified”.

Regarding classification counts and fully classified ratios, MWTC consistently outperformed UWTC at both the genus and species levels (Fig. S5, S6). With the V3-V4 amplicon sequences, AWTC and MWTC yielded lower classification counts than UWTC at the phylum and family levels when the GG2 and SILVA databases were used. Similarly, with the full-length sequences and the V3-V4 amplicon sequences, AWTC showed lower values in classification counts than UWTC at the genus and species levels when GG2 was used, with a 3.4–5.5% reduction at the genus level and a 5.6–9.5% reduction at the species level. Similarly, when used together with SILVA, AWTC showed a 0.9–9.8% decrease in classification counts at the genus level. In contrast, while MWTC demonstrated lower values in classification counts than UWTC, these were statistically comparable to UWTC (Tables S9 and S10). Interestingly, with the NCBI RefSeq and RDP, MWTC exhibited fully classified ratio at the species level, with improvements ranging from 19.2 to 64.6% in NCBI RefSeq and 22.8–58.3% in RDP, which exceeded that achieved by UWTC.

Error rate estimate

The error rates were thoroughly analyzed, considering variations across all databases for each amplicon dataset resulting from both in vivo and in vitro environments (Fig. 5). With the NCBI RefSeq, MWTC showed a relatively decreased error rate compared to UWTC, averaging by 33.9–62.6% at the genus level and 22.6–47.5% at the species level. Similarly, with RDP, reductions of 25.2–53.5% and 15.7–24.9% were observed at the genus and species levels, respectively. The GTDB database yielded a decrease of 1.7–4.4% at the genus level and 2.9–8.0% at the species level. With GG2 and SILVA, reductions were minimal, with a maximum decrease of only 1.2% at the species level. Overall, MWTC exhibited lower error rates than UWTC (P < 0.05), except at the phylum level with the V3-V4 amplicon sequences. Notably, the average error rate at the genus and species levels was higher with AWTC than with UWTC but lower with MWTC (Table S11).

Fig. 5.

Fig. 5

Error rate (% difference) plots for each classification count, including in vivo experimental datasets from full-length amplicon sequences (A) and the V3-V4 amplicon sequences (B) and in vitro datasets from the full-length amplicon sequences (C) and the V3-V4 amplicon sequences (D). Error rates were calculated by comparing the results with those of the BLAST 16S ribosomal RNA database, considering both unnamed and misnamed classifications as errors.

Discussion

Advancements in sequencing technologies and taxonomy classifiers have improved microbiome characterization, yet existing databases remain heavily human-centric, limiting analysis accuracies of non-human microbiomes like the rumen. While weighted taxonomy classifiers enhance microbial identification, their application to rumen samples may be less precise due to the absence of ruminal microbiota in EMPO habitat types and the unique microbial profiles influenced by ruminant species and diet. This underscores the need for ruminant-specific taxonomy classifiers to improve classification accuracy. To enhance the accuracy of taxonomic classification in rumen microbiome analyses, we suggested assigning weights during the classification process. However, since many microbiome databases do not contain sufficient rumen-specific data, we proposed the development of manually weighted datasets using both shotgun metagenomics and amplicon sequencing data from specific ruminant breeds.

MWTC achieved higher classification counts, fully classified ratios, and a lower error rate than UWTC, demonstrating its capability to provide more accurate annotations, particularly at lower taxonomic levels. This suggests that even with the same amplicon sequence variant (ASV) dataset, beta diversity analysis results could vary depending on the classifier used, which assigns different taxonomy. This further highlighted that classifiers with different taxonomy weighting methods influence the microbial community profiles differently.

MWTC preserves or increases the relative abundance of key microbial taxa present in the rumen samples. With NCBI RefSeq and RDP, AWTC reclassified Xylanibacter ruminicola, classified by UWTC, as Segatella copri. Segatella copri accounts for approximately 30% of the human gut microbiome14,15, but it is not dominant in the rumen. In contrast, MWTC retained the annotation of Xylanibacter ruminicola and further enhanced its relative abundance. Additionally, MWTC showed Xylanibacter, Succiniclasticum, Succinivibrio, Sodaliphilus, Ruminococcoides, Butyrivibrio, Treponema, and Fibrobacter, all of which are common genera found in the rumen3,10,16, without a decline in abundance compared to UWTC. These suggest that MWTC provides a more accurate classification for rumen microbiota compared to the weighting approach used in AWTC.

A previous study had shown that using shotgun metagenome data as a weighting factor increased the taxonomy detection rate9. In this study, by incorporating not only shotgun sequencing datasets but also amplicon sequencing datasets as taxonomic weights, MWTC demonstrated higher classification counts and fully classified ratios, along with reduced error rates, than UWTC and AWTC. This indicated that amplicon sequencing datasets can be utilized to provide taxonomic weights while maintaining high resolution. However, not all the amplicon datasets are suitable for this purpose. To construct the MWTC, amplicon datasets derived from DNA collected directly from the ruminal fluid of individual animals were utilized. In contrast, amplicon datasets obtained from the in vitro fermentation process may not be ideal for use as weighting datasets, since they originate from mixed cultures gathered from multiple individuals17,18 and may include microorganisms introduced during the process of preparing buffer solutions19,20.

When GG2 and SILVA were used, no differences in beta diversity or error rate were noted. This finding could be attributed to the similar taxonomies implemented in GG2 and SILVA. In GG2, ambiguous annotations, such as ‘CAG,’ ‘RUG (Rumen uncultured genus)’ and ‘sp.’ are frequently observed, while in SILVA, annotations like ‘uncultured,’ ‘unidentified,’ or ‘unknown organism’ are common, both presenting challenges in identifying microbes at lower taxonomic levels and compromising taxonomic and phylogenetic rigor (Supplementary data)21,22. Despite AWTC and MWTC showed higher classification count and fully classified ratio compared to UWTC, the error rate did not decrease in the GG2 and SILVA databases. The AWTC showed increased error rates compared to the UWTC at both the genus and species levels. However, the MWTC demonstrated lower error rates than the AWTC and overall showed similar or lower error rates than UWTC. This suggests that the MWTC maintained a resolution comparable to that of the UWTC when GG2 and SILVA were used.

The GTDB, which serves as the underlying database for GG2, shares the same classification framework as GG223. Similar to GG2, GTDB reflects the characteristics of the database with annotations like ‘CAG,’ ‘RUG (Rumen uncultured genus)’ and ‘spp.’ However, unlike GG2 and SILVA, GTDB did not yield lower classification counts at the phylum or family levels and showed a slight improvement in error rates. There were significant improvements in the V3-V4 amplicon sequences, which may have a lower resolution than the full-length amplicon sequences. Therefore, integration of MWTC with GTDB as the database could be used in future studies.

In the cases of NCBI RefSeq and RDP, both classification count and fully classified ratio showed a significant increase by MWTC, accompanied by a reduction in error rates up to 47.5% at the species level. Although the two databases showed lower classification counts by up to two to three times at the genus and species levels compared to other databases when UWTC was used, applying MWTC narrowed the gap in classification counts, improving to a level achieved by the other databases. Furthermore, the error rate was remarkably enhanced, exceeding the performance of the other databases. The NCBI RefSeq used in this analysis contained only 16S rRNA gene sequences annotated at the species level, excluding the genomes of uncultured microbes. Similarly, RDP was derived from genomic data provided by NCBI GenBank, European Molecular Biology Laboratory (EMBL), and DNA Database Bank of Japan (DDBJ)24, and its results were highly consistent with those of NCBI RefSeq. Thus, we confirmed that using a weighted taxonomy classifier is highly effective, especially in those that exclude uncultured genomes, and its impact becomes even more pronounced when applied to species-specific analyses, such as ruminants. Unlike other databases that predominantly contain uncultured microbes, these NCBI RefSeq and RDP enable the assignment of microbial individuals to lower taxonomic levels, thereby being suitable for increasingly detailed studies on microbes and their communities.

This study applied taxonomic weighting based on datasets derived exclusively from Hanwoo cattle. Given that microbial community composition can vary significantly across breeds25,26, further validation using datasets from diverse breeds is necessary to assess the generalizability of the MWTC for breed-specific microbiome analyses. Additionally, classifier performance in this study was evaluated using the V3–V4 hypervariable region of the 16S rRNA gene. As other hypervariable regions were not tested, additional validation across different regions is warranted. It is also important to acknowledge that the observed outcomes may be partially influenced by database-specific quality and sample-dependent biases inherent in the datasets used for evaluation. Nevertheless, the overall improvement in taxonomic resolution observed with MWTC suggests that its application would likely enhance classification performance.

Conclusion

With the continuous advancement of microbiome analysis techniques, understanding the composition and structure of the ruminal microbial community has become increasingly important. In this context, the development and implementation of ruminant breed-specific taxonomic classifiers have been recognized as a critical need. To minimize the influence of various environmental factors that may alter rumen microbiome composition, this study applied breed-specific weights for taxonomic classification. Overall, applying MWTC enhanced taxonomic classification resolution in 16S rRNA gene-based microbiome analysis compared with AWTC, which applies taxonomic weights based on the EMPO database. Notably, the current study confirmed that taxonomic weights could be applied using both shotgun metagenome and amplicon sequence datasets. Among the databases evaluated, the application of MWTC enhanced the performance of NCBI RefSeq, supporting its use for species-level classification. This enhancement facilitated species-level microbial analysis, making it more effective than other databases that rely on ambiguous annotations. Thus, we recommend using NCBI RefSeq for species-level classification, and its performance will significantly improve when applying MWTC.

Methods

Construction of unweighted and weighted taxonomy classifiers

All data were analyzed using the QIIME2-amplicon version 2024.1027. To construct the UWTC, reference databases for bacteria and archaea were retrieved using q2-rescript21, including the NCBI RefSeq (16S rRNA; October 14, 2024), SILVA (versions 138.1 and 138.2)28, GTDB (R220)29, GG2 (versions 2022.10 and 2024.09), and the RDP bacterial and archaeal hierarchy model (version 2.14; August 2023)30. These databases were subsequently used to construct the taxonomy classifier. Archaeal sequences shorter than 900 bp and bacterial sequences shorter than 1,200 bp were filtered from all databases to enhance taxonomic reliability and resolution. Additionally, all the database sequences were dereplicated within each database to eliminate redundancy.

To construct an AWTC, metadata were developed using the EMPO31 for assigning average taxonomic weights. Publicly available 16S rRNA gene sequences corresponding to the V4 region were retrieved from Qiita using q2-clawback9. In this resource, the sequences are represented as 150 bp Illumina reads derived from the ~ 250 bp V4 region, together with the associated EMPO3 metadata. The following keywords were used to search and download the EMPO3 metadata: ‘animal-corpus,’ ‘animal-distal-gut,’ ‘animal-proximal-gut,’ ‘animal-secretion,’ ‘animal-surface,’ ‘plant-corpus,’ ‘plant-rhizosphere,’ ‘plant-surface,’ ‘sediment-non-saline,’ ‘sediment-saline,’ ‘soil-non-saline,’ ‘surface-saline,’ ‘water-non-saline,’ ‘water-saline,’ ‘animal-non-saline,’ and ‘animal-saline’.

Metagenomic shotgun and 16S rRNA gene amplicon datasets derived from rumen fluid samples collected directly from Hanwoo cattle were used to construct MWTC. The characteristics of each dataset are presented in Table 1. The metagenomic shotgun dataset from Hanwoo steers (36 samples) and cows (5 samples) was preprocessed using fastp32 for quality control, followed by sequence filtering with Bowtie2 (version 2.5.4)33 to remove host-derived sequences by aligning reads to the Hanwoo genome (GCA_028973685.2) and feed-derived sequences by aligning to the genomes of feed ingredients consumed by each animal group. For Hanwoo steers, feed-derived sequences were filtered using the genomes of oat hay (GCA_916181665.1), corn (GCF_902167145.1), rice straw (GCF_001433935.1), wheat (GCF_018294505.1), and palm kernel meal (GCF_000442705.1). For cows, filtering was performed using the genomes of oat hay, corn, and rice straw. The filtering ensured the isolation of microbial DNA sequences by removing host and feed-related sequences. The filtered metagenomic datasets were processed using SortMeRNA (version 4.3.7)34 to extract the rRNA gene sequences. The resultant rRNA sequences were directly imported into QIIME2, and a feature table was generated after sequence dereplication. For the 16S rRNA gene amplicon sequencing data from Illumina Miseq runs, primer sequence were removed using Cutadapt35, and the paired-end reads were then merged using FLASH2 (version 2.2.00)36 with a minimum overlap of 20 bp. The processed reads were subsequently imported into QIIME2, followed by denoising with Deblur37 using default options. For the PacBio sequencing run, raw reads were directly imported into QIIME2 and denoised using DADA2 denoise-ccs38 with default options. The average number of raw reads and reads after each preprocessing step per sample are presented in Table 1.

Table 1.

Summary of the datasets used to generate the manually weighted taxonomy classifier, and number of reads after each preprocessing.

NCBI BioProject Platform Number of samples Number of raw reads1) Number of reads after fastp filtering Number of reads after SortMeRNA Number of reads after dereplication or denoising References
Shotgun metagenomic sequencing
PRJNA1242891 Illumina HiSeq 36 40,606,291 ± 1,705,341 40,242,984 ± 1,694,884 363,163 ± 59,436 363,025 ± 59,416 Unpublished
PRJNA1240938 Illumina HiSeq 5 23,221,950 ± 2,468,085 23,069,977 ± 2,427,414 389,962 ± 79,863 388,852 ± 79,987 Unpublished
Amplicon sequencing
PRJEB25166 Illumina MiSeq 9 116,553 ± 6,690 - - 30,085 ± 2,475 Song et al., 201852
PRJNA523867 Illumina MiSeq 25 222,464 ± 20,635 - - 54,052 ± 4,803 Kim et al., 202053
PRJNA725944 Illumina MiSeq 8 134,859 ± 37,939 - - 47,213 ± 13,503 Bharanidharan et al., 202154
PRJNA797687 Illumina MiSeq 184 106,030 ± 44,380 - - 33,261 ± 16,056 Unpublished
- PacBio Sequel IIe 15 143,702 ± 14,337 - - 81,209 ± 8,773

Kang et al., 202455

(Data availability: 10.5281/zenodo.15123162)

1) All the numbers of reads are shown as mean ± SD.

NCBI, national center for biotechnology information; SD, standard deviation; PacBio, pacific biosciences.

A custom Python script was used to combine the feature tables and the corresponding representative sequences to generate the input file for taxonomy weight generation. Sequence variants from combined datasets were then generated using q2-clawback. Sequences associated with chloroplasts and mitochondria were removed before generating the classifier to ensure accuracy. These datasets incorporated all databases used to create the classifier, including AWTC and UWTC generated with the naïve Bayes algorithm using the ‘feature-classifier’ plugin39. The process for constructing the weighted taxonomy classifier is detailed in Fig. 6 and detailed code and implementation can be accessed in GitHub repository (https://github.com/6seok/rumanclass).

Fig. 6.

Fig. 6

The workflow used to construct each weighted taxonomy classifier. The bioinformatics tools used were italicized.

Amplicon datasets for validation of taxonomy classifier

16S rRNA gene amplicon sequencing data were collected from rumen samples of Hanwoo cattle to validate the taxonomy classifiers. Briefly, rumen fluid samples were directly collected via stomach tubing from the rumen following an in vivo experiment and in vitro fermentation for 24 h. DNA samples extracted from all rumen fluid samples were isolated using the repeated bead-beating and column method40. The nearly full-length and the V3-V4 hypervariable region of the 16S rRNA gene were amplified using the primers 27F (5′-AGRGTTYGATYMTGGCTCAG-3′) and 1492R (5′-GYTACCTTGTTACGACTT-3′) and 341F (5′-CCTACGGGNGGCWGCAG-3′) and 805R (5′-GACTACHVGGGTATCTAATCC-3′), producing amplicons of full-length and V3-V4 regions, respectively. Full-length amplicon samples were sequenced using the PacBio Sequel platform (Pacific Biosciences, CA, USA), while V3–V4 amplicon samples were sequenced on the Illumina MiSeq platform (Illumina, San Diego, CA, USA). A total of 47 in vivo samples and five in vitro samples were used for validating the classifiers with the full-length amplicons, while 30 in vivo samples and four in vitro samples were analyzed for the V3-V4 amplicons.

Data processing

The amplicon datasets for validation were classified into 21 classifiers using scikit-learn41, employing three classification methods—UWTC, AWTC, and MWTC—applied across the seven databases. For comparisons of the overall microbiota composition between the databases and among UWTC, AWTC, and MWTC, principal component analysis (PCA) was conducted using Bray-Curtis dissimilarity matrices at both the genus and species levels for both the full-length and the V3-V4 amplicon sequences. Alpha diversity metrics, including observed features42, evenness43, Shannon44, Simpson, and inversed Simpson diversity indices45 were also computed at these levels. Prior to these analyses, the data from all samples were normalized using Total Sum Scaling (TSS). Classification counts were calculated for all ASVs only when there was no blank classification at each taxonomic level, from the phylum to the species. To address cases where certain ASVs were not fully classified at a specific taxonomic level, the ratio of fully classified taxa was calculated as the number of fully classified taxa at each level relative to the total taxa count. This metric, referred to as the fully classified ratio, was evaluated for each database at both the genus and species levels to assess classification efficiency.

Error rate estimation

Error rates for taxonomic classification were calculated following the approach outlined by Kaehler et al.9 to determine the proportion of ASVs misclassified by classifiers. The BLAST 16S ribosomal RNA database46 was used as a reference, and all ASVs were classified to the species level without considering confidence scores. Applying a confidence score may lead to the classification of certain ASVs at higher taxonomic levels only, whereas not using a confidence score ensures full classification down to the species level. However, significant updates in taxonomy occurred after the release of GG2 2022.10 and SILVA 138.1, and in these updates some taxa were reclassified at both the phylum (e.g., Firmicutes to Bacillota, Bacteroidetes to Bacteroidota, Proteobacteria to Pseudomonadota, and Actinobacteria to Actinomycetota) and the genus levels (e.g., certain Prevotella species reclassified to Xylanibacter and Segatella, and Propionibacterium acnes renamed as Cutibacterium acnes)47,48. Thus, error rates were not calculated for these databases.

Statistical analysis

The effects of classifier types on overall microbiota composition analysis were statistically evaluated using permutational multivariate analysis of variance (PERMANOVA). This analysis used the vegan and pairwiseAdonis packages in R (version 4.3.3) with 9,999 random permutations49,50. Multiple testing adjustments were applied using the Benjamini-Hochberg correction method51. Normality of alpha diversity metrics, classification counts, fully classified ratios, and error rates were assessed using the Shapiro-Wilk test, while homogeneity of variances was evaluated using Levene’s test. All data satisfied the assumption of homogeneity of variances. Statistical analysis was conducted using PROC MIXED in SAS 9.4 (SAS Institute Inc., Cary, NC, USA) for normally distributed data, whereas PROC GLIMMIX was used for the data that did not follow a normal distribution. Post hoc pairwise comparisons were performed based on least squares means and adjusted for multiple testing using Tukey’s method. Classifier type was treated as a fixed effect. For the alpha diversity metrics, classification counts and fully classified ratios, each amplicon dataset was included as a random effect, while for the error rates, both amplicon dataset and database types were treated as random effects. Statistical significance was declared at s threshold of P ≤ 0.05.

Supplementary Information

Below is the link to the electronic supplementary material.

Supplementary Material 1 (21.7MB, docx)

Author contributions

R.K.: Original draft, Conceptualization, Visualization, Software, Methodology, and Data curation. Z.Y.: Review & Editing, Validation. H.K.: Review & Editing, Validation. J.S.: Review & Editing. M.K.: Review & Editing. T.P.: Conceptualization, Review & Editing, Validation. All authors reviewed the final version.

Funding

This research was supported by the Chung-Ang University Research Grants in 2023.

Data availability

The sequencing data analyzed in this study have been deposited in the NCBI Sequence Read Archive (SRA) under the following BioProject accession numbers: PRJNA1244067 and PRJNA1244069 for in vivo long-read sequencing, PRJNA1242525 for in vivo V3-V4 sequencing, PRJNA1098854 for in vitro long-read sequencing, and PRJNA1240950 for in vitro V3-V4 sequencing. All codes and prebuilt classifiers used for data analysis are available at https://github.com/6seok/rumanclass.

Declarations

Competing interests

The authors declare no competing interests.

Footnotes

Publisher’s note

Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.

Change history

12/9/2025

The original online version of this Article was revised: The Funding section in the original version of this Article was incorrect. It now reads: This research was supported by the Chung-Ang University Research Grants in 2023.

References

  • 1.Xu, Q. et al. Gut microbiota and their role in health and metabolic disease of dairy cow. Front. Nutr.810.3389/fnut.2021.701511 (2021). [DOI] [PMC free article] [PubMed]
  • 2.Russell, J. B., Muck, R. E. & Weimer, P. J. Quantitative analysis of cellulose degradation and growth of cellulolytic bacteria in the rumen. FEMS Microbiol. Ecol.67, 183–197. 10.1111/j.1574-6941.2008.00633.x (2009). [DOI] [PubMed] [Google Scholar]
  • 3.Qi, W. et al. — Invited Review — Understanding the functionality of the rumen microbiota: searching for better opportunities for rumen microbial manipulation. Anim. Biosci.37, 370–384. 10.5713/ab.23.0308 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 4.Pollock, J., Glendinning, L., Wisedchanwet, T. & Watson, M. The madness of microbiome: attempting to find consensus best practice for 16S Microbiome studies. Appl. Environ. Microbiol.84, e02627–e02617. 10.1128/AEM.02627-17 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 5.Seshadri, R. et al. Cultivation and sequencing of rumen Microbiome members from the Hungate1000 collection. Nat. Biotechnol.36, 359–367. 10.1038/nbt.4110 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6.HoloRuminant Consortium. HoloRuminant – Understanding Microbiomes of the Ruminant holobiont. 2021–2025 (European Commission, 2021).
  • 7.Wilkinson, T. J. et al. CowPI: A rumen Microbiome focussed version of the PICRUSt functional inference software. Front. Microbiol.910.3389/fmicb.2018.01095 (2018). [DOI] [PMC free article] [PubMed]
  • 8.McGovern, E., Waters, S. M., Blackshields, G. & McCabe, M. S. Evaluating established methods for rumen 16S rRNA amplicon sequencing with mock microbial populations. Front. Microbiol.910.3389/fmicb.2018.01365 (2018). [DOI] [PMC free article] [PubMed]
  • 9.Kaehler, B. D. et al. Species abundance information improves sequence taxonomy classification accuracy. Nat. Commun.10, 4643. 10.1038/s41467-019-12669-6 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 10.Henderson, G. et al. Rumen microbial community composition varies with diet and host, but a core Microbiome is found across a wide geographical range. Sci. Rep.5, 14567. 10.1038/srep14567 (2015). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 11.Islam, M. et al. Holstein and Jersey steers differ in rumen microbiota and enteric methane emissions even fed the same total mixed ration. Front. Microbiol.1210.3389/fmicb.2021.601061 (2021). [DOI] [PMC free article] [PubMed]
  • 12.Menke, K. H. et al. The Estimation of the digestibility and metabolizable energy content of ruminant feedingstuffs from the gas production when they are incubated with rumen liquor in vitro. J. Agric. Sci.93, 217–222. 10.1017/S0021859600086305 (1979). [Google Scholar]
  • 13.Shaw, C. A. et al. A comparison of three artificial rumen systems for rumen Microbiome modeling. Fermentation9, 953 (2023). [Google Scholar]
  • 14.Xiao, X., Singh, A., Giometto, A. & Brito, I. L. Segatella copri strains adopt distinct roles within a single individual’s gut. bioRxiv, 2024.2005.2020.59501510.1101/2024.05.20.595015(2024).
  • 15.Panwar, D., Briggs, J., Fraser, A. S. C., Stewart, W. A. & Brumer, H. Transcriptional delineation of polysaccharide utilization loci in the human gut commensal Segatella Copri DSM18205 and co-culture with exemplar bacteroides species on dietary plant glycans. Appl. Environ. Microbiol.91, e0175924. 10.1128/aem.01759-24 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16.Buckner, A. M. et al. The selective culture and enrichment of major rumen bacteria on three distinct anaerobic culture media. bioRxiv, 2024.2008.2023.608987 (2025). 10.1101/2024.08.23.608987 [DOI] [PMC free article] [PubMed]
  • 17.Kang, R. et al. Effects of diets for three growing stages by rumen inocula donors on in vitro rumen fermentation and Microbiome. J. Anim. Sci. Technol.10.5187/jast.2023.e109 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 18.Jeong, J., Yu, C., Kang, R., Kim, M. & Park, T. Application of propionate-producing bacterial consortium in ruminal methanogenesis inhibited environment with bromoethanesulfonate as a methanogen direct inhibitor. Front. Vet. Sci.1110.3389/fvets.2024.1422474 (2024). [DOI] [PMC free article] [PubMed]
  • 19.McDougall, E. I. Studies on ruminant saliva. 1. The composition and output of sheep’s saliva. Biochem. J.43, 99–109 (1948). [PMC free article] [PubMed] [Google Scholar]
  • 20.Goering, H. Forage fiber analyses (apparatus, reagents, procedures, and some applications). (1970).
  • 21.Robeson, M. S. et al. RESCRIPt: Reproducible sequence taxonomy reference database management. PLoS Comput Biol 17, e1009581 (2021). 10.1371/journal.pcbi.1009581 [DOI] [PMC free article] [PubMed]
  • 22.Ceccarani, C. & Severgnini, M. A comparison between Greengenes, SILVA, RDP, and NCBI reference databases in four published microbiota datasets. BioRxiv2023.2004.2012.53586410.1101/2023.04.12.535864 (2023).
  • 23.McDonald, D. et al. Greengenes2 unifies microbial data in a single reference tree. Nat. Biotechnol.42, 715–718. 10.1038/s41587-023-01845-1 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 24.Cole, J. R. et al. The ribosomal database project (RDP-II): previewing a new autoaligner that allows regular updates and the new prokaryotic taxonomy. Nucleic Acids Res.31, 442–443. 10.1093/nar/gkg039 (2003). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 25.De Mulder, T. et al. Impact of breed on the rumen microbial community composition and methane emission of Holstein Friesian and Belgian blue heifers. Livest. Sci.207, 38–44 (2018). [Google Scholar]
  • 26.Li, F., Hitch, T. C., Chen, Y., Creevey, C. J. & Guan, L. L. Comparative metagenomic and metatranscriptomic analyses reveal the breed effect on the rumen Microbiome and its associations with feed efficiency in beef cattle. Microbiome7, 1–21 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 27.Bolyen, E. et al. Reproducible, interactive, scalable and extensible Microbiome data science using QIIME 2. Nat. Biotechnol.37, 852–857. 10.1038/s41587-019-0209-9 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 28.Quast, C. et al. The SILVA ribosomal RNA gene database project: improved data processing and web-based tools. Nucleic Acids Res.41, D590–D596. 10.1093/nar/gks1219 (2012). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 29.Parks, D. H. et al. GTDB: an ongoing census of bacterial and archaeal diversity through a phylogenetically consistent, rank normalized and complete genome-based taxonomy. Nucleic Acids Res.50, D785–D794. 10.1093/nar/gkab776 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 30.Wang, Q. & Cole, J. R. Updated RDP taxonomy and RDP classifier for more accurate taxonomic classification. Microbiol. Resour. Ann.13, e01063–e01023. 10.1128/mra.01063-23 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 31.Thompson, L. R. et al. A communal catalogue reveals earth’s multiscale microbial diversity. Nature551, 457–463. 10.1038/nature24621 (2017). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 32.Chen, S., Zhou, Y., Chen, Y. & Gu, J. Fastp: an ultra-fast all-in-one FASTQ preprocessor. Bioinformatics34, i884–i890. 10.1093/bioinformatics/bty560 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 33.Langmead, B. & Salzberg, S. L. Fast gapped-read alignment with bowtie 2. Nat. Methods. 9, 357–359. 10.1038/nmeth.1923 (2012). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 34.Kopylova, E., Noé, L. & Touzet, H. SortMeRNA: fast and accurate filtering of ribosomal RNAs in metatranscriptomic data. Bioinformatics28, 3211–3217. 10.1093/bioinformatics/bts611 (2012). [DOI] [PubMed] [Google Scholar]
  • 35.Martin, M. Cutadapt removes adapter sequences from high-throughput sequencing reads. EMBnet J.17, 3. 10.14806/ej.17.1.200 (2011). [Google Scholar]
  • 36.Magoč, T. & Salzberg, S. L. FLASH: fast length adjustment of short reads to improve genome assemblies. Bioinformatics27, 2957–2963. 10.1093/bioinformatics/btr507 (2011). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 37.Amir, A. et al. Deblur rapidly resolves single-nucleotide community sequence patterns. mSystems 2, 10.1128/msystems.00191 – 00116 (2017). 10.1128/msystems.00191-16 [DOI] [PMC free article] [PubMed]
  • 38.Callahan, B. J. et al. DADA2: High-resolution sample inference from illumina amplicon data. Nat. Methods. 13, 581–583. 10.1038/nmeth.3869 (2016). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 39.Wang, Q., Garrity, G. M., Tiedje, J. M. & Cole, J. R. Naive bayesian classifier for rapid assignment of rRNA sequences into the new bacterial taxonomy. Appl. Environ. Microbiol.73, 5261–5267. 10.1128/aem.00062-07 (2007). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 40.Yu, Z. & Morrison, M. Improved extraction of PCR-quality community DNA from digesta and fecal samples. BioTechniques36, 808–812. 10.2144/04365ST04 (2004). [DOI] [PubMed] [Google Scholar]
  • 41.Pedregosa, F. et al. Scikit-learn: machine learning in python. J. Mach. Learn. Res.12, 2825–2830 (2011). [Google Scholar]
  • 42.DeSantis, T. Z. et al. Greengenes, a chimera-checked 16S rRNA gene database and workbench compatible with ARB. Appl. Environ. Microbiol.72, 5069–5072. 10.1128/aem.03006-05 (2006). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 43.Pielou, E. C. The measurement of diversity in different types of biological collections. J. Theor. Biol.13, 131–144. 10.1016/0022-5193(66)90013-0 (1966). [Google Scholar]
  • 44.Shannon, C. E. A mathematical theory of communication. Bell Syst. Tech. J.27, 379–423. 10.1002/j.1538-7305.1948.tb01338.x (1948). [Google Scholar]
  • 45.Simpson, E. Measurement of diversity. Nature16310.1038/163688a0 (1949).
  • 46.Camacho, C. et al. BLAST+: architecture and applications. BMC Bioinform.10, 421. 10.1186/1471-2105-10-421 (2009). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 47.National Center for Biotechnology Information. Prokaryote phyla added to the NCBI taxonomy database, (2021). https://ncbiinsights.ncbi.nlm.nih.gov/2021/12/10/ncbi-taxonomy-prokaryote-phyla-added/
  • 48.Hitch, T. C. A. et al. A taxonomic note on the genus prevotella: description of four novel genera and emended description of the genera Hallella and xylanibacter. Syst. Appl. Microbiol.45, 126354. 10.1016/j.syapm.2022.126354 (2022). [DOI] [PubMed] [Google Scholar]
  • 49.Martinez Arbizu, P. pairwiseAdonis: Pairwise multilevel comparison using adonis, (2020). https://github.com/pmartinezarbizu/pairwiseAdonis
  • 50.Dixon, P. VEGAN, a package of R functions for community ecology. J. Veg. Sci.14, 927–930. 10.1111/j.1654-1103.2003.tb02228.x (2003). [Google Scholar]
  • 51.Benjamini, Y. & Hochberg, Y. Controlling the false discovery Rate - a practical and powerful approach to multiple testing. J. Royal Stat. Soc. Ser. B-Statistical Methodol.57, 289–300. 10.1111/j.2517-6161.1995.tb02031.x (1995). [Google Scholar]
  • 52.Song, J. et al. Effects of sampling techniques and sites on rumen Microbiome and fermentation parameters in Hanwoo steers. J. Microbiol. Biotechnol.28, 1700–1705. 10.4014/jmb.1803.03002 (2018). [DOI] [PubMed] [Google Scholar]
  • 53.Kim, M., Park, T., Jeong, J. Y., Baek, Y. & Lee, H. J. Association between rumen microbiota and marbling score in Korean native beef cattle. Animals10, 712. 10.3390/ani10040712 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 54.Bharanidharan, R. et al. Feeding systems and host breeds influence ruminal fermentation, methane production, microbial diversity and metagenomic gene abundance. Front. Microbiol.12, 701081. 10.3389/fmicb.2021.701081 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 55.Kang, R. et al. Impact of Forage Sources on Ruminal Bacteriome and Carcass Traits in Hanwoo Steers During the Late Fattening Stages. Microorganisms 12, 2082 (2024). [DOI] [PMC free article] [PubMed]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Supplementary Material 1 (21.7MB, docx)

Data Availability Statement

The sequencing data analyzed in this study have been deposited in the NCBI Sequence Read Archive (SRA) under the following BioProject accession numbers: PRJNA1244067 and PRJNA1244069 for in vivo long-read sequencing, PRJNA1242525 for in vivo V3-V4 sequencing, PRJNA1098854 for in vitro long-read sequencing, and PRJNA1240950 for in vitro V3-V4 sequencing. All codes and prebuilt classifiers used for data analysis are available at https://github.com/6seok/rumanclass.


Articles from Scientific Reports are provided here courtesy of Nature Publishing Group

RESOURCES