Skip to main content
Scientific Reports logoLink to Scientific Reports
. 2025 Jul 4;15:23841. doi: 10.1038/s41598-025-08859-6

ez-CAZy a reference annotation database for linking glycoside hydrolase sequence to enzymatic activity

Daniel S Erdody 1, Nicholas G Griffin 1, Renaud Berlemont 1,
PMCID: PMC12229491  PMID: 40610571

Abstract

Glycoside Hydrolases (GH) are carbohydrate-active enzymes that play a critical role in the degradation of carbohydrates, impacting ecosystem function, human health, and biotechnological applications. The functional annotation of GHs within the CAZy database is hindered by the lack of sequence-specific definitions, annotation tools, and reliance on generalized “majority rule” activity assumptions. Here, we introduce ez-CAZy, a custom reference database designed to link GH sequences, as well as other CAZy, to their enzymatic activities. By reannotating over 7,000 biochemically characterized GHs using Hidden Markov Model profiles and other publicly accessible tools, we provide detailed sequence metadata, domain architectures, and functional predictions. Our analysis reveals the clustered distribution of enzymatic activities and multi-domain architectures within GH families, facilitating more precise functional predictions for newly identified sequences. The predictive accuracy of ez-CAZy was validated using over 500 recently characterized GHs, demonstrating functional annotation. ez-CAZy addresses critical gaps in GH annotation pipelines and offers a publicly accessible tool to support sequence analysis and enzymatic research. This work underscores the need for standardized enzyme characterization and expanded substrate testing to enhance annotation accuracy in future studies.

Supplementary Information

The online version contains supplementary material available at 10.1038/s41598-025-08859-6.

Keywords: CAZy, Carbohydrate active enzyme, Glycoside hydrolase, Cellulase, Amylase, Annotation

Subject terms: Glycomics, Databases, Sequence annotation

Introduction

Carbohydrate active enzymes (CAZy) are essential proteins supporting the synthesis, modification and deconstruction of carbohydrates. CAZy proteins are central to ecosystem functioning, human health and have many biotechnological applications (e.g., biofuels and biotherapeutics). CAZy proteins correspond to functional protein domains in either single-domain (SD) or multi-domain (MD) proteins. CAZy domains are divided into five classes based on their biochemical reactions: glycoside hydrolases (GH), polysaccharide lyases (PL), carbohydrate esterases (CE), auxiliary activities (AA), and glycosyl transferases (GT)1. In addition, accessory non-catalytic carbohydrate binding modules (CBM) help the associated CAZy domains bind to specific substrates [e.g2,3]. Initially, many CAZy families were named according to their described activities. However, as more proteins were characterized, new activities were described in previously identified CAZy families leading to conflicting names. For example, GH families 5 and 18 were originally named “cellulase” and “chitinase”4,5 however these families now include many proteins which are de facto not cellulolytic nor chitinolytic enzymes respectively. For example, some GH5 are xylanases6. Eventually, a naming convention focusing on the type of reaction being catalyzed (e.g., glycoside hydrolase) was adopted and implemented in the carbohydrate active enzyme (CAZy) database (cazy.org)5,7.

Among others, the CAZy database lists the reported enzymatic activity of the biochemically characterized CAZy proteins1,7,8. As of 2024, although some CAZy families (e.g., GH22) are associated with only one reported enzymatic activity, most are associated with several activities. For each family, the CAZy database provides an updated list of associated activities in the form of “Enzyme Commission” (EC) numbers (e.g., > 30 EC numbers for GH5). In addition, the recent implementation of CAZac provides detailed information about the substrates and products of each reaction9. Yet, even in polyspecific CAZy families, individual enzymes are usually associated with a single reported function1 and variations in the substrate specificity and activity within CAZy families mirror diverging evolutionary processes and have been associated with subtle changes in the structure of the substrate binding cleft and active site.

The CAZy database also defines CAZy domains and classifies them into families based on conserved 3D-structures and amino acid sequences7. Therefore, Multiple Sequence Alignments (MSA) and Hidden Markov Model (HMM) profiles have been frequently used to identify CAZy domains [e.g1012]. However, while the CAZy database defines CAZy families, no explicit definitions (e.g., HMM-profiles, MSA) are provided and no sequence is directly available on the CAZy database. Before 2012, the CAZy database provided links to other relevant databases useful for domain definition, sequence identification, and characterization (e.g., Pfam, InterPro). However, these links have been systematically removed from the cazy.org website. Conversely, links redirecting the users to deprecated databases are still in place (e.g., PRINTS and HOMSTRAD). In this context, the reduction in the usability of the CAZy database prompted the development of alternative annotation systems for CAZy domain identification relying on publicly accessible information [e.g10,12,13]. For some CAZy families, sequence variations leading to polyspecificity were accounted for and integrated into sorting algorithms, resulting in the creation of subfamilies with somewhat conserved substrate specificity [e.g1417]. For example, per the CAZy database, the subfamily 7 in the GH5 family (i.e., GH5_7) is associated with just two reported activities.

Besides CAZy proteins with a single reported enzymatic activity, some CAZy proteins are associated with multiple activities. Briefly, “Tier 1” multifunctional enzymes are single-domain CAZy proteins that display multiple activities. These proteins generally display high activity toward one substrate and reduced activity on alternative promiscuous substrates18. As proteins are flexible 3D structures, it is expected that most CAZy can accommodate and process structurally related substrates albeit with variable affinity and kinetic efficiency18,19. Alternatively, multidomain enzymes with a single catalytic domain are called “Tier 2”, and those with multiple catalytic domains are “Tier 3” multifunctional enzymes18. These multi-domain proteins result from genetic recombination that shuffles sequences for protein domains20,21. While domain shuffling is essentially random and has the potential to generate many domain combinations12, most characterized multi-domain proteins listed on the CAZy database consist of a single catalytic domain associated with CBM(s) and perform a single function. Proteins with multiple catalytic domains are relatively scarce among CAZy12,20,21. However, the fusion of protein domains involved in the same process and displaying synergies might be under positive selection and few Tier 3 multifunctional CAZy have been identified and characterized [e.g12,21,22]. For most of these multi-domain multi-activity proteins, biochemical evidence linking specific activities to specific domains is generally missing. In consequence, for Tier 3 CAZy proteins, all the described activities are associated with each identified domain and incorporated in the list of activities found in the corresponding CAZy family. This lack of biochemical evidence ends up ballooning the number of activities associated with each CAZy domain and family1,8.

While no annotation system is directly available from the CAZy database, the compiled information on the CAZy database is used extensively to link sequences identified in (meta)-genomes to enzymatic function and to ecosystem performance and health [e.g2325]. However, as the CAZy database was not designed for sequence annotation and offered limited resources for that purpose, linking identified sequences to function frequently relies on a “majority rule”: in absence of biochemical evidence, the function of a newly identified CAZy sequence is assumed to be the dominant activity identified in the corresponding CAZy family or potential subfamily. For example, for the lack of a better solution, GH5 identified in metagenomes are generally assumed to be cellulases [e.g2628] however, many distinct activities are associated with GH5. This potentially misleading majority rule is also implemented in some annotation systems, at the GH family or sub-family level29. Overall, although the “majority rule” makes sense for monospecific CAZy families and families dominated by a single enzymatic activity, no rationale supports its application to mostly polyspecific CAZy families.

Here, we present ez-CAZy, a custom database for linking newly identified sequences to specific enzymatic activities. We retrieved and re-annotated the sequences of biochemically characterized GHs listed on the CAZy database. For each sequence, we identified the precise domain architecture using publicly accessible tools to link sequence variation to the identified enzymatic function. The data was consolidated in a single publicly available FASTA file. Next, we investigated how the enzymatic activity (i.e., EC) and the overall domain architecture of GHs relate to variation in the sequence of the catalytic domains. We hypothesized that more similar sequences would display similar enzymatic activities and that the clumped distribution of enzymatic activities among characterized CAZy would reflect their multi-domain architecture. Finally, we tested if the activity of newly identified and characterized GHs – not listed in our initial dataset – could be predicted by performing local sequence alignment. Here, we predicted that, since the EC activity and the multi-domain architecture are not randomly distributed among GHs, matching unknown sequences to characterized GHs will facilitate the prediction of the enzymatic activity of newly identified sequences. This work, based on publicly accessible data and bioinformatics tools, provides an easy way to (re)analyze available sequences and envision custom annotation of CAZy sequences from newly sequenced samples.

Results

Sequences retrieval and reannotation

We retrieved the sequence and functional information of 7,198 characterized GH proteins listed on CAZy in Spring 2023. The sequences were re-annotated using Pfam HMM-profiles (Supplementary data 1). A total of 14,841 protein domains from 445 distinct domain families were identified. Of the characterized GHs, 3,169 were single-domain proteins, while the remaining 4,029 multi-domain proteins contained 11,672 protein domains. Multi-domain CAZy with two (n = 2,081) or three (n = 1,229) domains made up the majority of multidomain proteins, but more than 170 proteins contained at least 6 identified domains and one GH70 protein (i.e., CDX66820.1) was associated with 19 domains. Our Pfam-based annotation matched most of the annotation from CAZy, except for the recently coined CAZy families for which no HMM-profile nor multiple sequence alignment were publicly available at the time of analysis. These proteins were generally annotated as members of previously defined CAZy families. For example, the b-galactosidase from Bacteroides ovatus ATCC 8483 (ALJ49469.1), although listed as a GH147 on the CAZy database, was annotated as a GH5 (PF00150), suggesting the structural proximity between these families. In addition, some of the unmatched proteins were assigned to domains not known for their involvement in carbohydrate metabolism. For example, the GH120 from Bifidobacterium adolescentis ATCC 15,703 (BAF39080.1) was annotated as a domain of unknown function (DUF1565 - PF07602) and a b-helix domain (PF13229).

The reannotated sequences were compiled in two separate FASTA files containing (i) the whole protein sequence (ez-CAZy_V1a.faa) and (ii) the sequence of the identified catalytic domain only (ez-CAZy_V1b.faa) available in supplementary data.

GH multi-domain architecture vs. enzymatic activity

Following the creation of the databases, we investigated the domain organization from N- to C-terminal across characterized GHs (Supplementary data 1). Several GH families contained only single-domain proteins. However, most of these families were recently created or contained few characterized proteins with some notable exceptions including GH22, GH34, and GH56 with 62, 30, and 29 single domain proteins, respectively. When considering GH families with ≥10 characterized proteins, some families were still dominated by single-domain proteins (e.g., GH8, GH51) but most contained proteins with variable and sometimes complex multi-domain architectures. Some large families consisted of mostly multi-domain proteins (e.g., GH3, GH9, GH13, GH32) (Fig. 1A). For these predominantly multi-domain GH families, the existence of specific CAZy subdomains increased the prevalence of multi-domain proteins. For example, many characterized GH3 (PF00933) contained a specific C-terminal subdomain (PF01915) only found in some GH3. Similarly, many GH9 (PF00759) were associated with a specific CelD_N subdomain (PF02927) at the N-terminal end (Fig. 2). Although these subdomains were only found in association with GH3 and GH9 respectively, they were not systematically identified. Besides family-specific subdomains, proteins from many GH families were associated with myriads of domains. For example, 61, 55, 21, and 13 distinct types of protein domains were associated with characterized GH5, GH13, GH9, and GH3 respectively. Most domains associated with GHs were small non-catalytic domains such as carbohydrate binding modules (CBM), dockerin domains, and domains of unknown function (DUF). Finally, some characterized GH proteins contained multiple catalytic domains sometimes in conjunction with accessory non-catalytic domains. For example, celA (ACM60955.1) from Caldicellulosiruptor bescii DSM 6725 consisted of N-GH9-3×(CBM3)-GH48-C22. Other characterized proteins consisted of a single GH catalytic domain and a non-CAZy catalytic domain. For example, RmNag3 (ACY47819.1) from Rhodothermus marinus DSM4252 was composed of typical “GH3-GH3c” at the N-terminal end associated with a functional ß-lactamase domain, for penicillin degradation, at the C-terminal end30.

Fig. 1.

Fig. 1

(A) Distribution of multi-domain GH architecture among families with ≥50 characterized proteins at the time of analysis. (Proteins are sorted by proportion of single-domain GHs). (B) Distribution of reported enzymatic activity among GH families with ≥50 characterized proteins at the time of analysis. Number of reported activities may not match the number of proteins as some proteins display multiple activities. (Protein families are sorted by the prevalence of their most frequent activity).

Fig. 2.

Fig. 2

Tree of the GH9 catalytic domains (PF00759, in red) highlighting the distribution of multi-domain architecture, from N- to C-terminal. All the displayed proteins exhibit only cellulolytic activity (EC.3.2.1.4) unless otherwise noted.

Most single-domain GH families (e.g., GH22, GH34, GH56) were affiliated with a reduced number of enzymatic activities. For example, all identified GH22 were described as lysozyme (EC.3.2.1.17). Conversely, most predominantly multi-domain GH families (e.g., GH3, GH13, GH5) were affiliated with many enzymatic activities (Fig. 1B) suggesting that the addition of accessory domains affected the activity of the catalytic domains. However, in some GH families, variable multi-domain architecture was not associated with variable enzymatic activity. Indeed, although characterized GH9 displayed many variable architectures, most (86%) acted as cellulase (i.e., EC.3.2.1.4) and eventually displayed additional activities (Fig. 2). Conversely, the mostly single-domain GH8 proteins displayed variable activities (e.g., EC.3.2.1.4, EC.3.2.1.6, EC.3.2.1.8).

To investigate how the discrete evolution of the catalytic domains and the addition of accessory domains impacted the reported enzymatic activity, we created trees using the sequence of the identified catalytic domains and investigated the distribution of the protein multi-domain architecture and the reported EC activity/ies. For example, when considering biochemically characterized GH9, our initial dataset included 151 proteins accounting for 10 distinct enzymatic activities, with a total of 374 domains from 21 domain families (Fig. 2). The distribution of non-GH9 domains relative to the sequence of the GH9 domain was strongly clustered. Specifically, apart from a few deeply branched GH9 sequences, all the proteins with an N-terminal “CelD_N” domain (PF02927) belonged in a single cluster. Within this cluster, some smaller clusters were composed of proteins with additional domains located before or after the core CelD_N-GH9 domains. Overall, proteins with complex multidomain architecture formed smaller homogenous clusters. The only notable exceptions to this clustered distribution are the presence of CBM_3 (PF00942), and the dockerin domain (PF00404) which were found in multiple clusters.

Regarding the reported enzymatic activity (EC) of biochemically characterized GH9 proteins, most acted as cellulase (EC.3.2.1.4), while enzymes with alternative activities were often clustered and displayed specific multi-domain architecture (Fig. 2). However, within these clusters, not all enzymes were reported to share the same activity.

The clustered distribution of the multidomain architecture of biochemically characterized GH9 along with the distribution of the reported enzymatic activities was further investigated using an ANOSIM test. Both the distribution of multidomain architecture and the reported enzymatic activity were positive, thus indicating a significantly clustered distribution (p < 0.01).

We then performed the same analysis for all the GH families associated with at least 10 biochemically characterized proteins (Fig. 3A). Across GH families with various accessory domains, the domain architecture was generally clustered (R-ANOSIMArchi>0, p < 0.01, Fig. 3A). In the few families where the domain architecture was not clustered, the distribution was generally random (neither significantly clustered nor scattered, ANOSIMArchi~0, p > 0.05). In addition, most of the GH families where the distribution of multi-domain architectures was randomly distributed were families largely dominated by single-domain proteins (e.g., GH8) or GH families with few characterized proteins (e.g., n = 17 in GH130) (Fig. 3A).

Fig. 3.

Fig. 3

(A) ANOSIM test investigating the clustering of reported EC activities and multi-domain architectures across GH families with > 10 biochemically characterized enzymes. (B) Frequency of predicted enzymatic activity of newly listed GHs matching the reported biochemically characterized activity.

The distribution of the reported enzymatic activities across biochemically characterized GH families, was significantly clustered in most protein families (R-ANOSIMEC>0, p < 0.01, Fig. 3A). However, in a few GH families with many characterized proteins (e.g., GH12, GH53, GH32), the reported enzymatic activity was essentially random. Interestingly, for most of these families, the architecture was also randomly distributed.

Overall, this suggested that the sequence of the catalytic domain alone can inform the precise enzymatic activity (EC) for most GH families.

The activity of new GH sequences

Next, we tested the possibility of using our custom database to infer the specific function of newly characterized GH proteins. We retrieved 513 sequences that were characterized after the compilation of our initial database and for which a “reported activity” was available on the CAZy database as of November 2024. The sequences were distributed among 67 of the 175 listed GH families. Focusing on GH families with ≥10 proteins in our initial database, we matched the new sequences (queries) with characterized catalytic domains (targets) using local alignment with an adjusted e-value cutoff of 10−3 to account for the small size of our database. We then compared the activity of the closest hit/target for each new query sequence with the reported enzymatic activity as reported in the CAZy database (Fig. 3B).

The activity of 411 newly characterized proteins matched some activity of their best hit. However, because of variation in how GH proteins are characterized in various labs (e.g., what substrates are available and/or tested) and how the activity is reported, the activity of some multidomain queries only partially matched the activity of characterized sequences in our custom database.

The number and frequency of matches between the predicted activity, derived from the best hit, and the reported activity varied across GH families. However, the activity of all newly characterized GH proteins for 27 families exactly matched the activity of their closest hit in our database. Although this was expected for monofunctional GH families (e.g., GH22, GH37, GH77), this was also observed in several GH families associated with multiple activities (e.g., GH20, GH27, GH44, and GH62). Moreover, at least 75% of the predicted activity matched the reported activity in 13 additional GH families. For example, the reported activity of 88.2% and 85.7% of the newly characterized GH5 and GH43 respectively, matched the activity of their closest hit (Fig. 3B).

In a few families, less than 75% of newly characterized enzymes matched the reported activity of their closest relative in our custom database. In some families, this reflected the low number of characterized sequences in our database or the low number of newly characterized proteins. For example, the reported activity of the unique, newly characterized GH4, GH33, GH49, and GH94 did not match the reported activity of their closest relative. But less than 75% of newly characterized proteins from some populated GH families including GH1, GH3, GH29, and GH32 matched the activity of their closest relative in our database. This discrepancy likely mirrors diverging evolutionary processes leading to distinct activities among related sequences within polyspecific GH families (e.g., GH1, GH3) and the uneven distribution of biochemically characterized proteins in some families (e.g., GH49).

ez-CAZy-V2

We combined the reannotated sequences from ez-CAZy-V1 and the additional sequences of recently characterized GH proteins analyzed above in a single publicly accessible file (i.e., ez-CAZy-V2). In addition, ez_CAZy-V2 also contains the Pfam based annotation with the domain architecture and the listed enzymatic activity of all the sequences from all the biochemically characterized CAZy (i.e., GT, CE, PL, AA) listed on CAZy database as of October 2024.

Discussion

The identification of CAZy proteins in sequenced datasets is essential to understand the diversity and physiology of individual lineages and the functioning of complex microbial communities. In addition, as there is an ever-growing need for effective catalysts to support biotechnological application, identifying new CAZy proteins is needed. In this context, the CAZy database (cazy.org) defining CAZy families, and compiling lists of biochemically characterized CAZy proteins has been an essential resource for biochemists, microbiologists, and microbial ecologists. In 2008, Davies and Sinnott noted that “... sequence annotations are one of the great banes of modern molecular biology, and the CAZy classifications could, and should, inform much more insightful annotations across the literature”. However, as of 2024, the continued lack of definitions (e.g., HMM-profiles), clearly annotated sequences, guidelines, and the many activities associated with most CAZy families complicate the investigation of enzymatic processing of carbohydrates in sequenced genomes and metagenomes. Despite its monopoly, the CAZy database offers no easy way to perform such comparison resulting in the creation of alternative annotation pipelines based on publicly accessible tools (e.g., dbCAN, GeneHunt)11,12,28,29. These pipelines, sometimes integrated in larger annotation infrastructures (e.g., KBase)31, have successfully identified CAZy sequence in assembled metagenomes32 and sequenced genomes33. However, these pipelines, when predicting the activity of newly identified GHs, continue to rely on the majority rule (e.g29).

Here, we compiled ez-CAZy to infer the function of GHs identified in sequenced genomes and metagenomes. The proposed approach relies on available reported enzymatic activity of characterized GHs and on the idea that closely related sequences share similar activities. However, the lack of standardized GH characterization and the binary nature of the reported activity can produce inconsistencies in the available data. Specifically, when considering new GHs, a priori knowledge of an enzyme and/or its sequence can limit the range of tested substrates. For example, according to the majority rule, knowing that most characterized GH5 are described as cellulase, newly identified GH5 would only be tested for their cellulolytic activity especially since testing other activities can be technically difficult, costly, and time-consuming.

In conclusion, the development of the ez-CAZy database addresses significant gaps in the annotation and functional prediction of glycoside hydrolases (GH). The clustered distribution of enzymatic activities and domain architectures across GH families highlights the evolutionary and structural underpinnings of functional diversity in these enzymes. By systematically reannotating sequences and leveraging publicly accessible tools, this resource enhances the ability to link sequence variations to specific enzymatic activities. While providing a streamlined framework for functional prediction of GHs, ez-CAZy bridges the limitations of the CAZy database, offering a publicly accessible, updated resource for researchers. In addition, the data for other carbohydrate-active enzymes was compiled and made publicly accessible. In the future, this approach will contribute not only to advancing microbial and enzymatic research but also to practical applications, such as identifying candidate enzymes for specific biotechnological applications.

Moving forward, the biochemical characterization of new GHs from families associated with few characterized enzymes or displaying new multi-domain architecture will help further elucidate the interplay between variation in the sequence of the catalytic domain, the multidomain architecture, and the detected enzymatic activity. In addition, the adoption of standardized characterization methods and broader testing of enzymatic activities could further improve the accuracy and utility of resources like ez-CAZy, driving innovation in both academia and industry.

Methods

Sequence collection and reannotation

Protein sequences from biochemically characterized GHs, listed on the CAZy database (http://www.cazy.org) were retrieved from the NIH website using eFETCH34. The sequences were reannotated using HMMscan35 and the entire set of HMM-profiles from PFam-A (V.37)4,36. Finally the annotation from the CAZy database, from the NIH, and our custom annotation were merged with the sequence and metadata were formatted as follows: CAZy identification (e.g., GH5), sequence accession number from the NIH (e.g., AAP56348.1), the domain architecture - from the N- to the C-terminal end (e.g., N-PF00150-PF00942-C), the position of the identified domains - from the N- to C-terminal (e.g., 47–379,476–550), the reported enzymatic activity per the CAZy database (e.g., 3.2.1.4), and the taxonomic origin if known (e.g., Thermobifida fusca TM51).

Sequences comparison

After retrieving and reannotating the complete sequences of biochemically characterized GH, we trimmed the sequences to only include the catalytic domains. Then, the sequences were aligned using Clustal Omega (V.1.2.3)37. Next, we used Phylip (V.3.695)38 to compute the evolutionary distance (JTT) between aligned sequences and created trees (Neighboor joining) which were visualized with iTOL39. Finally, ANOSIM tests with 9,999 permutations implemented in R Vegan package (V.2.6-8)40 were performed to investigate the distribution (i.e., clustered, random, scattered) of the reported enzymatic activities (EC) and the complete domain architecture from N- to C-terminal relative to the sequence of the aligned catalytic domains.

All data were analyzed using R statistical language (V.4.4.1) and visualized using ggplot2 (V.3.5.1).

New sequences characterization

In October 2024, we redownloaded the entire set of biochemically characterized proteins listed on the CAZy database and discarded the sequences already included in our initial dataset. These new sequences were treated as unknown queries for local alignment with our established dataset (i.e., ez-CAZy_V1b.faa), although these sequences were already associated with specific enzymatic activity (i.e., EC #). A similar procedure can be used to analyze new and unknown sequences (i.e., queries) in the terminal of a computer running a Linux OS. First, the database consisting of a single FASTA formatted file is retrieved from https://figshare.com/articles/dataset/Supplementary_data2_All_characterized_GH_catalytic_domains_retrieved_Spring_2023_-_ez-CAZy_V1b_faa/28017464?file=51132857. Next, using the NIH BLAST package (https://blast.ncbi.nlm.nih.gov/doc/blast-help/downloadblastdata.html), the initial FASTA file (i.e., ez-CAZy_V1b_faa) is converted to a custom database using the following command “makeblastdb ez-CAZy_V1b.faa -dbtype prot”. Finally, the protein queries in a single FASTA formatted file (e.g., NewQueries.txt) are aligned locally by entering “blastp -query NewQueries.txt -db ez-CAZy_V1b.faa -out Result.txt -outfmt 6”. This last command saves the detailed statistics for each of the queries and their matching hits (from best to worst) in Result.txt. We recommend unfamiliar users read the BLAST documentation in order to become familiar with the many options available within the program (e.g., e-value & score-based thresholding)41.

Electronic supplementary material

Below is the link to the electronic supplementary material.

Supplementary Material 1 (19.5KB, docx)

Acknowledgements

We thank all investigators who contributed to the characterization of the listed sequences and proteins. We thank Dr. J. Munoz and M. Akhtar for comment on earlier versions of the manuscript.The material generated and analyzed in this study is based on work supported by the National Science Foundation under Award Numbers #2138880 (RB).

Author contributions

D.S.E. & R.B. conceived the work, retrieved, developed, and analyzed the data and wrote the manuscript. N.G.G. contributed to the data retrieval and processing.

Data availability

The datasets generated and analyzed in the current study are available in the figshare repository https://figshare.com/projects/ez-CAZy/231242 and from the corresponding author on reasonable request.

Declarations

Competing interests

The authors declare no competing interests.

Footnotes

Publisher’s note

Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.

References

  • 1.Drula, E. et al. The carbohydrate-active enzyme database: functions and literature. Nucleic Acids Res.50, D571–D577 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 2.Ahmad, S. et al. Engineering processive cellulase of Clostridium thermocellum to divulge the role of the carbohydrate-binding module. Biotechnol. Appl. Biochem.10.1002/bab.2352 (2022). [DOI] [PubMed] [Google Scholar]
  • 3.Hervé, C. et al. Carbohydrate-binding modules promote the enzymatic Deconstruction of intact plant cell walls by targeting and proximity effects. Proc. Natl. Acad. Sci. U S A. 107, 15293–15298 (2010). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 4.Finn, R. D. et al. Pfam: the protein families database. Nucleic Acids Res.42, D222–D230 (2014). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 5.Henrissat, B. A classification of Glycosyl hydrolases based on amino acid sequence similarities. Biochem. J.280 (2), 309–316 (1991). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6.Dodd, D., Moon, Y. H., Swaminathan, K., Mackie, R. I. & Cann, I. K. O. Transcriptomic analyses of Xylan degradation by Prevotella bryantii and insights into energy acquisition by Xylanolytic bacteroidetes. J. Biol. Chem.285, 30261–30273 (2010). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 7.Davies, G. J. & Sinnott, M. L. Sorting the diverse: the sequencebased classifications of carbohydrateactive enzymes. Biochemist30, 26–32 (2008). [Google Scholar]
  • 8.Lombard, V., Ramulu, G., Drula, H., Coutinho, E., Henrissat, B. & P. M. & The carbohydrate-active enzymes database (CAZy) in 2013. Nucleic Acids Res.42, D490–D495 (2014). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 9.Lombard, V., Henrissat, B. & Garron, M. L. CAZac: an activity descriptor for carbohydrate-active enzymes. Nucleic Acids Res.gkae104510.1093/nar/gkae1045 (2024). [DOI] [PMC free article] [PubMed]
  • 10.Fraser, A. S. C. et al. SACCHARIS v2: streamlining prediction of Carbohydrate-Active enzyme specificities within large datasets. Methods Mol. Biol. Clifton NJ. 2836, 299–330 (2024). [DOI] [PubMed] [Google Scholar]
  • 11.Huang, L. et al. dbCAN-seq: a database of carbohydrate-active enzyme (CAZyme) sequence and annotation. Nucleic Acids Res.46, D516–D521 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12.Nguyen, S. et al. GeneHunt for rapid domain-specific annotation of glycoside hydrolases. Sci. Rep.9, 10137 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 13.Ausland, C. et al. dbCAN-PUL: a database of experimentally characterized cazyme gene clusters and their substrates. Nucleic Acids Res.49, D523–D528 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 14.Arumapperuma, T. et al. A subfamily classification to choreograph the diverse activities within glycoside hydrolase family 31. J. Biol. Chem.299, 103038 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15.Hornung, B. V. H. & Terrapon, N. An objective criterion to evaluate sequence-similarity networks helps in dividing the protein family sequence space. PLoS Comput. Biol.19, e1010881 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16.Mareček, F., Terrapon, N. & Janeček, Š. Two newly established and mutually related subfamilies GH13_48 and GH13_49 of the α-amylase family GH13. Appl. Microbiol. Biotechnol.108, 415 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 17.Viborg, A. H. et al. A subfamily roadmap of the evolutionarily diverse glycoside hydrolase family 16 (GH16). J. Biol. Chem.294, 15973–15986 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 18.Glasgow, E., Vander Meulen, K., Kuch, N. & Fox, B. G. Multifunctional cellulases are potent, versatile tools for a renewable bioeconomy. Curr. Opin. Biotechnol.67, 141–148 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19.Schmitz, E., Leontakianakou, S., Adlercreutz, P., Nordberg Karlsson, E. & Linares-Pastén, J. A. Novel function of CtXyn5A from Acetivibrio thermocellus: dual arabinoxylanase and feruloyl esterase activity in the same active site. Chembiochem Eur. J. Chem. Biol.24, e202200667 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 20.Berlemont, R., Fuller, D. A. & Sudarshan, A. Modularity of Cellulases, Xylanases, and Other Glycosyl Hydrolases Relevant for Biomass Degradation. in Handbook of Biorefinery Research and Technology (ed. Bisaria, V.) 1–35Springer Netherlands, Dordrecht, (2022). 10.1007/978-94-007-6724-9_24-1
  • 21.de García-Paz, F. Multidomain chimeric enzymes as a promising alternative for biocatalysts improvement: a minireview. Mol. Biol. Rep.51, 410 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22.Brunecky, R. et al. Revealing nature’s cellulase diversity: the digestion mechanism of Caldicellulosiruptor bescii CelA. Science342, 1513–1516 (2013). [DOI] [PubMed] [Google Scholar]
  • 23.Cabral, L. et al. Gut Microbiome of the largest living rodent harbors unprecedented enzymatic systems to degrade plant polysaccharides. Nat. Commun.13, 629 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 24.Cantarel, B. L., Lombard, V. & Henrissat, B. Complex carbohydrate utilization by the healthy human Microbiome. PloS One. 7, e28742 (2012). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 25.Turnbaugh, P. J. et al. A core gut Microbiome in obese and lean twins. Nature457, 480–484 (2009). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 26.Berlemont, R. et al. Cellulolytic potential under environmental changes in microbial communities from grassland litter. Front. Microbiol.5, 639 (2014). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 27.Klimek, D., Herold, M. & Calusinska, M. Comparative genomic analysis of planctomycetota potential for polysaccharide degradation identifies biotechnologically relevant microbes. BMC Genom.25, 523 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 28.Shaffer, M. et al. DRAM for distilling microbial metabolism to automate the curation of Microbiome function. Nucleic Acids Res.48, 8883–8900 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 29.Zheng, J. et al. dbCAN3: automated carbohydrate-active enzyme and substrate annotation. Nucleic Acids Res.51, W115–W121 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 30.Ara, K. Z. G. et al. Characterization and diversity of the complete set of GH family 3 enzymes from Rhodothermus marinus DSM 4253. Sci. Rep.10, 1329 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 31.Arkin, A. P. et al. KBase: the united States department of energy systems biology knowledgebase. Nat. Biotechnol.36, 566–569 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 32.Zheng, J. et al. dbCAN-seq update: cazyme gene clusters and substrates in microbiomes. Nucleic Acids Res.51, D557–D563 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 33.Berlemont, R. The supragenic organization of glycoside hydrolase encoding genes reveals distinct strategies for carbohydrate utilization in bacteria. Front. Microbiol.14, 1179206 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 34.Kans, J. Entrez Direct: E-utilities on the Unix Command Line. (2013).
  • 35.Eddy, S. R., Accelerated Profile, H. M. M. & Searches PLoS Comput. Biol.7, e1002195 (2011). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 36.Potter, S. C. et al. InterPro in 2019: improving coverage, classification and access to protein sequence annotations. Nucleic Acids Res.47, D351–D360 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 37.Sievers, F. & Higgins, D. G. Clustal Omega for making accurate alignments of many protein sequences. Protein Sci. Publ Protein Soc.27, 135–145 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 38.Retief, J. D. Phylogenetic analysis using PHYLIP. Methods Mol. Biol. Clifton NJ. 132, 243–258 (2000). [DOI] [PubMed] [Google Scholar]
  • 39.Letunic, I. & Bork, P. Interactive tree of life (iTOL) v5: an online tool for phylogenetic tree display and annotation. Nucleic Acids Res.49, W293–W296 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 40.Oksanen, J. et al. Vegan: community ecology package. R Package Version 1 R Package Version 2 0–4. 10.4135/9781412971874.n145 (2012). [Google Scholar]
  • 41.Altschul, S. F. et al. Gapped BLAST and PSI-BLAST: a new generation of protein database search programs. Nucleic Acids Res.25, 3389–3402 (1997). [DOI] [PMC free article] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Supplementary Material 1 (19.5KB, docx)

Data Availability Statement

The datasets generated and analyzed in the current study are available in the figshare repository https://figshare.com/projects/ez-CAZy/231242 and from the corresponding author on reasonable request.


Articles from Scientific Reports are provided here courtesy of Nature Publishing Group

RESOURCES