Abstract
Catalytic bioparts are fundamental to the design, construction and optimization of biological systems for specific metabolic pathways. However, the functional characterization information of these bioparts is frequently dispersed across multiple databases and literature sources, posing significant challenges to the effective design and optimization of specific chassis or cell factories. We developed the Registry and Database of Bioparts for Synthetic Biology (RDBSB), a comprehensive resource encompassing 83 193 curated catalytic bioparts with experimental evidences. RDBSB offers their detailed qualitative and quantitative catalytic information, including critical parameters such as activities, substrates, optimal pH and temperature, and chassis specificity. The platform features an interactive search engine, visualization tools and analysis utilities such as biopart finder, structure prediction and pathway design tools. Additionally, RDBSB promotes community engagement through a catalytic bioparts submission system to facilitate rapid data sharing and utilization. To date, RDBSB has supported the contribution of >1000 catalytic bioparts. We anticipate that the database will significantly enhance the resources available for pathway design in synthetic biology and serve essential tools for researchers. RDBSB is freely available at https://www.biosino.org/rdbsb/.
Graphical Abstract
Graphical Abstract.
Introduction
A primary goal of synthetic biology is the construction of artificial biological systems with novel functions archived through bottom-up forward engineering principles (1). Bioparts, as the fundamental building blocks, are indispensable and serve as the cornerstone of synthetic biology research and application (2,3). Catalytic bioparts refer to not only enzymes or other biomolecules that can catalyze biochemical reactions, but also their performance data and nuclear acid sequences in specific chassis (4,5). The selection and utilization of catalytic bioparts directly influence the efficiency and specificity of biosynthetic systems (6). By precise designing and assembling various catalytic bioparts, novel biological systems can be created for the production of valuable chemicals, pharmaceuticals, biofuels and other products (7–9).
Constructing biosynthetic systems needs to assemble target biosynthetic pathways within specific chassis in a ‘plug-and-play’ manner based on standardized catalytic bioparts (10,11). Thus, a reliable database of experimentally validated catalytic bioparts is essential for providing the resources required to design and optimize these biosynthetic systems. The Registry of Standard Biological Parts includes over 300 catalytic bioparts developed to address specific challenges in iGEM competition (12). While databases such as Swiss-Prot, MetaCyc, KEGG, BRENDA, CAZy and GTDB focus on curated protein (13), biosynthetic pathways (14,15) or enzyme (16–18) in non-engineered organism, they lack crucial information on the performance of catalytic biopart within chassis, which is vital for synthetic biology applications.
To address the gap, we developed the Registry and Database of Bioparts for Synthetic Biology (RDBSB), which contained 83 193 experimentally validated catalytic bioparts. We have also systematically reviewed and compiled comprehensive information for the application of catalytic bioparts, including optimal pH and temperature, as well as chassis. RDBSB provides an interactive data search engine and visualization interface for the convenience of users. Additionally, it offers a suite of efficient tools, including catalytic biopart finders such as BiopartFinder and MapView, structure prediction tools such as AlphaFold (19) and PVQD (20) and pathway design tool PathFinder, all of which provide robust support for discovering and utilizing catalytic biopart data. Furthermore, we have designed an effective submission system for sharing of catalytic biopart information. RDBSB is freely available at https://www.biosino.org/rdbsb/.
Materials and methods
The methodology for developing RDBSB is illustrated in Figure 1.
Figure 1.
The overview of RDBSB workflow. This workflow contains three steps, including data source and integration, dataset construction and online analysis.
Data source and integration
MetaCyc Protein, BRENDA, Swiss-Prot, KEGG GENES and NCBI GenBank served as the primary resources for integrating catalytic bioparts, providing crucial information on proteins, coding sequences (CDSs), reactions and functions (13–15,17,21). Reactions were sourced from MetaCyc Reactions, MetaCyc Enzrxns, KEGG REACTION, RHEA and BKMS (22,23), while substrate and product annotations were derived from MetaCyc Compounds, KEGG COMPOUND, PubChem and ChEBI (24,25). KEGG Enzyme, ENZYME, Pfam, MetaCyc pathways, KEGG PATHWAY, AlphaFoldDB and NCBI Taxonomy were used to annotate the enzyme, domain, pathway, structure and organism associated with each biopart (26–29) (Figure 1). Following data integration and deduplication, 390 708 catalytic bioparts were extracted.
The screening and curation of experimental validated reaction
All bioparts were screened for experimental validation based on reaction information. First, bioparts lacking reaction information in any source database were excluded, leaving 323 328 catalytic bioparts associated with reactions for further analysis (13–15,17). Second, bioparts marked with the ‘Evidence at protein level’ designation by Swiss-Prot were immediately classified as having experimental validation. Third, the remaining bioparts were evaluated for experimental evidence based on the presence of literature-supported reaction information in any of the source databases. Additionally, the reactions catalyzed by cytochrome P450 (P450) and glycosyltransferase GT1 (GT1) family enzymes were directly curated from the literature. In total, 83 193 of these bioparts were confirmed as experimentally validated.
The curation of experimental condition
We employ a ‘two-curator, one-reviewer’ method for manual review. In this process, two curators independently conduct assessments. If their conclusions align, the result is immediately confirmed. However, if discrepancies arise, a reviewer is brought in to facilitate discussions between the curators, and the final outcome is based on the consensus reached. For chassis review, enzyme method descriptions from BRENDA were curated using this approach, leading to the standardization of 2630 chassis description. Similarly, to determine the optimum temperature and pH for catalytic bioparts, literatures from Swiss-Prot were curated and structured into 2636 temperature records for 3049 bioparts and 4325 pH records for 8637 bioparts.
Additionally, RDBSB incorporated 10 535 structured optimum pH records and 9452 optimum temperature records from BRENDA, along with 590 optimum pH records and 319 optimum temperature records from MetaCyc. During the curation and incorporation process, pH and temperature values were standardized to ranges of 0–14 for pH and 0–125°C for temperature. As a result, 27 789 bioparts were associated with pH or temperature or chassis information.
The curation of P450 and GT1 family enzymes
Given the critical roles of P450 in oxygenation and GT1 in glycosylation of natural products, our curation efforts were specifically focused on these bioparts (21,22). To identify potential P450 and GT1 bioparts, we constructed a hidden Markov model library based on sequence profiles using HMMER (version 3.1b2) (30). The profile of GT1 family was obtained from dbCAN (version 8) (31), and the profile of P450 sourced from Pfam (version 32) (27). We then mined potential P450 and GT1 protein sequences from Swiss-Prot and the NCBI Non-Redundant Protein Sequence Database (NR, January 2020 download) (32) using hmmscan (pipe-0.0.1 r2, with parameters: e-value 1e−5, coverage 0.4) (33). For the proteins identified in Swiss-Prot, we reviewed associated biochemical reactions, kinetic parameters and evidence levels. To further refine the dataset, we filtered the literature associated with proteins from NR using keywords such as SDS–PAGE, HPLC, LC–MS, gas chromatography, GC–MS and TLC, ultimately selecting 276 P450 papers and 224 GT1 papers for detailed curation.
In curating P450 and GT1 bioparts, we aimed for comprehensive coverage by adhering to previously established curation methods. To address issues related to incomplete or non-standardized reactions in the literature, we supplemented the reactions with established chemical principles, reconstructed structural formulas for substrates and products, and ensured accurate naming and database ID assignment. Additionally, substrates and products were standardized via PubChem and ChEBI. For entries that could not be matched, structures were generated in SMILES format using ChemDraw (34), ensuring consistency and reliability across our curated data.
Database implementation
The RDBSB dataset can be visualized through web browsers and was developed using the Spring Boot framework (https://spring.io/projects/spring-boot), with core JavaScript libraries including jQuery (https://jquery.com/) and Echarts (https://echarts.apache.org/). Elasticsearch (https://www.elastic.co) is utilized to optimize data retrieval performance. Additionally, we used NCBI BLAST (v.2.13.0+) (35) for sequence alignment, AlphaFold2 (v2.0.1) (19) for tertiary structure prediction and PVQD (20) for multidimensional conformational prediction of catalytic biopart. The 3D structures predicted by AlphaFold and PVQD were displayed by Mol* Viewer (36).
PathFinder development
PathFinder is a tool we developed to facilitate route design from substrate to product via possible pathways. We construct a directed graph G
, where the vertex set
represents individual substrates, reaction and individual products, while the edge set
represents the chemical reactions or combination relationships between these vertices. If a reaction
transforms substrate
into product
, the edge
represents this transformation. To accelerate query speed, we excluded common compounds such as H2O, ATP, cofactor NADPH and coenzyme CoA from the network during graph construction. After building the graph for the RDBSB dataset, we use the allShortestPaths algorithm in Neo4j for path searching (https://neo4j.com/docs/cypher-manual/current/patterns/reference/), which efficiently traverses the shortest path in a directed graph with weighted edges.
Results
Overview of RDBSB
The integrity of biopart information is classified into four levels:
Level 1 refers to catalytic bioparts that contain protein sequences or CDSs.
Level 2 includes bioparts from Level 1 that have associated reactions.
Level 3 comprises bioparts from Level 2 with experimentally validated reactions.
Level 4 consists of bioparts from Level 3 that also have optimum pH, optimum temperature and chassis information.
In total, 390 708 catalytic bioparts were integrated from various database sources, including 83 193 that have been experimentally validated, which far exceed the coverage of other enzyme databases such as BRENDA in terms of experimentally validated catalytic bioparts (17). Of these, 3200 experimentally validated catalytic bioparts include curated data on optimum temperature, optimum pH and chassis (Figure 1).
The top three enzyme categories of bioparts are transferases, hydrolases and oxidoreductases (Figure 2A). The optimum temperature and pH for these bioparts are predominantly within the ranges of 20–40°C and pH 6–9, respectively (Figure 2B), which align with the optimal conditions for commonly used laboratory chassis, such as Saccharomyces cerevisiae, Escherichia coli, Bacillus subtilis, Corynebacterium glutamicum and Nicotiana tabacum. Most source organisms of these bioparts belong to the Bacteria, Metazoa and Viridiplantae groups (Figure 2C).
Figure 2.
Data overview of RDBSB. (A) Distribution of catalytic bioparts across enzyme categories classified according to ENZYME database. Level 1: catalytic bioparts containing protein sequences or CDSs; Level 2: bioparts with reactions identified in Level 1; Level 3: bioparts with experimentally validated reactions from Level 2; and Level 4: bioparts from Level 3 with optimum pH, optimum temperature and chassis information. (B) Distribution of experimentally validated bioparts across enzyme category, chassis, optimum pH and temperature. (C) Distribution of source organisms at varying levels of information integrity categorized according to NCBI Taxonomy.
Web user interface
RDBSB offers a user-friendly interface with versatile tools to visualize and retrieve catalytic bioparts for cell factory construction. Key features include (i) a quick navigation tool for locating bioparts with varying levels of information integrity; (ii) a keyword search function for retrieving bioparts by names, identifiers of public resources, compounds, chassis and pathways; (iii) filtering options to combine different experimental conditions to find suitable biopart, such as chassis, optimum pH and temperature and integrity level; and (iv) BiopartFinder, which enables amino acid sequence searches to identify relevant bioparts and potential alternatives or reaction conditions for pathway design in chassis.
On the detailed page for each biopart, users can access comprehensive information on catalytic functions, supported by literature references to ensure accuracy and reliability. The page also displays alternative bioparts involved in the same reaction, aiding in cell factory construction and optimization. The MapView tool allows users to interact and visually explore relationships between bioparts, chassis, organisms, compounds and literature (Figure 3A). At the bottom of the page, users can access amino acid and CDSs for laboratory use, along with structure prediction tools such as AlphaFold and PVQD. AlphaFold is used to predict precise structures for catalytic bioparts (19), while PVQD offers predictions of various conformations, providing structural insights for designing more efficient and specific bioparts (20) (Figure 3B).
Figure 3.
Online analysis served by RDBSB. (A) MapView example of GmOMT2. The MapView provides on related organism, chassis, reference and compound information of a specific biopart. Hovering over the biopart name displays brief details. (B) User interface for 3D conformation and structure analysis of GmOMT2 by PVQD and AlphaFold. Two conformations of GmOMT2 generated by PVQD are displayed in green and the structure of GmOMT2 predicted by AlphaFold is shown in blue. (C) PathFinder application showcasing examples of artificial biosynthetic pathway: one for drug icaritin pathways in yeast, and another for insect sex pheromones pathways in yeast.
Application of RDBSB for natural products synthetic biology research
The construction of biosynthetic pathways to produce value-added compounds is a primary objective in metabolic engineering and synthetic biology (4,37). With over 80 000 experimentally validated catalytic bioparts available, it is now possible to identify suitable biosynthetic pathways for desired natural products or even design artificial pathways to synthesize target compounds with previously unknown pathways. We utilize PathFinder for pathway design, as demonstrated by two examples (Figure 3C).
Biosynthesis of icaritin
Icaritin is a prenylflavonoid present in the Chinese herbal medicinal plants Epimedium spp., and is currently used in the treatment of advanced hepatocellular carcinoma (38). Structurally, icaritin was synthetized from kaempferol via two steps, the prenylation at C8 and the methylation at C4′-OH of kaempferol. Since both prenyltransferase and methyltransferase are superfamily enzymes and their function and sequence relationship are not clear, it is very challenging to screen suitable prenyltransferase and methyltransferase with desired function from certain plant to fulfill this pathway. We then test PathFinder for the designing of this pathway, when kaempferol and icaritin were input as the substrate and product, respectively, PathFinder designed a biosynthetic pathway for icaritin. Kaempferol was first converted to 8-prenylkaempferol by EsPT2 and then 8-prenylkaempferol was converted to icaritin by GmOMT2. However, directly reconstructing this pathway in yeast cytoplasm led to no icaritin production. As deposited in RDBSB, the cytoplasmic pH of yeast is typically ∼5.5–7, while the biopart GmOMT2 lost function at pH < 6.5; hence, the incompatibility between the suitable pH of biopart and the cytosol pH of yeast chassis led to the failure of de novo icaritin biosynthesis. Complete biosynthesis of icaritin has been successfully achieved by reconstructing this pathway through relocating GmOMT2 into the mitochondria with relatively higher pH (∼7.5) of yeast (39).
Biosynthesis of cis-11-hexadecenal
Another example for pathway design is the biosynthesis of cis-11-hexadecenal (Z11-16:Ald), a key component of sex pheromones in several notorious agricultural pests. Z11-16:Ald was taken as a biosynthesis target to broaden field application at low cost. However, the natural biosynthetic pathway of Z11-16:Ald in Helicoverpa armigera remains largely unknown. A proposed biosynthetic pathway was from palmitoyl-CoA through three-step conversion: palmitoyl-CoA was first converted to (11Z)-hexadecenoyl-CoA by fatty acyl desaturase, and then (Z)-11-hexadecenol was generated by fatty acyl reductase, which was then converted to Z11-16:Ald by alcohol oxidase, while the fatty acyl desaturase and alcohol oxidase remain unknown. When palmitoyl-CoA and Z11-16:Ald were input as the substrate and product, respectively, PathFinder proposed four possible pathways: palmitoyl-CoA was first converted to (11Z)-hexadecenoyl-CoA by desaturase from four different sources, and then thioesterase releases the CoA of (11Z)-hexadecenoyl-CoA to generate the corresponding hexadecanoate product, which was subsequently converted to Z11-16:Ald by carboxylic acid reductase. Notably, all these four pathways were ‘unnatural’; the route from palmitoyl-CoA to Z11-16:Ald was different from the proposed biosynthetic pathway in H. armigera. Results indicated that PathFinder could design artificial pathways to synthesize target compounds with unknown pathways. All the designed pathways were successfully reconstructed in budding yeast to construct cell factories; two of the designed pathways (pathways 1 and 2) resulted in the successful production of the target Z11-16:Ald, with pathway 1 giving the highest production level (40).
Application of RDBSB for the deposition of catalytic bioparts
RDBSB supports the online submission of catalytic bioparts, enabling researchers to submit according to the provided guidelines. Upon review, RDBSB assigns accession numbers to the bioparts. Additionally, RDBSB facilitates the deposition and sharing of related plasmids and strains, which can be made available to the research community upon request. Since its launch in 2019, RDBSB has supported the publication of over 200 catalytic bioparts in high-quality journals and has distributed numerous bioparts to the synthetic biology community (39–45). Through this system, manually curated 142 P450 bioparts and 120 GT1 bioparts in this study have been submitted to RDBSB.
Discussion
RDBSB is a comprehensive platform designed for the collection, storage and sharing of detailed qualitative and quantitative data on catalytic bioparts. It also offers practical tools for applying catalytic bioparts in natural product synthesis, contributing to the advancement of synthetic biology. To the best of our knowledge, we have aggregated as many catalytic bioparts as possible from public resources, with over 80 000 supported by experimental evidence through manual curation and literature mining. RDBSB emphasizes the collection and curation of experimental conditions, including optimum pH and temperature, and compatible chassis. Additionally, our system supports and encourages the submission of new catalytic bioparts, continually enriching the repository with experimental validated data.
To date, our database has received >500 000 visits. Its impact is becoming increasingly evident, with over 200 catalytic bioparts published in well-known journals such as Green Chemistry and Science Bulletin. We also provide tools such as PathFinder for designing biosynthetic pathways, which have successfully replicated two recently reported synthetic pathways. Looking ahead, we aim to incorporate energy requirements into pathway optimization. By introducing an energy threshold
, we hope to filter synthetic routes to identify those that are more efficient and energy-saving. This could enhance the economic feasibility of biosynthetic technologies and contribute to the development of green and sustainable chemistry.
In summary, RDBSB expands the resources available for synthetic biology research and applications, while offering tools for studying the functions of catalytic bioparts. We hope that RDBSB will become a valuable resource for synthetic biology and look forward to making further contributions to the field through ongoing development and collaboration.
Acknowledgements
We would like to thank the Ruijin Luo team from Ezhou Industrial Technology Research Institute, Huazhong University of Science and Technology for their support of website development and Peng Zhang from Bio-Med Big Data Center for the discussion.
Contributor Information
Wan Liu, National Genomics Data Center & Bio-Med Big Data Center, CAS Key Laboratory of Computational Biology, Shanghai Institute of Nutrition and Health, University of Chinese Academy of Sciences, Chinese Academy of Sciences, 320 Yueyang Road, Shanghai 200031, China.
Pingping Wang, CAS-Key Laboratory of Synthetic Biology, CAS Center for Excellence in Molecular Plant Sciences, Chinese Academy of Sciences, 300 Feng Lin Road, Shanghai 200032, China.
Xinhao Zhuang, National Genomics Data Center & Bio-Med Big Data Center, CAS Key Laboratory of Computational Biology, Shanghai Institute of Nutrition and Health, University of Chinese Academy of Sciences, Chinese Academy of Sciences, 320 Yueyang Road, Shanghai 200031, China.
Yunchao Ling, National Genomics Data Center & Bio-Med Big Data Center, CAS Key Laboratory of Computational Biology, Shanghai Institute of Nutrition and Health, University of Chinese Academy of Sciences, Chinese Academy of Sciences, 320 Yueyang Road, Shanghai 200031, China.
Haiyan Liu, School of Life Sciences, University of Science and Technology of China, 443 Huangshan Road, Hefei, Anhui 230026, China.
Sheng Wang, Shanghai Zelixir Biotech Company Ltd., 4/F, Youyue Building, No. 298, Xiangke Road, Pudong New District, Shanghai 200030, China.
Haihan Yu, School of Life Sciences, University of Science and Technology of China, 443 Huangshan Road, Hefei, Anhui 230026, China.
Liangxiao Ma, National Genomics Data Center & Bio-Med Big Data Center, CAS Key Laboratory of Computational Biology, Shanghai Institute of Nutrition and Health, University of Chinese Academy of Sciences, Chinese Academy of Sciences, 320 Yueyang Road, Shanghai 200031, China.
Yuguo Jiang, Shanghai Key Laboratory of Plant Functional Genomics and Resources, Shanghai Chenshan Botanical Garden, and Chenshan Science Research Center, CAS Center for Excellence in Molecular Plant Sciences (CEMPS), Chinese Academy of Sciences (CAS), 3888 Chenhua Road, Shanghai 201602, China.
Guoping Zhao, National Genomics Data Center & Bio-Med Big Data Center, CAS Key Laboratory of Computational Biology, Shanghai Institute of Nutrition and Health, University of Chinese Academy of Sciences, Chinese Academy of Sciences, 320 Yueyang Road, Shanghai 200031, China; School of Life Science, Hangzhou Institute for Advanced Study, University of Chinese Academy of Sciences, 1 Sub-lane Xiangshan, Hangzhou 310024, China.
Xing Yan, CAS-Key Laboratory of Synthetic Biology, CAS Center for Excellence in Molecular Plant Sciences, Chinese Academy of Sciences, 300 Feng Lin Road, Shanghai 200032, China.
Zhihua Zhou, CAS-Key Laboratory of Synthetic Biology, CAS Center for Excellence in Molecular Plant Sciences, Chinese Academy of Sciences, 300 Feng Lin Road, Shanghai 200032, China.
Guoqing Zhang, National Genomics Data Center & Bio-Med Big Data Center, CAS Key Laboratory of Computational Biology, Shanghai Institute of Nutrition and Health, University of Chinese Academy of Sciences, Chinese Academy of Sciences, 320 Yueyang Road, Shanghai 200031, China.
Data availability
RDBSB is freely available at https://www.biosino.org/rdbsb/.
Funding
National Key Research and Development Program of China [2018YFA0900700, 2023YFA0915500]; Strategic Biological Resources Service Network Plan of the Chinese Academy of Sciences [KFJ-BRP-009]; Biological Resources Programme of the Chinese Academy of Sciences [KFJ-BRP-017-79]; Major Research Plan of the National Natural Science Foundation of China [92251307]; Self-supporting Program of Guangzhou Laboratory [SRPG22-007]; Major Project of Guangzhou National Laboratory [GZNL2024A01002].
Conflict of interest statement. None declared.
References
- 1. Jia H., Schwille P.. Bottom-up synthetic biology: reconstitution in space and time. Curr. Opin. Biotechnol. 2019; 60:179–187. [DOI] [PubMed] [Google Scholar]
- 2. Wang Y.-H., Wei K.Y., Smolke C.D.. Synthetic biology: advancing the design of diverse genetic systems. Annu. Rev. Chem. Biomol. Eng. 2013; 4:69–102. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 3. Canton B., Labno A., Endy D.. Refinement and standardization of synthetic biological parts and devices. Nat. Biotechnol. 2008; 26:787–793. [DOI] [PubMed] [Google Scholar]
- 4. Keasling J.D. Manufacturing molecules through metabolic engineering. Science. 2010; 330:1355–1358. [DOI] [PubMed] [Google Scholar]
- 5. Nielsen J., Keasling J.D.. Engineering cellular metabolism. Cell. 2016; 164:1185–1197. [DOI] [PubMed] [Google Scholar]
- 6. Ellis T., Adie T., Baldwin G.S.. DNA assembly for synthetic biology: from parts to pathways and beyond. Integr. Biol. 2011; 3:109–118. [DOI] [PubMed] [Google Scholar]
- 7. Liang J., Luo Y., Zhao H.. Synthetic biology: putting synthesis into biology. Wiley Interdiscip. Rev. Syst. Biol. Med. 2011; 3:7–20. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 8. Carothers J.M., Goler J.A., Keasling J.D.. Chemical synthesis using synthetic biology. Curr. Opin. Biotechnol. 2009; 20:498–503. [DOI] [PubMed] [Google Scholar]
- 9. Smanski M.J., Zhou H., Claesen J., Shen B., Fischbach M.A., Voigt C.A.. Synthetic biology to access and expand nature’s chemical diversity. Nat. Rev. Microbiol. 2016; 14:135–149. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 10. Frasch H.J., Medema M.H., Takano E., Breitling R.. Design-based re-engineering of biosynthetic gene clusters: plug-and-play in practice. Curr. Opin. Biotechnol. 2013; 24:1144–1150. [DOI] [PubMed] [Google Scholar]
- 11. Ren H., Hu P., Zhao H.. A plug-and-play pathway refactoring workflow for natural product research in Escherichia coli and Saccharomyces cerevisiae. Biotechnol. Bioeng. 2017; 114:1847–1854. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 12. Moon H. iGEM 2021: a year in review. Biodes. Res. 2022; 2022:9794609. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 13. UniProt C. UniProt: the universal protein knowledgebase in 2023. Nucleic Acids Res. 2023; 51:D523–D531. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 14. Caspi R., Billington R., Keseler I.M., Kothari A., Krummenacker M., Midford P.E., Ong W.K., Paley S., Subhraveti P., Karp P.D.. The MetaCyc database of metabolic pathways and enzymes—a 2019 update. Nucleic Acids Res. 2020; 48:D445–D453. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 15. Kanehisa M., Furumichi M., Sato Y., Kawashima M., Ishiguro-Watanabe M.. KEGG for taxonomy-based analysis of pathways and genomes. Nucleic Acids Res. 2023; 51:D587–D592. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 16. Drula E., Garron M.L., Dogan S., Lombard V., Henrissat B., Terrapon N.. The carbohydrate-active enzyme database: functions and literature. Nucleic Acids Res. 2022; 50:D571–D577. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 17. Chang A., Jeske L., Ulbrich S., Hofmann J., Koblitz J., Schomburg I., Neumann-Schaal M., Jahn D., Schomburg D.. BRENDA, the ELIXIR core data resource in 2021: new developments and updates. Nucleic Acids Res. 2021; 49:D498–D508. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 18. Zhou C., Xu Q., He S., Ye W., Cao R., Wang P., Ling Y., Yan X., Wang Q., Zhang G.. GTDB: an integrated resource for glycosyltransferase sequences and annotations. Database. 2020; 2020:baaa047. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 19. Jumper J., Evans R., Pritzel A., Green T., Figurnov M., Ronneberger O., Tunyasuvunakool K., Bates R., Zidek A., Potapenko A.et al.. Highly accurate protein structure prediction with AlphaFold. Nature. 2021; 596:583–589. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 20. Yufeng Liu L.C., Liu H.. Diffusion in a quantized vector space generates non-idealized protein structures and predicts conformational distributions. 2023; bioRxiv doi:18 November 2023, preprint: not peer reviewed 10.1101/2023.11.18.567666. [DOI]
- 21. Sayers E.W., Cavanaugh M., Clark K., Pruitt K.D., Sherry S.T., Yankie L., Karsch-Mizrachi I.. GenBank 2024 update. Nucleic Acids Res. 2024; 52:D134–D137. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 22. Bansal P., Morgat A., Axelsen K.B., Muthukrishnan V., Coudert E., Aimo L., Hyka-Nouspikel N., Gasteiger E., Kerhornou A., Neto T.B.et al.. Rhea, the reaction knowledgebase in 2022. Nucleic Acids Res. 2022; 50:D693–D700. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 23. Lang M., Stelzer M., Schomburg D.. BKM-react, an integrated biochemical reaction database. BMC Biochem. 2011; 12:42. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 24. Kim S., Chen J., Cheng T., Gindulyte A., He J., He S., Li Q., Shoemaker B.A., Thiessen P.A., Yu B.et al.. PubChem 2023 update. Nucleic Acids Res. 2023; 51:D1373–D1380. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 25. Hastings J., de Matos P., Dekker A., Ennis M., Harsha B., Kale N., Muthukrishnan V., Owen G., Turner S., Williams M.et al.. The ChEBI reference database and ontology for biologically relevant chemistry: enhancements for 2013. Nucleic Acids Res. 2013; 41:D456–D463. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 26. Bairoch A. The ENZYME database in 2000. Nucleic Acids Res. 2000; 28:304–305. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 27. Mistry J., Chuguransky S., Williams L., Qureshi M., Salazar G.A., Sonnhammer E.L.L., Tosatto S.C.E., Paladin L., Raj S., Richardson L.J.et al.. Pfam: the protein families database in 2021. Nucleic Acids Res. 2021; 49:D412–D419. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 28. Varadi M., Bertoni D., Magana P., Paramval U., Pidruchna I., Radhakrishnan M., Tsenkov M., Nair S., Mirdita M., Yeo J.et al.. AlphaFold Protein Structure Database in 2024: providing structure coverage for over 214 million protein sequences. Nucleic Acids Res. 2024; 52:D368–D375. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 29. Schoch C.L., Ciufo S., Domrachev M., Hotton C.L., Kannan S., Khovanskaya R., Leipe D., McVeigh R., O’Neill K., Robbertse B.et al.. NCBI Taxonomy: a comprehensive update on curation, resources and tools. Database. 2020; 2020:baaa062. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 30. Potter S.C., Luciani A., Eddy S.R., Park Y., Lopez R., Finn R.D.. HMMER web server: 2018 update. Nucleic Acids Res. 2018; 46:W200–W204. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 31. Zheng J., Hu B., Zhang X., Ge Q., Yan Y., Akresi J., Piyush V., Huang L., Yin Y.. dbCAN-seq update: CAZyme gene clusters and substrates in microbiomes. Nucleic Acids Res. 2023; 51:D557–D563. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 32. Sayers E.W., Beck J., Bolton E.E., Bourexis D., Brister J.R., Canese K., Comeau D.C., Funk K., Kim S., Klimke W.et al.. Database resources of the National Center for Biotechnology Information. Nucleic Acids Res. 2021; 49:D10–D17. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 33. Eddy S.R. Accelerated profile HMM searches. PLoS Comput. Biol. 2011; 7:e1002195. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 34. Kerwin S.M. ChemBioOffice Ultra 2010 suite. J. Am. Chem. Soc. 2010; 132:2466–2467. [DOI] [PubMed] [Google Scholar]
- 35. Boratyn G.M., Camacho C., Cooper P.S., Coulouris G., Fong A., Ma N., Madden T.L., Matten W.T., McGinnis S.D., Merezhuk Y.et al.. BLAST: a more efficient report with usability improvements. Nucleic Acids Res. 2013; 41:W29–W33. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 36. Sehnal D., Bittrich S., Deshpande M., Svobodova R., Berka K., Bazgier V., Velankar S., Burley S.K., Koca J., Rose A.S.. Mol* Viewer: modern web app for 3D visualization and analysis of large biomolecular structures. Nucleic Acids Res. 2021; 49:W431–W437. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 37. Backman J.T., Filppula A.M., Niemi M., Neuvonen P.J.. Role of cytochrome P450 2C8 in drug metabolism and interactions. Pharmacol. Rev. 2016; 68:168–241. [DOI] [PubMed] [Google Scholar]
- 38. Fan Y., Li S., Ding X., Yue J., Jiang J., Zhao H., Hao R., Qiu W., Liu K., Li Y.. First-in-class immune-modulating small molecule icaritin in advanced hepatocellular carcinoma: preliminary results of safety, durable survival and immune biomarkers. BMC Cancer. 2019; 19:279. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 39. Wang P., Li C., Li X., Huang W., Wang Y., Wang J., Zhang Y., Yang X., Yan X., Wang Y.. Complete biosynthesis of the potential medicine icaritin by engineered Saccharomyces cerevisiae and Escherichia coli. Sci. Bull. 2021; 66:1906–1916. [DOI] [PubMed] [Google Scholar]
- 40. Jiang Y., Ma J., Wei Y., Liu Y., Zhou Z., Huang Y., Wang P., Yan X.. D e novo biosynthesis of sex pheromone components of Helicoverpa armigera through an artificial pathway in yeast. Green Chem. 2022; 24:767–778. [Google Scholar]
- 41. Shao Y., Lu N., Wu Z., Cai C., Wang S., Zhang L.-L., Zhou F., Xiao S., Liu L., Zeng X.. Creating a functional single-chromosome yeast. Nature. 2018; 560:331–335. [DOI] [PubMed] [Google Scholar]
- 42. Yang C., Li C., Wei W., Wei Y., Liu Q., Zhao G., Yue J., Yan X., Wang P., Zhou Z.. The unprecedented diversity of UGT94-family UDP-glycosyltransferases in Panax plants and their contribution to ginsenoside biosynthesis. Sci. Rep. 2020; 10:15394. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 43. Fan Z., Wang Y., Yang C., Zhou Z., Wang P., Yan X.. Identification of a novel multifunctional oxidosqualene cyclase from Zea mays sheds light on the biosynthetic pathway of three pentacyclic triterpenoids. Synth. Syst. Biotechnol. 2022; 7:1167–1172. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 44. Li C., Yan X., Xu Z., Wang Y., Shen X., Zhang L., Zhou Z., Wang P.. Pathway elucidation of bioactive rhamnosylated ginsenosides in Panax ginseng and their de novo high-level production by engineered Saccharomyces cerevisiae. Commun. Biol. 2022; 5:775. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 45. Xing H., Zou G., Liu C., Chai S., Yan X., Li X., Liu R., Yang Y., Zhou Z.. Improving the thermostability of a GH11 xylanase by directed evolution and rational design guided by B-factor analysis. Enzyme Microb. Technol. 2021; 143:109720. [DOI] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Data Availability Statement
RDBSB is freely available at https://www.biosino.org/rdbsb/.




