Skip to main content
ACS AuthorChoice logoLink to ACS AuthorChoice
. 2025 Feb 14;14(3):756–770. doi: 10.1021/acssynbio.4c00692

minChemBio: Expanding Chemical Synthesis with Chemo-Enzymatic Pathways Using Minimal Transitions

Mohit Anand , Vikas Upadhyay , Costas D Maranas †,*
PMCID: PMC11934129  PMID: 39951253

Abstract

graphic file with name sb4c00692_0013.jpg

Chemo-enzymatic pathway design aims to combine the strengths of enzymatic with chemical synthesis to traverse biomolecular design space more efficiently. While chemical reactions often struggle with regioselectivity and stereoselectivity, enzymatic conversions often encounter limitations of low enzyme activity or availability. Optimally integrating both approaches provides an opportunity to identify efficient pathways beyond the capabilities of either modality. Recently, studies have shown the advantage of leveraging enzymatic steps into industrial-scale chemical processes, such as for the blood sugar regulator Sitagliptin (Merck) and the HIV protease inhibitor Darunavir (Prozomix). Designing optimal chemo-enzymatic pathways is a complex task. It requires navigating a high-dimensional search space of potential reactions that combine individual chemical and biochemical steps while at the same time minimizing transitions between chemical catalysis and bioreactions. Here, we introduce an algorithmic approach, minChemBio, that relies on solving a mixed-integer linear programming (MILP) problem by optimally searching through known chemical and enzymatic steps extracted from the United States Patent Office (USPTO) and MetaNetX databases, respectively. minChemBio allows for the minimization of transitions between chemical and biological reactions in the pathway, thus reducing the need for costly separation and purification steps required. minChemBio was benchmarked on three case studies involving the synthesis of 2–5-furandicarboxylic acid, terephthalate, and 3-hydroxybutyrate. Identified designs included both established literature pathways as well as unexplored ones which were compared against pathways identified by existing retrosynthetic tools. minChemBio fills a current gap in the space of pathway retrosynthesis tools by controlling and minimizing the transitions between chemical catalysis and biocatalytic steps. It is accessible to users through open-source code (https://github.com/maranasgroup/chemo-enz).

Keywords: Computational Synthesis Planning, Chemo-enzymatic Pathways, Minimal Transitions, Mixed Integer Linear Programming

Introduction

Synthesis planning often begins with the design of a reaction pathway from a starting precursor to a target molecule. Recently, asymmetric synthesis of a therapeutic drug (+)-Xyloketal B was achieved by leveraging a chemo-enzymatic pathway, which included a biocatalytic hydroxylation step to synthesize the key stereoselective intermediate ortho-quinone methide with high yield.1 In another example, a biosynthetic step was used to introduce site-selective fluorination to complex polyketides.2 These hybrid chem/bio pathways offer opportunities to improve yield, provide milder reaction conditions, achieve stereo- and regio-selectivity, avoid toxic intermediates, or minimize transitions between process conditions. This motivates the development of customized computational tools to optimally leverage these synergies.

Retrosynthesis tools rely on recursively applying reaction templates or rules, which are derived from existing reactions to encapsulate their underlying chemistry, in reverse from the target product until an acceptable starting precursor(s) is identified. There has been significant progress in the development of such automated tools.3 For example, Chematica (or Synthia)4 accesses hundreds of thousands of expert-derived reaction rules organized in the form of a decision tree algorithm for multistep organic synthesis planning of complex natural products. The ASCKOS software suite5 demonstrated pathway planning for 15 medicinally relevant small molecule drugs, with a semiautomatic robotic platform performing the experiments in a continuous flow reactor. Levin et. al6 ranked enzymatic and chemical steps simultaneously generated from a combination of 7,984 reaction templates extracted from 37,000 biological reactions for the BKMS data set7 and 163,723 templates, borrowed from the retrosynthetic tool ASCKOS5 which extracted 12.5 million synthetic chemical reactions from the Reaxys data set.8 Sankarnarayanan et. al9 employed a different approach, using 105 expertly coded enzymatic reaction rules from RetroBioCat,10 a retro-biosynthetic tool, to substitute one or more steps of chemical reactions proposed by ASCKOS,5 thereby generating chemo-enzymatic pathways. One of the challenges with retrosynthesis tools is that pathway design criteria are applied a posteriori after candidate pathways are generated. In contrast, the biosynthetic pathway design tool optStoic11 imposes carbon efficiency criteria at the scoping phase encoded as optimal overall stoichiometries using cataloged reactions in the KEGG database.12 Expanding upon optStoic, novoStoic13 uses reaction rules to encode novel reactions and SMILES (Simplified Molecular Input Line Entry System) representations for metabolites. This allows for the seamless blending of cataloged reactions with novel steps. Tools such as dGPredictor14 can be used to assess the thermodynamic feasibility of the directionality of a biochemical step. In general, enzymatic reactions are carbon-efficient, stereoselective,15 high-yielding,16 and can be performed under mild conditions within a single production host.17 However, acceptable enzymatic alternatives are not always found to form a complete pathway. These gaps are often filled using chemical catalysis leveraging unit operations developed over many decades in the chemical process industry (CPI). For instance, small-molecule active pharmaceutical ingredient (APIs) synthesis, necessary for wide-ranging applications in pharmaceutical industry, primarily employs chemical catalysis reactions18 such as Suzuki and other cross coupling reactions such as transfer hydrogenation reactions.19 What is lacking in the current armamentarium of computation tools is a way of seamlessly blending both enzymatic with chemical catalysis steps in an efficient manner that minimizes transitions associated with costly purification/separation steps.

Chemical reactions rely on inorganic catalysts, while biological reactions require enzymes within the cell to produce the desired target molecule. Biological reactions offer stereospecificity and regioselectivity, whereas chemical reactions usually provide high throughput and rate advantages. For instance, enzymatic steps with poor enzymatic performance can be replaced by chemical catalysis analogues. Enzymatically catalyzed reactions can either be integrated in a pathway within a single production host without the need for intermediates separation (i.e., one pot synthesis),20,21 or in the presence of toxic intermediates the pathway can be assembled in vitro, and the product can be reintroduced into the cell.22 However, separation and purification steps between successive chemical reactions (see Figure 1) can significantly increase the synthesis cost. In chemical synthesis and manufacturing, a perfect atom economy, where all atoms in the reactants are retained in the final product, is rarely achieved, often resulting in side-products. Chemical processes often require solvents that must be removed and recycled. Separation steps are thus required when transitioning from a chemical to a biological step, and vice versa, and even between two consecutive chemical catalysis steps. minChemBio directly minimizes all such transitions.

Figure 1.

Figure 1

A schematic representing transitions in chemo-enzymatic pathway. The starting molecule undergoes a chemical reaction to form an intermediate, which requires a separation step before being used as a reactant in subsequent biological reactions. Finally, the target molecule is separated.

Herein, we introduce minChemBio, an optimization-based search method that finds reaction pathways from a starting precursor to a target molecule with the minimal number of transitions. minChemBio tracks the main carbon path of the reaction pathway to link reaction steps. Enzyme-catalyzed reactions are extracted from the MetaNetX database23 whereas chemical reactions are ported from the USPTO24 data set. We chose to trace the main carbon path instead of capturing the exact stoichiometry because most reaction entries from the USPTO database lack elemental balance. Entries from the USPTO database are first preprocessed to remove duplicates, incorrectly annotated reactions, and reactions with nonspecific (generic) reactant descriptions. Since entries in the USPTO database are not elementally and charge balanced, only a one-to-one mapping between the main reactant and product molecules are tracked using the tanimoto similarity coefficient from RDkit python package25 that uses molecular fingerprint-based similarity. Similarly, one-to-one mappings between main reactant to product molecules are generated for elementally balanced enzymatic reaction entries from MetaNetX23 upon removal of cofactors and small coreactants such as CO2, water, ammonia, H2O2, etc. Reactions in the MetaNetX23 are collated from various sources such as BIGG,26 ChEBI,27 KEGG,12 MODELSEED,28 RHEA,29 and SABIO-RK30 but with nonunanimous directionality across the sources, thus generating more than one one-to-one mapping from a single entry.

The tool was demonstrated by addressing the synthesis of high-volume platform chemicals (i.e., 2,5-furandicarboxylic acid) as well as important polymer monomers (i.e., Terephthalate and (R)-3-hydroxybutyrate). In this study, we used the web platforms provided by ASKCOS and RetroBioCat to compare the solutions for three case studies against the results from minChemBio. Unlike ASKCOS and RetroBioCat, which do not aim to prioritize known reactions or customizable precursors, minChemBio exclusively employs known reactions, enabling the design of cost-effective pathways using both starting and target molecules. In the case study of 2,5-furandicarboxylic acid, minChemBio identified two chemo-enzymatic pathways, among them one is previously unexplored, that begin with the abundant precursor glucose. Note that both ASKCOS and RetroBioCat did not generate any pathway starting from glucose. In the case study of Terephthalate, minChemBio found five previously unexplored chemo-enzymatic pathways using known reactions. Similarly, for (R)-3-hydroxybutyrate, two of five biological pathways identified from pyruvate were previously reported in literature but seemed inaccessible by ASKCOS and RetroBioCat.

Results

In this study, we put forth a set of target chemicals (see Table 1) and potential starting chemicals from a set of four common metabolic precursors: glucose, pyruvate, p-xylene, and glycerol. To determine the most plausible synthetic routes, we calculated structural similarities between the target molecules and each precursor using Tanimoto similarity coefficient. Pathways based on precursors with the highest similarity scores to the product were computationally designed. Literature searches were conducted to compare proposed pathways with previously reported routes. Existing retrosynthetic tools primarily rely on single-step retrosynthetic models, either template-based or template-free, and employ methods such as Monte Carlo Tree Search (MCTS)31 or AND-OR Tree Search32,33 to navigate through recursive retrosynthetic steps to reach commercially available precursors. The choice of ASKCOS and RetroBioCat for comparison was driven by their availability as free web platforms. In contrast, other chemo-enzymatic pathway search models,6,9 which also leverage ASKCOS for chemical steps and either RetroBioCat or reaction rules from the BKMS database for biological steps respectively, lack freely available web platforms.6,9 minChemBio uncovered additional pathways not reported in previous studies. In addition, thermodynamic feasibility analyses were performed for each enzymatic step using dGPredictor,14 a computational tool that estimates the change in standard Gibbs energy change at pH 7.0 and ionic strength of 0.1 M.

Table 1. Three Target and Starting Molecules Used to Test the minChemBio Tool with the Number of Solutions in Each Case.

target molecule starting precursor(s) # solutions by minChemBio
2,5-furandicarboxylic acid glucose 23
terephthalic acid p-xylene 21
(R)-3-hydroxy butyrate pyruvate, glycerol 21, 24

Case Study 1: 2,5-Furandicarboxylic Acid (FDCA)

FDCA is a versatile Bioplastic precursor for polyesters, polyamides, and valuable furanic chemicals, and offers a sustainable alternative to petroleum-based materials such as polyethylene terephthalate (PET).34 Traditionally, single-step FDCA production relied on 5-hydroxymethylfurfural (5-HMF), with various approaches emerging through microbial conversion35 such as sequential enzymatic oxidations36 including 5-hydroxymethylfurfural oxidase,37 and an immobilized laccase catalyst with 2,2,6,6-tetramethyl-piperidin-1-yl)oxyl (TEMPO) as mediator.34

minChemBio explored 23 chemo-enzymatic pathways (see Supporting Information) for FDCA production using glucose as a starting precursor and two potential solutions are shown in Figure 2a. Pathway 1 (1–4–5 see Figure 2a) is a three-step pathway which begins with the isomerization of glucose to keto-D-fructose, followed by conversion of keto-D-fructose to 5-hydroxy furfural (5-HMF) through a chemical reaction. The corresponding patent literature informs the use of high fructose maize syrup as the raw material to produce 5-HMF.38 Pathway 2 (2–3–4–5 see Figure 2a) incorporates an additional step to reach the intermediate keto-D-fructose, while sharing the final two steps with pathway 1. Previously explored pathways for FDCA production align with minChemBio’s solution steps 4 and 5 (see Figure 2a), as they utilize fructose as the starting precursor due to its furanose structure, followed by two successive synthetic chemical reactions30, 31.

Figure 2.

Figure 2

Pathways identified for the target molecule FDCA. (a) Two pathways were found using the chemo-enzymatic MILP pathway for target chemical FDCA from the source chemical glucose. The green arrows and the red arrows depict biological and chemical reactions, respectively. Pathway 1 is the shortest pathway found with three steps, and pathway 2 involves adding one step to get the same intermediate, fructose, and follow the same pathway for the rest of the way as pathway 1. Here, in pathway 1, step 1 is an isomerase reaction (MNXR192050); step 4, a USPTO chemical reaction (US09206148B2), is a cyclization reaction; step 5 is a 5-(hydroxymethyl) furfural oxidase (MNXR114172). In pathway 2, step 2 is a D-sorbitol reductase reaction (MNXR192036), and step 3 is a D-sorbitol 2-dehydrogenase reaction (MNXR146383). (b) Retrosynthetic ASKCOS solution shown by choosing the pathway with highest plausibility score. This reaction was based on the reaction template, which converts furfural to FDCA. One pathway from the set of solutions found by RetroBioCat is shown here, where the aldehyde oxidation template is used.

We analyzed the thermodynamic feasibility of these pathways found shown in Figure 1a using dGPredictor.14 The initial glucose isomerization reaction, (i.e., step 1 in Figure 2a), has an estimated standard Gibbs energy change of 1.21 kJ/mol ±0.2 kJ/mol. Step 2, catalyzed by D-sorbitol reductase, has an estimated standard Gibbs energy change of −9.75 kJ/mol ±0.11 kJ/mol. The D-sorbitol 2-dehydrogenase reaction, (i.e., step 3 in Figure 2a) which produces keto-D-fructose from D-sorbitol has an estimated Gibbs energy of 13.4 kJ/mol ±0.2 kJ/mol. Finally, step 5, the 5-HMF oxidase reaction has a substantially negative predicted Gibbs energy of −453 kJ/mol ±2.91 kJ/mol. Pathway 1 (steps 1–4–5 see Figure 2a) appears more thermodynamically favorable overall due to the lack of endergonic steps, while Pathway 2 (steps 2–3–4–5 see Figure 2a) has a positive Gibbs energy change in step 3 that could be a potential challenge, requiring additional considerations to maintain feasibility.

ASCKOS identified 33 reaction trees after analyzing 708 chemicals and 761 reactions using MCTS tree builder search and Reaxys model for reaction rules. Each identified tree represents a pathway containing either one or two steps, yet none of these pathways utilized a precursor that was less expensive than the target molecule (see Figure 2b). Similarly, RetroBioCat generated 327 pathway solutions, ranging from one to five steps, but again, none of the identified starting materials were less costly than the target molecule. Most pathways in RetroBioCat’s analysis began with a furan derivative, leading to the formation of FDCA through successive oxidation steps (see Figure 2b). In contrast, minChemBio identified a cyclization reaction at Step 4, where glucose served as the starting molecule.

Case Study 2: Terephthalate/Terephthalic Acid (TPA)

Terephthalic acid (TPA), the precursor to polyethylene terephthalate (PET), is widely produced via the AMOCO process. This process uses a combination of cobalt (Co), manganese (Mn), and bromide (Br) as liquid metal catalysts, air as an oxidant, and acetic acid as a solvent at high temperatures. Despite high p-xylene (PX) conversion and TPA yield,39,40 the AMOCO process poses environmental concerns due to possible acetic acid loss, heavy metal contamination, and high energy demands.40 This has motivated the search for sustainable alternatives, notably biocatalytic methods. Bramucci et al.41 demonstrated successive enzymatic oxidation from PX to 4-carboxybenzyl alcohol (4-CBAL) using xylene monooxygenase. However, subsequent enzymatic oxidation of 4-CBAL to TPA remained low-yielding. Recent work by Luo et al.40 suggests that bioengineering pathways to suppress back-reactions to p-toluate (PTA), the intermediate before 4-CBAL, could significantly enhance biocatalytic TPA yields, potentially rivaling those of conventional synthetic chemical processes.

minChemBio identified eight potential routes for TPA production among a total of 21 solutions (see Supporting Information), starting from PX as shown in Figure 3a. Pathway 1 (steps 1–2–3–4–5–6 see Figure 3a) is fully enzymatic with 6 steps of successive oxidation, while the remaining pathways incorporate at least one chemical step that bypasses one or more of these oxidation steps. To assess the feasibility of pathway 1, we analyzed the thermodynamic favorability of each enzymatic step using dGPredictor.14 Step 2, an aldehyde reductase reaction, and step 5, an alcohol dehydrogenase/aldehyde reductase reaction both possess positive estimates of Gibbs energy change values at 18.4 kJ/mol ±0.11 kJ/mol. The remaining enzymatic steps 1, 3, 4, and 6 have negative estimates of standard Gibbs energy change values.

Figure 3.

Figure 3

Pathways identified for the target molecule TPA. The green arrows and the red arrows depict biological and chemical reactions, respectively. MetaNetX reaction IDs and United States Patent numbers are mentioned for biological and chemical steps, respectively, in the description. (a) Eight pathway solutions from minChemBio are shown using 11 reactions. Successive oxidation of pX is the essence of these steps. Pathway 1, using steps 1–2–3–4–5–6, is a fully enzymatic pathway. The remaining seven pathways have at least one chemical step. Among the biological reactions, step 1 is a monooxygenase reaction (MNXR146191); step 2 is an aldehyde reductase (MNXR130059); step 3 is a p-methylbenzaldehyde:NADP+ oxidoreductase (MNXR109459); step 4 is 4-toluene carboxylate monooxygenase (MNXR123901); step 5 is an aldehyde reductase (MNXR132891); and step 6 is 4-CBA dehydrogenase (MNXR115792). Among the chemical reactions, step 11 is similar to the conventional AMOCO process with liquid metal Co-catalyst (US07115541B2); step 8 represents the minor product of the reaction in step 11 where the yield of pTA is 2% and for TPA is 90%; step 10 is the oxidation of pTA to TPA with Co-acetate as a catalyst (US04245078); step 12 also represents the same patent literature as step 10 but the starting substrate is p-TALD instead of pTA; step 7 shows partial oxidation of pX to p-TALC (ethylene as coreactant in the presence of air and ethylene oxide as coproduct where selectivity toward ethylene oxide; US04046782); step 9 is another oxidation reaction that converts p-TALC to p-TA (US07064226B2). (b) ASCKOS and RetroBioCat solutions are shown here. Six pathways, each containing one retrosynthetic step starting from a cheaper precursor are chosen from the list of top pathways found by ASKCOS. Four pathways, each containing one retrosynthetic step starting from a precursor found by RetroBioCat are shown here.

Pathway 2 (steps 8–4–5–6 see Figure 3a) avoids thermodynamically infeasible step 2, by producing PTA (step 8 in Figure 3a)5 and bypassing the first three steps and sharing the final three enzymatic steps with pathway 1. Pathways 3 (steps 1–2–12 see Figure 3a) and 4 (steps 1–2–3–10 see Figure 3a) also contain one transition from biological to the chemical environment where p-tolualdehye (p-TALD) is oxidized to terephthalate through step 12 and PTA is oxidized to TPA through step 10 respectively. Both steps 10 and 12 map to the same USPTO chemical reaction42 that uses cobalt-acetate as the catalyst. Pathway 5 (steps 8–10 see Figure 3a) contains a chemical-to-chemical transition and thus requires one separation and purification step to isolate product p-toluate. Pathways 6 (steps 1–9–4–5–6 see Figure 3a) and 7 (steps 7–9–4–5–6 see Figure 3a) involve the same three intermediate molecules (p-TALC, pTA, 4-CBAL), but the former contains only one chemical reaction, step 9 in Figure 3a, whereas the latter contains two chemical catalysis reaction steps 7 and 9. Step 7 is an oxidation reaction where PX is converted to 4-methyl benzyl alcohol (p-TALC) and ethylene serves as a coreactant which generates ethylene oxide as a coproduct with high selectivity.43 Therefore, the highly exergonic enzymatic reaction (i.e., step 1 in Figure 3a) with an estimated standard Gibbs energy change of −393.22 kJ/mol ±1.34 kJ/mol, is more favorable over the chemical reaction, (i.e., step 7 in Figure 3a) since there is no coreactant which has higher selectivity than PX. Pathway 8 (steps 7–9–10 see Figure 3a) involves only chemical reactions and overcomes the thermodynamically challenging enzymatic reactions of Pathway 1 but involves two transitions requiring additional separation and purification steps. Ultimately, a comprehensive evaluation of these pathways would require a techno-economic analysis to determine the most promising strategies for terephthalate production.

ASCKOS identified 96 reaction trees with a depth of one and two after exploring a set of 206 chemicals and 214 reactions using Retro-star tree builder and Reaxys model for reaction rules. Most of the starting materials or intermediates were derivatives of benzene or PX. Among the pathways identified, six are illustrated in Figure 3b, where the starting precursors are less expensive compared to the target molecule, TPA. Notably, the conversion of p-Xylene to terephthalate is the only known reaction and coincides with step 11 in the minChemBio solutions; the remaining reactions rely on reaction templates. RetroBioCat generated 150 pathway solutions, but only four of these pathways included precursor molecules present in the PubChem database,44 and thus, the rest were excluded from further consideration. Of these four, the pathway beginning with 4-carboxybenzaldehyde (4-CBA) features a known aldehyde oxidation step, which corresponds to step 6 in the minChemBio solutions. The remaining pathways involve novel reaction steps based on reaction rules.

Case Study 3: (R)-3-Hydroxybutyrate (3HB)

Polyhydroxyalkanoates (PHAs) are a class of microbially produced biopolyesters accumulated under conditions of nutrient limitation and excess carbon.45 Of these, poly hydroxybutyrate (PHB) is a promising Bioplastic candidate due to its versatility in applications such as packaging, medical implants, and drug delivery.46 PHAs offer desirable properties including biodegradability, biocompatibility, water resistance, and oxygen impermeability.47 Despite their potential, PHAs have limited market share (1.8% of global plastics) due to production costs, scalability challenges, and limitations of the primary PHB homopolymer.48 High substrate costs are a major factor, contributing up to 40% of the final product price.49 PHB biosynthesis begins with β-ketothiolase catalyzed condensation of excess acetyl-CoA to acetoacetyl-CoA, followed by reduction and CoA removal to yield the PHB precursor, 3-hydroxybutyryl-Co.50

minChemBio identified 21 potential pathways for (R)-3-hydroxybutyrate (3HB) production (see Supporting Information) starting from pyruvate, five among them are shown in Figure 4a. Pathways 1 through 4 are fully enzymatic, while pathway 5 incorporates a chemical step. Except for pathway 4, all pathways share the same final step, (i.e., step 3) shown in Figure 4a, the 3-hydroxybutyrate dehydrogenase reaction that converts acetoacetate to 3HB. Pathway 1 (steps 8–3 see Figure 4a) consists of two steps, including an aminotransferase reaction, (i.e., step 8 in Figure 4a), that requires 3-aminobutanoate as a cosubstrate. Notably, this pathway was not found among previously reported50 enzymatic pathways. This finding highlights the importance of considering the full spectrum of known reactions when exploring potential biosynthetic routes. Pathway 2 (9–10–3 see Figure 4a) begins with a pyruvate-oxime:acetone oximinotransferase reaction, i.e. step 9 in Figure 4a, followed by a thermodynamically unfavorable acetone carboxylation reaction, (i.e., step 10 in Figure 4a), with an estimated Gibbs energy change of 14.8 kJ/mol ±0.17 kJ/mol. Pathways 3 (steps 7–6–11–3 in Figure 4a) and 4 (steps 7–6–5–4 in Figure 4a) are four step pathways, previously reported in the literature.50 They share two initial steps and contain a thermodynamically unfavorable acetyl-CoA C-acetyltransferase reaction, (i.e., step 6 in Figure 4a), with an estimated Gibbs energy change of 26.31 kJ/mol ±0.11 kJ/mol. These pathways differ in their last two steps, with pathway 3 relying on a hydrolase reaction, (i.e., step 11 in Figure 4a), followed by 3HB dehydrogenase, whereas pathway 4 employs an aldehyde reductase (i.e., step 5 in Figure 4a) and an acyl-CoA thioesterase (i.e., step 4 in Figure 4a).

Figure 4.

Figure 4

Pathways identified for the target molecule 3HB. (a) Five pathways were found to R-3-hydroxybutyrate (3HB) by minChemBio using pyruvate as a starting precursor. The green arrows and the red arrows depict biological and chemical reactions, respectively. Pathways 1 to 4 are fully enzymatic pathways, and pathway 5 is a chemo-enzymatic pathway with one chemical reaction, step 2. Here, step 1 is a pyruvate dehydrogenase reaction (MNXR182167); step 2 converts acetate to acetoacetate (alternatively achieved by the biochemical reaction MNXR190492); step 3 is the 3-HB dehydrogenase reaction (MNXR96232); step 4 is an acyl-CoA thioesterase reaction (MNXR118585); step 5 is an acetoacetyl-CoA reductase reaction (MNXR145174); step 6 is an acetyl-CoA acetyltransferase reaction (MNXR190494); step 7 is a pyruvate dehydrogenase reaction (MNXR106432); step 8 is an aminotransferase reaction (MNXR194248); step 9 is a pyruvate oxime:acetone oxime transferase reaction (MNXR148359); step 10 is the acetoacetate decarboxylase reaction (MNXR95442); and step 11 is an acetoacetyl-CoA hydrolase reaction (MNXR106949). (b) ASKCOS and Retrobiocat solutions were found using their web platform. Four pathways are shown found by ASCKOS where the starting molecule is cheaper than the target molecule. One pathway shown was found by RetroBioCat where the starting molecule is pyruvate.

Pathway 5 (steps 1–2–3 in Figure 4a) is a hybrid pathway, which includes a chemical step converting acetate to acetoacetate. Examining the patent literature suggested this step was likely misannotated in the USPTO data set. While a biochemical alternative for step 2 using acetoacetyl-CoA: acetate CoA-transferase reaction is possible, this example highlights the need for careful verification of predicted hybrid pathways through manual inspection of the patent literature. In summary, pathway 1 has no predicted thermodynamically unfavorable steps whereas pathways 2, 3, and 4 each contain one unfavorable step.

ASCKOS identified 200 reaction trees of depths one and two after exploring a total of 756 chemicals and 867 reactions using MCTS tree builder search and Reaxys model for reaction rules. Figure 4b presents the top four pathways discovered by ASCKOS, where the starting chemicals are cheaper than the target molecule. In contrast, RetroBioCat identified 5,142 pathways, with number of steps ranging from one to six. However, only a small subset of these pathways begins with pyruvate, a starting material that is less expensive than the target molecule. One such pathway, where pyruvate is the starting molecule, is shown in Figure 4b that does not contain any known reactions.

Using glycerol as a precursor, minChemBio identified 24 pathways (see Supporting Information) among which two potential synthesis routes for 3-hydroxybutyrate (3-HB) are shown in Figure 5. In pathway 1, (steps 1–2–3–4 in Figure 5), glycerol first undergoes dehydration through glycerol dehydratase reaction, (i.e., step 1 in Figure 5) to form 3-hydroxypropanal. This intermediate is then oxidized by 3-hydroxypropanal dehydrogenase, (i.e., step 2 in Figure 5), to produce 3-hydroxypropanoate. A limitation arises in the subsequent CoA transferase reaction, (i.e., step 3 in Figure 5), where a transferase reaction occurs. Here, the main carbon path from the intermediate 3-hydroxy propanoate leads to 3-hydroxypropanoyl-CoA but acetoacetate is treated as the primary intermediate in the pathway. Transferase reactions lead to one-to-one mappings sometimes that do not track the main carbon path which is a limitation of the existing tool (see MetaNetX biological reaction database in Methods). Pathway 2 (steps 5–6–7–4 in Figure 5) employs a three-step chemo-enzymatic approach. An initial chemical step converts glycerol to acetaldehyde. Because the USPTO reaction data set lacks major/minor product distinctions, acetaldehyde is chosen over the primary product acrolein as the primary product in the one-to-one mapping. Acetaldehyde oxidation by aldehyde oxidase in step 6 yields acetate. However, it is followed by the acetoacetyl-CoA: acetate CoA-transferase reaction, (i.e., step 7 in Figure 5) which deviates from the primary reactant’s carbon path to produce the coproduct acetyl-CoA. Therefore, a cautionary tale with the use of minChemBio is that transferase reactions in MetaNetX database can propagate in the results by arriving at solutions that do not follow the main carbon path.

Figure 5.

Figure 5

Two pathways identified for the target molecule 3HB by minChemBio using glycerol as the starting molecule. Step 1 is a glycerol hydrolase dehydratase reaction (MNXR100324). Step 2 is 3-hydroxyproponal dehydratase reaction (MNXR177707). Step 3 is the 3-hydroxypoponoate:acetoacetate CoA transferase reaction (MNXR130270). Step 4 is 3-hydroxybutyrate dehydrogenase reaction (MNXR96232). Step 5 is chemical dehydration reaction producing acetaldehyde and acrolein. Step 6 is an aldehyde oxidase reaction (MNXR122134), and step 7 is an acetoacetyl-CoA:acetate CoA-transferase reaction (MNXR190492).

Discussion

Using minChemBio, we demonstrated several example synthesis challenges and showed that it can propose established as well as unexplored chemo-enzymatic pathways using USPTO chemical reactions24 and MetaNetX biological reactions. minChemBio extracts one-to-one mappings from chemical and biological reactions while allowing for the reversible use of enzymatic reactions whenever the thermodynamics is favorable (i.e., ΔG negative or near zero). It significantly streamlines a process that traditionally requires hours of literature search, enabling the design of candidate pathways using known reactions between source and target chemicals within minutes. By minimizing transitions, the tool suggests pathways with the fewest separation and purification steps and thus prioritizing fully enzymatic pathways and successively including chemical steps as alternate solutions. We analyzed the thermodynamics of the enzymatic reactions in the solution to compare the alternative pathways identified by minChemBio, as the directionality of reactions is not clearly specified in the MetaNetX database. Since USPTO chemical reactions are assumed to occur only in the forward direction, the Gibbs energy change (ΔG) was not considered in this analysis. We contrasted minChemBio solutions with existing retrosynthetic tools such as ASKCOS5 and RetroBioCat10 that use reaction templates but do not prioritize known reactions. Since, they do not allow for control over the starting molecule, these tools can suggest pathways starting with costlier molecule as compared to the target whereas in minChemBio the user can specify the starting molecule.

There can be potential limitations to the solutions provided by minChemBio. The current implementation cannot distinguish between major and minor products in chemical reactions due to a lack of reaction elemental balancing information in the USPTO reaction SMILES strings. Furthermore, considering all combinations of one-to-one mappings in transferase reactions in MetaNetX database can sometimes result in pathways which deviate from the main carbon path of the primary reactant, and instead follow a cosubstrate’s carbon path. While an alternative data set (USPTO-MIT)51 with elementally balanced entries exists, its lack of associated patent identifiers limits its utility.

minChemBio only considers known reactions with no opportunities for proposing novel steps as in novoStoic.13 This can be performed a posteriori either by sequence-to-sequence translational models52 or by extracting templates53 from the USPTO database. As evident from the results obtained, relying on reaction SMILES strings of USPTO database does not guarantee feasible chemo-enzymatic pathways due to the presence of incorrect reaction SMILES annotations. As shown in Figure 3a, upon manually inspecting the chemical reaction (step 3), we found that the reactants and products are incorrectly annotated in the reaction SMILES (see Figure 6).54 Therefore, there is a need to eliminate incorrect reaction SMILES strings from the data set. Herein, the use of Large Language Models (LLM) to query the patent literature text and validate the reaction SMILES present in the USPTO database is of relevance. This could enable the elimination of incorrectly annotated reactions and avoid spurious pathway designs.

Figure 6.

Figure 6

Incorrect annotations in the USPTO database result in erroneous one-to-one mappings. In this case, the mapping shows acetic acid transforming into acetoacetate. However, a review of the patent literature reveals that this reaction does not occur, indicating a misannotation. Consequently, the extracted one-to-one mapping is also incorrect. Such inconsistencies can be resolved by cross-referencing with the original patent literature.

Methods

A detailed description of minChemBio is presented in this section. First, we present the streamlit platform webpage, that shows how inputs are given to generate solutions. Then, we describe the preprocessing steps taken in the USPTO chemical reaction database and MetaNetX biological reaction database to form the network of known reactions. Finally, we illustrate the mathematical formulation of the minChemBio algorithm.

Web Platform: minChemBio

We begin by identifying the chemical identifiers (IDs) of both the source and target molecules. If a molecule is present in the MetaNetX database, its MetaNetX ID is used. Otherwise, the chemical ID (see File S1.txt in 10.26207/tbg0-gr88) is determined using the molecule’s SMILES string. The users can input SMILES strings of source and target molecule to obtain the respective molecule IDs.

For the following example, p-Xylene (MetaNetX ID: MNXM3685) is used as the starting molecule, and terephthalate (MetaNetX ID: MNXM2734) is the target molecule as shown in Figure 7. A mixed-integer linear programming (MILP) formulation is solved, yielding 23 solutions within 2 h. the Reaction IDs (see File S2.csv in 10.26207/tbg0-gr88) and Chemical IDs (see File S1.txt in 10.26207/tbg0-gr88) and their corresponding SMILES strings, obtained from the USPTO database, which participate in the chemoenzymatic pathways. The biological reactions within the pathways are associated with their respective MetaNetX IDs.

Figure 7.

Figure 7

Streamlit platform for minChemBio. Two inputs are required: starting molecule ID in place of “reactant” and target molecule ID in place of “product”.

USPTO Chemical Reaction Data Set

The initial data set consisted of around 1,808,938 chemical reactions curated from the USPTO chemical reaction database.24 There are 1,271,398 unique molecules used to describe these reactions. Compared to commercially available options such as Pistachio55 and Reaxys,8 the USPTO data set is publicly available but suffers from inherent limitations, such as lack of stoichiometry information, imbalanced reactions, duplicate reactions, erroneous or missing reactants/products, and often misidentified reagents/catalysts.24 The reaction data consists of SMILES strings (see Figure 8a) which include the reactant(s), the agent(s), and the product(s). Agents are molecules that do not contribute atoms to the product or accept atoms from the reactants. They are commonly used to denote catalysts, solvents, and other indirect participants to the reaction. To address these issues, the data set was preprocessed. Each reaction was assigned a unique identifier starting with “RXN” followed by 8 digits (see Supporting Information). Similarly, each chemical was assigned a unique identifier starting with “CHEM” followed by 8 digits (see Supporting Information). The first stage of preprocessing consisted of three steps. Reactions with the same SMILES string were identified using string matching in python and duplicate reactions were removed from the database. Only about 20,000 reactions were removed this way. Therefore, a tanimoto similarity between reactants of different reactions and between products of different reactions was found to remove duplicate reactions, which led to 512,644 duplicate reactions being removed. Therefore, we get a list of reactions each with a reaction id where reactant, product and agent are available in form of molecule ids. Since, some molecules in chemical reaction data set overlap in the biological reaction data set. The next step involved replacing all the CHEM ids with MNXM ids wherever a molecule was present in both data sets. Subsequently, reactions involving only such molecules with less than or equal to one carbon atom were excluded, as our focus lies on synthesis planning for hydrocarbons. These steps resulted in a set of unique chemical reactions with at least one reactant and one product having more than one carbon atom.

Figure 8.

Figure 8

Preprocessing of the USPTO data set includes removing duplicates, assigning a unique identifier to chemicals overlapping in both databases, and removing reactions that do not contain a hydrocarbon in either the reactant set or the product set. USPTO data set consists of 1.8 million SMILES reaction strings and 1.2 million molecules as part of these reaction strings. The initial preprocessing includes the dollowing: (i) Removing duplicate reactions: first by string matching, and finally by comparing all the reactants and products of every reaction with every other reaction. At the end of this step, ∼500 000 reactions are removed. (ii) Assigning a unique identifier to a molecule appearing in both USPTO and MetaNetX databases: This is achieved by calculating the Tanimoto coefficient of every pair of molecules (one taken from the USPTO database, one taken from MetaNetX). This identifies 13 532 molecules which overlap in both the databases. (iii) Removing reactions that do not have hydrocarbon molecules either in the reactant set or in the product set: If the highest number of carbons in any reactant or product or molecule was ≤1, then we removed that reaction. This removed 3480 reactions.

Previous works56,57 have included only those reactions which balanced all the carbon atoms in the reactants and the products. In this process, about 75% of the total reactions in USPTO data set are ignored, but as shown in Figure 9, we need to identify the inconsistencies of the reaction to be able to use them in our reaction network.

Figure 9.

Figure 9

One-to-one mapping tracks the main carbon path by identifying the primary reactant and product by selecting the pair with highest Tanimoto similarity coefficient. (a) A reaction RXN00300525 where N–N dimethylaniline is listed as one of reactants, but it is not contributing any atoms to the product. Therefore, it is not one of the coreactants. Similarly, in reaction RXN00605309, di-isopropylethylamaine is not one of the coreactants, since it does not contribute any atom to the product. We identify the reactant that is most similar to the product in both cases using the Tanimoto coefficient of the pair. This transformation effectively represents the chemical reaction, which we call one-to-one mapping.

Due to the often-missing coreactants and/or coproducts, it is not possible to balance the stoichiometries of molecules involved in the reaction. Consequently, one-to-one reaction mappings were extracted that trace the main carbon path.56 These represent the primary reactant and product for each reaction as shown in Figure 9. This description captures the core transformation without introducing errors from an attempt to fix stoichiometric imbalances. By tracking the main carbon path, a MILP (Mixed Integer Linear Programming) optimization formulation is put forth to find a connected path from the starting precursor to the target molecule. The identification of the primary reactant(s) and product(s) for each reaction involves the following steps. First, all carbon-containing reactants and products were extracted, and SMILES string canonicalization was employed to match entries from a biological reaction database enabling the assignment of identical chemical identifiers wherever applicable. Canonicalization of SMILES strings allows for comparisons between any two molecules by simply matching the characters of their respective SMILES strings. Canonicalization is first required so that the order of characters in the SMILES strings follows the same parsing convention. The reactant-product pair exhibiting the highest Tanimoto similarity coefficient, calculated using the RDKit python library,25 is flagged as the reactant-product pair which traces the main carbon path. The identified primary reactant-product pair is checked against the possibility that they are solvents or catalysts if they are present on both sides of the reaction. In this case, the pair with the second highest Tanimoto similarity coefficient is chosen as the main carbon path as shown in Figure 10. This data preprocessing and simplification process yielded a final data set of 1,135,706 one-to-one reaction mappings.

Figure 10.

Figure 10

Incorrect one-to-one mapping using the Tanimoto coefficient is obtained when an agent is misannotated as a coreactant or coproduct. Here, acetic acid is part of the product set. When acetic acid is removed from the set of products, as it is also part of set of agents, the new primary reactant of the one-to-one mapping is 5-tert-butyl-m-xylene, and the main product of the one-to-one mapping is 4-tert-butyl-2,6-dimethylmandelic acid.

MetaNetX Biological Reaction Data Set

A one-to-one mapping approach between main carbon path reactants and products was similarly applied to the biological reactions, sourced from the MetaNetX database,12 which aggregates biochemical reactions and metabolite information from various sources such as KEGG,12 MetaCyc,58 RHEA,27,29 and ChEBI.27 It encompasses 57,541 reactions and nearly 1 million metabolites. However, only a fraction of these metabolites is involved in the reactions. Biological reactions commonly utilize cofactors, which are typically large carbon-containing metabolites generated during metabolic processes and act as essential part of an active site. A curated list of 47 cofactors, including small molecules such as CO2, H2O, and H+, was compiled and removed from the reactions before creating one-to-one mappings.

From the initial set of 57,541 reactions, a set of 13,306 unique reactions was generated after removing duplicates and generic reactions. SMILES strings representing metabolites were canonicalized using the RDKit python library.25 It is important to note that the biological reaction data set represents a smaller subset compared to the chemical reaction data set. Moreover, biological reactions often involve multiple primary reactants and products, particularly in the case of transferase enzymes. It would be reductive to select one pair of main reactant and product. As a result, all possible combinations were considered, and a pair was discarded only if the Tanimoto similarity coefficient was less than 0.1, and the difference in the number of carbons between the reactant and product was more than 90% of the smaller molecule’s carbon count (see Figure 11). Considering the reversible nature of biological reactions, one-to-one mappings were extracted for both the forward and reverse directions. This resulted in a final list of 67,622 unique one-to-one mappings for biological reactions. To avoid thermodynamic inconsistencies in the pathway, pairs of opposite one-to-one mappings from the same biological reactions were flagged to ensure they do not appear together in the same pathway. For instance, in Figure 11, four such pairs of opposite one-to-one mappings can be identified: (1,2), (1,4), (3,2) and (3,4).

Figure 11.

Figure 11

Generated four forward one-to-one mappings from a transferase reaction in the MetaNetX biological reaction database. (a) Oximinotransferase reaction. (b) In this transferase reactions, te Tanimoto similarity coefficients of any two pairs of reactant and product differ by <10%, and therefore, all four possible combinations (in the gray box) of one-to-one mappings are extracted. Along with forward mappings, reverse mappings are also extracted, and unique numbers of mappings are included in the final data set.

Transitions

The objective is to identify reaction pathways with minimum transitions converting a starting precursor to a target molecule including both chemical and biological reactions. When a chemical reaction is succeeded by a biological reaction in a pathway and vice versa, separation and purification steps are required before moving on to the next reaction. Similarly, even when a chemical reaction is succeeded by a chemical reaction, change in reaction conditions require additional separation and purification steps. Fewer transitions will avoid these additional steps and reduce the cost and complexity of the overall chemo-enzymatic pathway.

Mathematical Formulation

Sets.

graphic file with name sb4c00692_m001.jpg
graphic file with name sb4c00692_m002.jpg
graphic file with name sb4c00692_m003.jpg

Parameters.

graphic file with name sb4c00692_m004.jpg
graphic file with name sb4c00692_m005.jpg

Rij represents a one-to-one mapping extracted from each reaction in either of the databases. It is defined for every chemical i present in every one-to-one mapping j (see Figure 12). Exchange one-to-one mappings are added on top of the extracted one-to-one mappings to indicate the source and sink(target) of the flux in the pathway.

Figure 12.

Figure 12

A representation of the chemo-enzymatic reaction network and identification of transitions along pathways for target molecule from starting molecule a. Each chemical reaction results in a one one-to-one mapping, whereas every biological reaction provides at least a forward and a backward one-to-one mapping since biological reactions in the database do have the information about directionality. For example: Two one-to-one connections are generated from biological reaction between chemical e and chemical i. Mapping 12 signifies a connection from i to e, whereas mapping −12 signifies a connection from e to i. When the target chemical “c” is chosen and the starting chemical “a” is chosen, then five pathways are identified. For every active one-to-one mapping j, vj = 1, and for every biological to chemical transition at the intermediate i, yi = 1. For every chemical-to-chemical transition at intermediate i, zi = 1.

Variables:

graphic file with name sb4c00692_m006.jpg

vj is flux of the one-to-one mapping j. The flux vj is nonzero if the one-to-one mapping j is present in the pathway.

graphic file with name sb4c00692_m007.jpg
graphic file with name sb4c00692_m008.jpg

ybi and yci are defined to enable the identification of a biological to chemical transition, i.e. when a chemical intermediate i is produced by biological one-to-one mapping and consumed by a chemical one-to-one mapping. ybi is nonzero when a chemical intermediate i is produced by a biological reaction and yci is nonzero when a chemical intermediate i is consumed by a chemical reaction. When ybi and yci are both nonzero simultaneously, it indicates that there is a biological to chemical transition at the position of the chemical intermediate i.

Similarly, xbi and xci are defined to identify a transition but a chemical to biological one. To identify such a transition, we need to ascertain that a chemical i is produced using a chemical one-to-one mapping and is consumed using a biological one-to-one mapping. xbi is nonzero when a chemical intermediate i is consumed by a biological reaction and xci is nonzero when a chemical intermediate i is produced by a chemical reaction. Therefore, a chemical to biological transition requires both xbi and xci to be nonzero simultaneously at a chemical intermediate i.

graphic file with name sb4c00692_m009.jpg
graphic file with name sb4c00692_m010.jpg
graphic file with name sb4c00692_m011.jpg

yi is the product of the variables of ybi and yci. If the latter variables are simultaneously nonzero, then yi becomes nonzero at the position of chemical intermediate i (see Figure 12). Therefore, yi identifies if there is a chemical to biological transition present in the pathway. xi is the product of the variables of xbi and xci. If the latter variables are simultaneously nonzero, then xi becomes nonzero at the position of chemical intermediate i (see Figure 12).

Whenever an intermediate is produced by a chemical one-to-one mapping in the pathway and consumed by a chemical one-to-one in the pathway, there exists a chemical-to-chemical transition. Two successive chemical one-to-one mappings in the pathway require additional separation and a purification step(s) in between. This transition is identified by the variable zi, which becomes nonzero for a chemical intermediate i, if yci and xci are both nonzero. Using the above variables, we can find out if the intermediate molecules of the pathway are results of any transition. Total number of transitions is the sum of xi, yi and zi over the set of chemicals i.

Objective Function. The expression of objective function minimizes the total number of transitions present in the pathway. The total of number of transitions is found out using the sum of the all the transitions associated with all the intermediates.

graphic file with name sb4c00692_m012.jpg

Constraints.

graphic file with name sb4c00692_m013.jpg 1

Constraint 1 satisfies the molecule balance of the pathway, which signifies that every intermediate produced must be consumed. Here, exchange one-to-one mappings are added on top of extracted one-to-one mappings so that the R.H.S. of the equation goes to 0. Exchange one-to-one mappings are half mappings for both primary reactant and primary product, they are present to ensure molecule balance equation adds up to 0. The fluxes of these exchange one-to-one mappings are nonzero to reflect that the source and sink reactions are active in the overall pathway.

graphic file with name sb4c00692_m014.jpg 2
graphic file with name sb4c00692_m015.jpg 3

Constraints 4 and 5 determine the value of variables ybi and yci. If j is a biological one-to-one mapping (Cj = 0) and chemical i is the product for one-to-one mapping j (Rij > 0) and j is active (vj = 1), then chemical intermediate i is produced through a biological one-to-one mapping (extracted from a biological reaction) in the pathway. An intermediate chemical i can be a product only once in the pathway, the maximum value of this sum is 1, and the minimum value is 0. This sum is bound by the binary variable ybi. A similar constraint is presented in 5, where for every chemical intermediate i present in the pathway, a sum is defined which determines whether the intermediate molecule i is consumed through a chemical one-to-one mapping(extracted from a chemical reaction). This sum is bound by the binary variable yci. Constraints 6, 7 and 8 are linear constraints used to denote that binary variable yi is the product of binary variables, ybi and yci. Whenever, yi is nonzero, there exists a biological to chemical transition at intermediate chemical i.

graphic file with name sb4c00692_m016.jpg 4
graphic file with name sb4c00692_m017.jpg 5
graphic file with name sb4c00692_m018.jpg 6
graphic file with name sb4c00692_m019.jpg 7
graphic file with name sb4c00692_m020.jpg 8

To identify a chemical to biological transition at chemical intermediate i, we would need to check whether intermediate chemical i is product of chemical one-to-one mapping and a reactant of a biological one-to-one mapping which are active in the pathway. xbi and xci are the binary variables used to determine those, respectively. xi is the product of xbi and xci expressed in terms of linear constraints. Constraints 9 through 13 determine the value of these variables.

graphic file with name sb4c00692_m021.jpg 9
graphic file with name sb4c00692_m022.jpg 10
graphic file with name sb4c00692_m023.jpg 11
graphic file with name sb4c00692_m024.jpg 12
graphic file with name sb4c00692_m025.jpg 13

The following constraint is allowing the total number of transitions in the pathway. The number of transitions is sum of xi and yi over all chemicals i. If t is the number of allowed transitions, then,

graphic file with name sb4c00692_m026.jpg 14

Further, all the forward and reverse mappings of biological reactions are extracted, and these pairs should not simultaneously appear into the pathway to maintain thermodynamic feasibility. Constraint 15 avoids all such pairs of one-to-one mappings which are opposite of each other and extracted from the same biological rection. Constraint 16 ensures that all variables used in the MILP formulation are binary.

graphic file with name sb4c00692_m027.jpg 15
graphic file with name sb4c00692_m028.jpg 16

Alternate Optimal Solutions

To identify alternate optimal and suboptimal solutions, an earlier solution k (set of positive fluxes) is prevented from reappearing. Constraint 17 removes every previous optimal solution by removing the all the one-to-one mappings active together in previous pathways. Often, suboptimal solutions can give a better pathway, because suboptimality means more steps involved in the overall pathway, and if all the steps are thermodynamically feasible and do not require cosubstrates in biological steps then they are worth considering.

graphic file with name sb4c00692_m029.jpg 17

Acknowledgments

This work is supported by the U.S. National Science Foundation funded Molecule Maker Lab Institute (MMLI), award number 2019897, supported by National AI Research Institutes Program of the Directorate for Computer and Information Science and Engineering (CISE), in collaboration with the Division of Chemistry (CHE) and the Division of Chemical, Bioengineering, and Environmental Transport Systems (CBET) awarded to C.D.M. The funders had no role in study design, data collection and analysis, decision to publish, or preparation of the manuscript

Data Availability Statement

minChemBio is available https://github.com/maranasgroup/chemo-enz under an MIT license. All the codes and data are also available at 10.26207/tbg0-gr88.

Supporting Information Available

The Supporting Information is available free of charge at https://pubs.acs.org/doi/10.1021/acssynbio.4c00692.

  • Figures of each individual pathway solution found by minChemBio for the four case studies discussed (PDF)

Author Contributions

M.A.: Conceptualization, data curation, formal analysis, investigation, methodology, software, writing—original draft, visualization. V.U.: Conceptualization, data curation, writing—review and editing. C.D.M.: supervision, funding acquisition, writing—review and editing.

The authors declare no competing financial interest.

Supplementary Material

sb4c00692_si_001.pdf (2.9MB, pdf)

References

  1. Romero E. O.; Perkins J. C.; Burch J. E.; Delgadillo D. A.; Nelson H. M.; Narayan A. R. H. Chemoenzymatic Synthesis of (+)-Xyloketal B. Org. Lett. 2023, 25, 1547–1552. 10.1021/acs.orglett.3c00334. [DOI] [PubMed] [Google Scholar]
  2. Sirirungruang S.; Ad O.; Privalsky T. M.; Ramesh S.; Sax J. L.; Dong H.; Baidoo E. E. K.; Amer B.; Khosla C.; Chang M. C. Y. Engineering site-selective incorporation of fluorine into polyketides. Nat. Chem. Biol. 2022, 18, 886–893. 10.1038/s41589-022-01070-y. [DOI] [PMC free article] [PubMed] [Google Scholar]
  3. Tu Z.; Stuyver T.; Coley C. W. Predictive chemistry: machine learning for reaction deployment, reaction development, and reaction discovery. Chem. Sci. 2023, 14, 226–244. 10.1039/D2SC05089G. [DOI] [PMC free article] [PubMed] [Google Scholar]
  4. Klucznik T.; Mikulak-Klucznik B.; McCormack M. P.; Lima H.; Szymkuć S.; Bhowmick M.; Molga K.; Zhou Y.; Rickershauser L.; Gajewska E. P.; Toutchkine A.; Dittwald P.; Startek M. P.; Kirkovits G. J.; Roszak R.; Adamski A.; Sieredzińska B.; Mrksich M.; Trice S. L. J.; Grzybowski B. A. Efficient Syntheses of Diverse, Medicinally Relevant Targets Planned by Computer and Executed in the Laboratory. Chem. 2018, 4, 522–532. 10.1016/j.chempr.2018.02.002. [DOI] [Google Scholar]
  5. Coley C. W.; Thomas D. A.; Lummiss J. A. M.; Jaworski J. N.; Breen C. P.; Schultz V.; Hart T.; Fishman J. S.; Rogers L.; Gao H.; Hicklin R. W.; Plehiers P. P.; Byington J.; Piotti J. S.; Green W. H.; Hart A. J.; Jamison T. F.; Jensen K. F. A robotic platform for flow synthesis of organic compounds informed by AI planning. Science (1979) 2019, 365, eaax1566 10.1126/science.aax1566. [DOI] [PubMed] [Google Scholar]
  6. Levin I.; Liu M.; Voigt C. A.; Coley C. W. Merging enzymatic and synthetic chemistry with computational synthesis planning. Nat. Commun. 2022, 13, 7747. 10.1038/s41467-022-35422-y. [DOI] [PMC free article] [PubMed] [Google Scholar]
  7. Lang M.; Stelzer M.; Schomburg D. BKM-react, an integrated biochemical reaction database. BMC Biochem 2011, 12, 42. 10.1186/1471-2091-12-42. [DOI] [PMC free article] [PubMed] [Google Scholar]
  8. Reaxys. https://www.reaxys.com.
  9. Sankaranarayanan K.; Jensen K. F. Computer-assisted multistep chemoenzymatic retrosynthesis using a chemical synthesis planner. Chem. Sci. 2023, 14, 6467–6475. 10.1039/D3SC01355C. [DOI] [PMC free article] [PubMed] [Google Scholar]
  10. Finnigan W.; Hepworth L. J.; Flitsch S. L.; Turner N. J. RetroBioCat as a computer-aided synthesis planning tool for biocatalytic reactions and cascades. Nat. Catal 2021, 4, 98–104. 10.1038/s41929-020-00556-z. [DOI] [PMC free article] [PubMed] [Google Scholar]
  11. Chowdhury A.; Maranas C. D. Designing overall stoichiometric conversions and intervening metabolic reactions. Sci. Rep 2015, 5, 16009 10.1038/srep16009. [DOI] [PMC free article] [PubMed] [Google Scholar]
  12. Kanehisa M.; Goto S. KEGG: Kyoto Encyclopedia of Genes and Genomes. Nucleic Acids Res. 2000, 28, 27–30. 10.1093/nar/28.1.27. [DOI] [PMC free article] [PubMed] [Google Scholar]
  13. Kumar A.; Wang L.; Ng C. Y.; Maranas C. D. Pathway design using de novo steps through uncharted biochemical spaces. Nat. Commun. 2018, 9, 184. 10.1038/s41467-017-02362-x. [DOI] [PMC free article] [PubMed] [Google Scholar]
  14. Wang L.; Upadhyay V.; Maranas C. D. dGPredictor: Automated fragmentation method for metabolic reaction free energy prediction and de novo pathway design. PLoS Comput. Biol. 2021, 17, e1009448. 10.1371/journal.pcbi.1009448. [DOI] [PMC free article] [PubMed] [Google Scholar]
  15. Mu R., Wang Z., Wamsley M. C., Duke C. N., Lii P. H., Epley S. E., Todd L. C., Roberts P. J. (2020) Application of Enzymes in Regioselective and Stereoselective Organic Reactions. Catalysts 10.832. 10.3390/catal10080832 [DOI] [Google Scholar]
  16. Siddiqui K. S., Ertan H., Poljak A., Bridge W. J. (2022) Evaluating Enzymatic Productivity—The Missing Link to Enzyme Utility. Int. J. Mol. Sci. 23.6908. 10.3390/ijms23136908 [DOI] [PMC free article] [PubMed] [Google Scholar]
  17. France S. P.; Hepworth L. J.; Turner N. J.; Flitsch S. L. Constructing Biocatalytic Cascades: In Vitro and in Vivo Approaches to de Novo Multi-Enzyme Pathways. ACS Catal. 2017, 7, 710–724. 10.1021/acscatal.6b02979. [DOI] [Google Scholar]
  18. Zetzsche L. E.; Yazarians J. A.; Chakrabarty S.; Hinze M. E.; Murray L. A. M.; Lukowski A. L.; Joyce L. A.; Narayan A. R. H. Biocatalytic oxidative cross-coupling reactions for biaryl bond formation. Nature 2022, 603, 79–85. 10.1038/s41586-021-04365-7. [DOI] [PMC free article] [PubMed] [Google Scholar]
  19. Ruck R. T.; Strotman N. A.; Krska S. W. The Catalysis Laboratory at Merck: 20 Years of Catalyzing Innovation. ACS Catal. 2023, 13, 475–503. 10.1021/acscatal.2c05159. [DOI] [Google Scholar]
  20. Sperl J. M.; Sieber V. Multienzyme Cascade Reactions—Status and Recent Advances. ACS Catal. 2018, 8, 2385–2396. 10.1021/acscatal.7b03440. [DOI] [Google Scholar]
  21. Wang Z.; Sundara Sekar B.; Li Z. Recent advances in artificial enzyme cascades for the production of value-added chemicals. Bioresour. Technol. 2021, 323, 124551 10.1016/j.biortech.2020.124551. [DOI] [PubMed] [Google Scholar]
  22. Chainani Y.; Bonnanzio G.; Tyo K. E. J.; Broadbelt L. J. Coupling chemistry and biology for the synthesis of advanced bioproducts. Curr. Opin Biotechnol 2023, 84, 102992 10.1016/j.copbio.2023.102992. [DOI] [PubMed] [Google Scholar]
  23. Moretti S.; Tran V. D. T.; Mehl F.; Ibberson M.; Pagni M. MetaNetX/MNXref: unified namespace for metabolites and biochemical reactions in the context of metabolic models. Nucleic Acids Res. 2021, 49, D570–D574. 10.1093/nar/gkaa992. [DOI] [PMC free article] [PubMed] [Google Scholar]
  24. Lowe D.Chemical reactions from US patents (1976–2016). Dataset, 2017.
  25. RDKit: Open-Source cheminformatics. https://www.rdkit.org.
  26. Norsigian C. J.; Pusarla N.; McConn J. L.; Yurkovich J. T.; Dräger A.; Palsson B. O.; King Z. BiGG Models 2020: multi-strain genome-scale models and expansion across the phylogenetic tree. Nucleic Acids Res. 2019, 48, D402–D406. 10.1093/nar/gkz1054. [DOI] [PMC free article] [PubMed] [Google Scholar]
  27. Hastings J.; de Matos P.; Dekker A.; Ennis M.; Harsha B.; Kale N.; Muthukrishnan V.; Owen G.; Turner S.; Williams M.; Steinbeck C. The ChEBI reference database and ontology for biologically relevant chemistry: enhancements for 2013. Nucleic Acids Res. 2012, 41, D456–D463. 10.1093/nar/gks1146. [DOI] [PMC free article] [PubMed] [Google Scholar]
  28. Seaver S. M. D.; Liu F.; Zhang Q.; Jeffryes J.; Faria J. P.; Edirisinghe J. N.; Mundy M.; Chia N.; Noor E.; Beber M. E.; Best A. A.; DeJongh M.; Kimbrel J. A.; D’haeseleer P.; McCorkle S. R.; Bolton J. R.; Pearson E.; Canon S.; Wood-Charlson E. M.; Cottingham R. W.; Arkin A. P.; Henry C. S. The ModelSEED Biochemistry Database for the integration of metabolic annotations and the reconstruction, comparison and analysis of metabolic models for plants, fungi and microbes. Nucleic Acids Res. 2021, 49, D575–D588. 10.1093/nar/gkaa746. [DOI] [PMC free article] [PubMed] [Google Scholar]
  29. Bansal P.; Morgat A.; Axelsen K. B.; Muthukrishnan V.; Coudert E.; Aimo L.; Hyka-Nouspikel N.; Gasteiger E.; Kerhornou A.; Neto T. B.; Pozzato M.; Blatter M.-C.; Ignatchenko A.; Redaschi N.; Bridge A. Rhea, the reaction knowledgebase in 2022. Nucleic Acids Res. 2022, 50, D693–D700. 10.1093/nar/gkab1016. [DOI] [PMC free article] [PubMed] [Google Scholar]
  30. Wittig U.; Rey M.; Weidemann A.; Kania R.; Müller W. SABIO-RK: an updated resource for manually curated biochemical reaction kinetics. Nucleic Acids Res. 2018, 46, D656–D660. 10.1093/nar/gkx1065. [DOI] [PMC free article] [PubMed] [Google Scholar]
  31. Probst D.; Manica M.; Nana Teukam Y. G.; Castrogiovanni A.; Paratore F.; Laino T. Biocatalysed synthesis planning using data-driven learning. Nat. Commun. 2022, 13, 964. 10.1038/s41467-022-28536-w. [DOI] [PMC free article] [PubMed] [Google Scholar]
  32. Koch M.; Duigou T.; Faulon J.-L. Reinforcement Learning for Bioretrosynthesis. ACS Synth. Biol. 2020, 9, 157–168. 10.1021/acssynbio.9b00447. [DOI] [PubMed] [Google Scholar]
  33. Schwaller P.; Petraglia R.; Zullo V.; Nair V. H.; Haeuselmann R. A.; Pisoni R.; Bekas C.; Iuliano A.; Laino T. Predicting retrosynthetic pathways using transformer-based models and a hyper-graph exploration strategy. Chem. Sci. 2020, 11, 3316–3325. 10.1039/C9SC05704H. [DOI] [PMC free article] [PubMed] [Google Scholar]
  34. Wang K.-F.; Liu C.; Sui K.; Guo C.; Liu C.-Z. Efficient Catalytic Oxidation of 5-Hydroxymethylfurfural to 2,5-Furandicarboxylic Acid by Magnetic Laccase Catalyst. ChemBioChem. 2018, 19, 654–659. 10.1002/cbic.201800008. [DOI] [PubMed] [Google Scholar]
  35. Yang C.-F.; Huang C.-R. Biotransformation of 5-hydroxy-methylfurfural into 2,5-furan-dicarboxylic acid by bacterial isolate using thermal acid algal hydrolysate. Bioresour. Technol. 2016, 214, 311–318. 10.1016/j.biortech.2016.04.122. [DOI] [PubMed] [Google Scholar]
  36. McKenna S. M.; Leimkühler S.; Herter S.; Turner N. J.; Carnell A. J. Enzyme cascade reactions: synthesis of furandicarboxylic acid (FDCA) and carboxylic acids using oxidases in tandem. Green Chem. 2015, 17, 3271–3275. 10.1039/C5GC00707K. [DOI] [Google Scholar]
  37. Dijkman W. P.; Groothuis D. E.; Fraaije M. W. Enzyme-Catalyzed Oxidation of 5-Hydroxymethylfurfural to Furan-2,5-dicarboxylic Acid. Angew. Chem., Int. Ed. 2014, 53, 6515–6518. 10.1002/anie.201402904. [DOI] [PubMed] [Google Scholar]
  38. Ku C. J.; Jaewon J.; Sangyong K.; Bora K.; Baek-Jin K.; Seunghan S.; Dohoon L.. Method for producing 5-hydroxymethyl-2-furfural from maize syrup containing fructose. US9206148B2, 2015.
  39. Ghiaci M.; Mostajeran M.; Gil A. Synthesis and Characterization of Co–Mn Nanoparticles Immobilized on a Modified Bentonite and Its Application for Oxidation of p-Xylene to Terephthalic Acid. Ind. Eng. Chem. Res. 2012, 51, 15821–15831. 10.1021/ie3021939. [DOI] [Google Scholar]
  40. Luo Z. W.; Lee S. Y. Biotransformation of p-xylene into terephthalic acid by engineered Escherichia coli. Nat. Commun. 2017, 8, 15689 10.1038/ncomms15689. [DOI] [PMC free article] [PubMed] [Google Scholar]
  41. Bramucci M. G.; Nagarajan V.; Thomas S. M.. Use of xylene monooxygenase for the oxidation of substituted monocyclic aromatic compounds. US2003/0073206A1, 2003.
  42. Takashi S., Akira T., Susumu N.. Process for producing terephthalic acid. US 4245078 A, 1981.
  43. Kollar J.Process for the industrial production of ethylene oxide and aromatic acid. US4046782A, 1977.
  44. Kim S.; Chen J.; Cheng T.; Gindulyte A.; He J.; He S.; Li Q.; Shoemaker B. A.; Thiessen P. A.; Yu B.; Zaslavsky L.; Zhang J.; Bolton E. E. PubChem. 2023 update. Nucleic Acids Res. 2023, 51, D1373–D1380. 10.1093/nar/gkac956. [DOI] [PMC free article] [PubMed] [Google Scholar]
  45. Ahn W. S.; Park S. J.; Lee S. Y. Production of Poly(3-Hydroxybutyrate) by Fed-Batch Culture of Recombinant Escherichia coliwith a Highly Concentrated Whey Solution. Appl. Environ. Microbiol. 2000, 66, 3624–3627. 10.1128/AEM.66.8.3624-3627.2000. [DOI] [PMC free article] [PubMed] [Google Scholar]
  46. Lin Z.; Zhang Y.; Yuan Q.; Liu Q.; Li Y.; Wang Z.; Ma H.; Chen T.; Zhao X. Metabolic engineering of Escherichia coli for poly(3-hydroxybutyrate) production viathreonine bypass. Microb Cell Fact 2015, 14, 185. 10.1186/s12934-015-0369-3. [DOI] [PMC free article] [PubMed] [Google Scholar]
  47. Tokiwa Y.; Ugwu C. U. Biotechnological production of (R)-3-hydroxybutyric acid monomer. J. Biotechnol. 2007, 132, 264–272. 10.1016/j.jbiotec.2007.03.015. [DOI] [PubMed] [Google Scholar]
  48. Amabile C.; Abate T.; De Crescenzo C.; Sabbarese S.; Muñoz R.; Chianese S.; Musmarra D. Sustainable Process for the Production of Poly(3-hydroxybutyrate-co-3-hydroxyvalerate) from Renewable Resources: A Simulation Study. ACS Sustain Chem. Eng. 2022, 10, 14230–14239. 10.1021/acssuschemeng.2c04111. [DOI] [PMC free article] [PubMed] [Google Scholar]
  49. Kim B. S. Production of poly(3-hydroxybutyrate) from inexpensive substrates. Enzyme Microb Technol. 2000, 27, 774–777. 10.1016/S0141-0229(00)00299-4. [DOI] [PubMed] [Google Scholar]
  50. Mierziak J.; Burgberger M.; Wojtasik W. 3-Hydroxybutyrate as a Metabolite and a Signal Molecule Regulating Processes of Living Organisms. Biomolecules 2021, 11, 402. 10.3390/biom11030402. [DOI] [PMC free article] [PubMed] [Google Scholar]
  51. Jin W., Coley C., Barzilay R., Jaakkola T.. Predicting Organic Reaction Outcomes with Weisfeiler-Lehman Network. In Advances in Neural Information Processing Systems; Guyon I., Von Luxburg U., Bengio S., Wallach H., Fergus R., Vishwanathan S., Garnett R., Eds.; Curran Associates, Inc, 2017. [Google Scholar]
  52. Liu B.; Ramsundar B.; Kawthekar P.; Shi J.; Gomes J.; Luu Nguyen Q.; Ho S.; Sloane J.; Wender P.; Pande V. Retrosynthetic Reaction Prediction Using Neural Sequence-to-Sequence Models. ACS Cent Sci. 2017, 3, 1103–1113. 10.1021/acscentsci.7b00303. [DOI] [PMC free article] [PubMed] [Google Scholar]
  53. Coley C. W.; Green W. H.; Jensen K. F. RDChiral: An RDKit Wrapper for Handling Stereochemistry in Retrosynthetic Template Extraction and Application. J. Chem. Inf Model 2019, 59, 2529–2537. 10.1021/acs.jcim.9b00286. [DOI] [PubMed] [Google Scholar]
  54. Knobelsdorf J. A., Shepherd T. A., Tromiczak E. G., Zarrinmayeh H., Zimmerman D. M.. Sulfonamide derivatives. US6703425B2, 2004.
  55. Nextmove Software Pistachio. http://www.nextmovesoftware.com/pistachio.
  56. Zhang C.; Lapkin A. A. Reinforcement learning optimization of reaction routes on the basis of large, hybrid organic chemistry–synthetic biological, reaction network data. React. Chem. Eng. 2023, 8, 2491–2504. 10.1039/D2RE00406B. [DOI] [Google Scholar]
  57. Jacob P.-M.; Yamin P.; Perez-Storey C.; Hopgood M.; Lapkin A. A. Towards automation of chemical process route selection based on data mining. Green Chem. 2017, 19, 140–152. 10.1039/C6GC02482C. [DOI] [Google Scholar]
  58. Caspi R.; Billington R.; Keseler I. M.; Kothari A.; Krummenacker M.; Midford P. E.; Ong W. K.; Paley S.; Subhraveti P.; Karp P. D. The MetaCyc database of metabolic pathways and enzymes - a 2019 update. Nucleic Acids Res. 2020, 48, D445–D453. 10.1093/nar/gkz862. [DOI] [PMC free article] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

sb4c00692_si_001.pdf (2.9MB, pdf)

Data Availability Statement

minChemBio is available https://github.com/maranasgroup/chemo-enz under an MIT license. All the codes and data are also available at 10.26207/tbg0-gr88.


Articles from ACS Synthetic Biology are provided here courtesy of American Chemical Society

RESOURCES