Abstract
Monoterpene synthases (mTSs) are a large family of enzymes, which have promising industrial applications, yet remain difficult to engineer due to complex and poorly understood sequence–function relationships. Here, we present a structure-based machine learning (ML) framework that accurately predicts whether a mTS is likely to produce linalool or limonene, chosen as canonical examples of linear and cyclic monoterpene products, respectively. Our approach identifies active site properties, which are used as ML features, enabling functional predictions that go beyond sequence alone. As residue positioning is important for monoterpene synthesis, we created an algorithm to identify structurally conserved residues in the active sites of mTSs, identifying new and existing motifs essential for both general catalysis and specific cyclisation steps. This integrated workflow thus provides insights into terpene synthases that were previously inaccessible and may offer a generalizable strategy for probing and engineering other poorly understood enzyme families. Future work could see this approach used to guide the rational design of mTSs and other hard-to-engineer enzyme families.
An interpretable structure-based ML framework uses active-site properties to predict linalool or limonene (linear/cyclic) production by monoterpene synthases identifying conserved motifs linked to cyclisation to inform future rational engineering.
Introduction
Deep learning models have led to an increased understanding of biological systems and have applications in drug discovery, fundamental structural biology and protein function/annotation.1–3 General ML approaches have been used in biology for decades. However, due to the limited availability of data, ML models were mainly represented in sequence-based approaches, with only a few examples using structural information, typically to supplement a sequence-based model.4–8 Thus, these approaches were less applicable with enzyme families which exhibit poor sequence–function relationships, and those that rely on structural information for engineering. One such example is cytochrome P450 enzymes,9–11 which require exploration of very large sequence spaces because of poor sequence–function mapping, and where ML methods have been developed to address this challenge.12 Another example is β-lactamases,13,14 where structural features account for the largest variation in functional data,15 and serine proteases,16,17 which share neither sequence homology nor overall fold, yet perform the same proteolytic function through conservation of catalytic residues within their active sites.18 Recent sequence-based enzyme-function models have continued to improve functional annotation from sequence alone, but such methods may not be suitable for enzymes with poor-sequence function relationships.19–21 However, the recent advancements in deep learning structure prediction offer the possibility of applying ML approaches to complex structural systems, due to the availability of fairly accurate structures and libraries like the AlphaFold database.22 Structural information has recently been used to build supervised learning models to understand enzyme specificity in other enzyme families with poor sequence–function relationships, including the use of AlphaFold-derived structural features to predict donor specificity in GT-B glycosyltransferases.23 Transfer learning and protein language models have also been applied to enzyme-function prediction, including substrate prediction in RiPP biosynthetic enzymes24 and the prediction of enzyme-catalysed reactions, enzyme functions and EC classifications.25
Terpene synthases (TS) are generally considered to be difficult to engineer, marked by weak sequence–function relationships and high sequence plasticity, where even minor amino acid changes can cause major shifts in product profile.26,27 Structural insights are further hampered by the scarcity of X-ray crystal structures, limiting understanding of product preferences and complicating rational engineering efforts.26–28 Although investigation into single, double and more additive mutations has been performed previously, rational design is limited as the complexity of interactions increases exponentially with each additional mutation.29–31 Despite this, the monoterpene synthase (mTS) family of enzymes remains an attractive engineering target as they typically catalyse the conversion of a single substrate (geranyl pyrophosphate; GPP) into a wide range of monoterpene products (Fig. 1). These products have widespread applications in fragrances, pharmaceuticals, and biofuels, including jet biofuel.32 Traditional monoterpene production methods face significant drawbacks, such as high environmental costs, water use, pesticide dependency, and the complexity of resulting product mixtures requiring extensive purification.33,34 Enzymatic synthesis potentially offers a cleaner, more selective alternative. This, together with the fact that terpene synthases catalyse some of nature's most complex chemistry,35 makes them an important system for both fundamental study and engineering.
Fig. 1. Simplified reaction scheme showing the mTS-catalysed formation of linalool and limonene from GPP. Adapted from ref. 35.

While mTSs exhibit weak sequence–function relationships, mechanistic details are known. Key conserved motifs, such as the DDxxD metal-binding motif, are known to coordinate divalent metal ions and the phosphate moiety of the substrate which facilitates carbocation formation.36 Steric effects and pocket geometry can influence cyclisation.35,36 Specific active site histidine residues have been shown to promote α-terpinyl cation formation,29,37 and cysteine is linked to differences in cyclisation type rather than the initial 1,6-ring closure.31,38 Polar residues have been shown to affect the number of cyclisation events,31,39 and charged residues have been shown to affect product specificity.40–42 Hydrophobic residues have also been linked to controlling water during the formation of hydroxylated monoterpenes (such as linalool),35,42,43 and the role of active site aromatic residues has been well studied and mutational studies have shown that removing aromatic residues often results in the loss of cyclisation.29,35,36,39,42 For example, in a study on (4S)-limonene synthase, the W324A active site mutation led to a significant decrease in enzyme activity and altered product profile from (cyclic) limonene to linear products, but replacing the W324 residue with other aromatic residues did not rescue activity.29 Histidine residues have also been shown to affect linear product formation29 and phenylalanine has also been shown to play an important role in linalool production where the addition of phenylalanine to the active site of a bicyclic (1,8-cineole)-producing mTS causes loss of cyclisation and the production of linalool. However, when a separate phenylalanine was removed it also lost cyclic production capability.42 Overall, aromatic residues have been shown to prevent early quenching of the carbocation by stabilising the positive charge through cation–π interaction,36 which is thought to be important in formation of cyclic monoterpenes (Fig. 1).42 However, these mechanistic insights are typically derived from individual enzymes or small sets of variants and do not readily translate into more general predictive rules. In particular, small changes in active site composition or residue positioning can lead to large shifts in product profile, making it difficult to infer function directly from sequence.26,31,39 As a result, predicting mTS product specificity remains a non-trivial problem, even in the presence of established mechanistic knowledge.
ML methods present a promising solution to elucidate complex, non-linear relationships between enzyme structure, sequence, and product profiles, thereby enabling more precise and scalable engineering strategies. Given their complexity of understanding mTSs, ML is particularly attractive for terpene synthase research, and various ML/AI methods have been used in preliminary analysis,44 molecular replacement,45 and as the basis of screening platforms.46–48 In fact, previous work on sesquiterpene synthases showed that homology modelling data could enhance sequence-based approaches and reduce the species-related bias in an XGBoost based model.8 This model used whole-structural data. However, active site information plays a more important role in terpene synthase catalysis, given the lack of sequence–functional relationship and that mutational ‘hot-spots’ are around the active site.26,27,31,36 It follows that a model based on these features would be more accurate as it would have less ‘noise’. There has also been pre-print published which aims to use ML to understand TSs.49
A particular challenge in engineering mTS enzymes involves the generation and control of the α-terpinyl cation, which is necessary for the cyclisation step that unlocks increased chemical diversity of mTSs (Fig. 1).50 In this study, we investigate features of the mTS structure–function relationship that give rise to linear vs. cyclic product formation, selecting linalool and limonene as canonical examples of linear and cyclic monoterpenes, respectively due to their chemical similarity and the availability of characterized data. We developed an interpretable workflow that uses active-site structural information to distinguish linalool- and limonene-producing mTSs. Unlike uninterpretable protein language models,24 models targeting broader enzyme classes,25 whole-structure,23 broader terpene classifiers,49 and sequence-focused sesquiterpene models,8 our workflow predicts a specific difference in terpene end-product formation: cyclisation. To do this we use AlphaFold2,51 to predict structures of mTSs with experimentally characterized product profile data showing linalool or limonene as a major product,52–105 Fpocket106 to identify the active site, statistical methods to explore functional motifs, and we build a predictive ML model that can discriminate between mTSs that catalyse the formation of linalool or limonene as a proof of principle that links mTS sequence to a specific function (cyclisation).
Methods
All code, AlphaFold structures, Fpocket output and ML training and evaluation data used in this paper is freely available at https://github.com/cathaloraghallaigh/ATC. The dataset was started using data provided by Dr Nicole Leferink (University of Manchester), which had 18 suitable sequences. It was expanded by searching through all ‘linalool’ and ‘limonene’ annotated sequences on UniProt,107 to find those with published GC-MS data confirming their classification. An exhaustive manual literature search was also carried out to find any mention of linalool or limonene synthases, and the papers were then checked to find if the enzymes had been experimentally characterized. A0A8H4VHP2 and O22340 were not included, as their inclusion reduced model performance. This dataset of sequences, their host organism and the relevant papers associated with them can be found in Table S1.
AlphaFold2.1.1 (ref. 51) was used for all structure predictions. The ‘best’ relaxed model (highest pLDDT) was chosen in each case. Active site information was generated using Fpocket106 with standard settings, which provides pocket-level descriptors and identifies the residues lining each predicted pocket. These outputs were used to derive active-site features describing pocket geometry, physicochemical properties and amino acid composition; the full feature list is provided in Table S2.108 This workflow (from original protein sequence to active site characteristics in a dataset) is mostly automated. The only manual section of this process is identifying which ‘pocket’ is the active site by visual inspection, as Fpocket typically identifies multiple pockets. This was performed based on the conserved architecture of monoterpene synthase active sites, including conserved active site motifs such as the DDxxD metal binding motif, and by comparison with the active site location in X-ray crystal structures through structural alignment. Fpocket106 was selected because it is an open-source and widely used cavity detection algorithm,109–111 which can be easily integrated into automated workflows, as its outputs are provided directly as PDB files. This allowed the entire pipeline from structure prediction to feature extraction to be as automated and reproducible as possible. In addition, Fpocket provides a relatively large number of quantitative pocket descriptors, and the algorithm also identifies the residues lining each pocket, enabling extraction of active-site amino acid composition for downstream analysis and model training. Structural sensitivity was assessed by repeating pocket identification and feature extraction using the second- and third-ranked AlphaFold2 models. Structures with pocket residue-count similarity below 75% of the original model were excluded to remove pockets affected by pinching or extension beyond the active site. Remaining structures were compared using feature similarity and leave-one-protein-out classification with the original model hyperparameters (Table S9).
Computational analyses were carried out in Python. Statistical analyses were performed using SciPy, pandas, and the built-in collections module. Data visualization and sequence logo generation were conducted with Matplotlib, seaborn, Logomaker, and Biopython. Machine-learning models were implemented with NumPy, scikit-learn, and XGBoost. Aligned protein structures (using the cealign PyMOL command) were analysed by extracting Cα coordinates from multi-model CIF files. Reference positions were defined as the Cα coordinates of active-site residues identified by Fpocket in the reference structure. For each reference position, the nearest Cα atom in every other protein was identified within a distance cutoff (4.0 Å). To maintain one-to-one correspondence, an assignment procedure resolved conflicts: if two positions competed for the same residue, the closer match was retained, and the displaced position was reassigned to its next best candidate, if there is one. This prevented residues from being reused across targets within a protein. Proteins were then grouped according to functional labels, and residue frequencies were calculated at each reference position for every group. Positions showing residues at ≥15% frequency were summarized with their group differences (Table S10), and sequence logos were generated to visualize residue distributions across groups. In addition, the script outputs per-protein tables listing the nearest residues to each reference position together with their distances, so individual proteins of interest can be investigated.
ML models were built to use active site features for binary classification (linear vs. cyclic products) using the XGBoost algorithm.112 XGBoost classifiers were trained with largely default hyperparameters: log-loss evaluation metric, histogram-based tree construction, L1 regularization α = 0.1, L2 regularization λ = 1.0 and learning rate of 0.1. Computations were accelerated on CUDA-enabled GPUs for the Bayesian optimization XGBoost models but CPUs were used instead for the rest, to ensure reproducibility. Feature importance was assessed using the gain metric in XGBoost, which measures how much each feature helps improve the model's predictions. Class imbalance was addressed by applying inverse class-frequency sample weights during training. Model selection targeted leave-one-out cross-validated (LOO-CV) accuracy while using scikit-optimize. Training used inverse-frequency sample weighting, where each observation was weighted by the reciprocal of its class prevalence to balance contributions from the linear and cyclic classes. Accuracy was optimized using leave-one-out cross-validation (LOO-CV) accuracy as the objective within a scikit-optimize Bayesian optimization (gp_minimize) framework over categorical feature-group choices. The acquisition function was set to Lower Confidence Bound (acq_func = “LCB”) with an exploration parameter (kappa = 10.0) to promote exploration of uncertain regions. Early stopping was applied when improvements in LOO-CV accuracy fell below 1 × 10−4 for 100 successive evaluations. Although the search allowed up to 500 evaluations, peak performance was reached within the first 100 iterations. Following feature selection, the FS1 features and hyperparameters were fixed. The model was trained on the 80% test/train set and evaluated on the independent holdout set, with performance assessed using accuracy, balanced accuracy, MCC, and per-class precision, recall, specificity and F1 score. Given the small dataset, model robustness was further assessed across the complete dataset using repeated stratified five-fold cross-validation with 10 repeats. To provide insight into feature importance and the direction of individual feature contributions across full dataset, fold-level XGBoost gain and out-of-fold SHAP113 values were also calculated.
Further model generalisation was tested using enzymes selected from a published sesquiterpene synthase dataset.8,114 Enzymes classified as linear or producing exclusively monocyclic products through a single 1,6-cyclisation were selected. AlphaFold structures were generated as described above and analysed using Fpocket. Only complete, structurally plausible models for which Fpocket identified a coherent pocket at the expected active-site location were retained, giving 66 enzymes comprising 39 linear and 27 cyclic synthases. The 12 features comprising FS1 and the XGBoost hyperparameters established using the mTS dataset were retained without further feature selection or optimisation. Balanced sample weighting was used to account for class imbalance. Performance was assessed using repeated stratified five-fold cross-validation with 10 repeats.
Results and discussion
We first compiled a library of 88 experimentally characterized mTSs, containing 29 cyclic and 59 linear producing enzymes (Table S1); each enzyme has a published product profile determined using GC-MS.52–105 A UniProt search of any proteins labelled linalool or limonene synthase was combined with a literature search of limonene and linalool producing enzymes, and all enzymes in this library have been reported to produce either linalool or limonene as a major product, which we define here as where ≥70% of the reported products belong to either the linear or cyclic class. This threshold was selected because the reported product distributions clustered above 70% for either the linear or cyclic class. We graphed the structural and phylogenetic relationship between the enzymes in tree format (Fig. 2). The sequence-based tree reflects evolutionary relationships inferred from protein sequence, whereas the structure-based tree groups enzymes by similarity in their overall protein structures. There seems to be some clustering between the groups. However, clear discrimination between the groups is not apparent, which is understandable, as previously stated, the sequences are ‘plastic’ and only small residue changes (mostly within the active site) cause a change in product profile. This is further exemplified when we carried out homology-aware evaluation of sequence-based nearest-neighbour classification, which achieved 66.7% accuracy with BLAST115 and 61.6% with phmmer,116 whereas structure-based classification using Foldseek117 achieved 78.2% accuracy (Table S18).
Fig. 2. Phylogenetic analysis of the 88 monoterpene synthases (mTSs) in our curated library, colored by product preference: linalool (linear, orange) and limonene (cyclic, blue). (A) Sequence-based tree generated from a multiple sequence alignment using Clustal Omega,118 representing evolutionary relationships inferred from protein-sequence similarity. (B) Structure-based tree constructed from pairwise structural alignments obtained with ColabAlign,119 grouping enzymes by similarity in their overall three-dimensional structures. Both trees were visualized using the interactive tree of life (iTOL) software.120 Enlarged versions of both trees, labelled with UniProt accessions and scientific species names, are provided in Fig. S1.

Due to the lack of availability of experimental structures for most enzymes in the library (available X-ray crystal structures found in Table S3), structures were predicted for all enzymes using AlphaFold2.51 As predicted structures are increasingly used as inputs to downstream ML pipelines, benchmarking structure-prediction performance on curated experimental datasets remains an important prerequisite for reliable biological inference.121 Where X-ray crystal structures were available, comparison with predicted structures generally showed good agreement, including when structures were predicted when the corresponding X-ray crystal structure(s) were excluded (“blacklisted”) from the template selection in AlphaFold, and also when the enzymes were predicted to be homodimers (we have modelled them all as monomers here; Tables S4–S8). Analysis of one X-ray crystal structure, PDB: 8GY0 (enzyme: A0A8H4VHP2) showed poor agreement between the predicted and experimental structures (Table S7) and inclusion of this enzyme in subsequent analysis resulted in a decrease in model accuracy. Consequently, we identified this as an outlier and have not included it in our dataset.
Previous structural analysis has shown that mTS enzymes can undergo an open to closed conformational change upon substrate binding, which involves the inward movement of the H-α1 helix and repositioning of a conserved active-site tyrosine.76 Comparison of the position of this residue in the AlphaFold models used in this study shows that they almost all (85/88) adopt closed-like conformations similar to PDB: 2ONG (Fig. S3B). The tendency of AlphaFold2 to predict ‘closed’ protein structures has been reported previously.122 Using the predicted structures, we next found active site information using the Fpocket106 cavity detection algorithm, which is based on Voronoi tessellation to identify protein cavities. Along with active site volume, charge and scores of other properties, this algorithm identifies residues within the active site pocket (Fig. 3).
Fig. 3. Statistical analysis of the active sites of the linear (orange) and cyclic (blue) mTS groups. Shown are the amino acid frequency and number of residues (top), density plots of amino acid and pocket based volume scores, hydrophobicity and polarity (middle, measured in density) and aromatic residue statistics (bottom, measured in proportion per protein). A larger version of this figure is found in the supporting information (Fig. S5).

Fpocket identifies residues that line the active site/cavity. Structural comparison of the active site residues identified by Fpocket highlighted relatively larger differences between AlphaFold models and X-ray crystal structures (Tables S4–S8). The worst agreement was observed for A0A1C9J6A7, where >4 Å RMSD in these residues was observed between the X-ray crystal structure 5UV0 and the AlphaFold model (Table S5). This X-ray crystal structure is in an open conformation76 and further inspection of these data showed that much of the variance comes from specific residues that are found within disordered regions that are not well resolved in the other A0A1C9J6A7 X-ray crystal structures. When excluding these residues, the active site heavy-atom RMSD falls to ≈2.6 Å. Crucially, Fpocket identifies most residues within the mTS active site despite considerable uncertainty in their positions.
The volume of the active site has been shown to play an important part in product profile, mostly for substrate specificity.27 However, the volume of amino acids within the active site have also been shown to be correlated to end product due to the contour causing different steric effects on the carbocation.123 As expected, the amino acid-based volume score significantly differs between the active sites of the cyclic vs. linear producing enzymes in our library (Fig. 3). The average van der Waals volume contributions of the residues lining the pocket is higher in the cyclic (limonene) group: mean cyclic scores 4.50 Å3 (SEM of 0.04), mean linear scores 4.36 Å3 (SEM of 0.02), p = 0.004. The overall size of the active site pocket appears to be more rigidly conserved within the cyclic group of enzymes, with the active site of the linear group on average being larger and more variable: mean cyclic scores 802 Å3 (SEM of 35 Å3), mean linear scores 1134 Å3 (SEM of 71 Å3), p = 0.0001. More generally, most active site features were found to be more variable in the linalool producing enzymes, but this may be due to the larger size of this dataset. Alternatively, it may be that mutations in cyclic mTSs are more likely to cause loss-of-function, as the α-terpinyl cation (Fig. 1) must be stabilised for longer and it must perform more complex chemistry, thus there is a stronger evolutionary pressure for cyclic monoterpene synthesis.35,36
The role of aromatic residues in the active site of TS has been well studied.35,36,39,42 On average, tryptophan was found with equal frequency in the active sites of limonene and linalool producing enzymes (Fig. 3). Conversely, there are significant differences between the frequency of other aromatic residues identified in the active sites of these two enzyme groups. However, when normalized by the number of residues identified in each active site, there is only a marginally higher proportion of aromatic residues in the limonene producing enzymes: mean cyclic ratio 0.225 (SEM of 0.008), mean linear ratio 0.200 (SEM of 0.006), p = 0.095. Hydrophobic residues have also been linked to controlling water during the formation of hydroxylated monoterpenes (such as linalool),35,42 which may explain why there is not a clearer difference between the hydrophobicity scores for two enzyme groups (Fig. 3). More generally, the hydrophobicity of TS active sites has been stated to play an important role in cyclisation, where, along with a more aromatic pocket, a more hydrophobic pocket would prevent premature quenching of the cation intermediate(s).35 However, our results here show that the hydrophobicity score (mean of hydrophobicity of active site residues124) is higher for the linear group: mean cyclic score 18.2 (SEM of 0.845), mean linear score 22.5 (SEM of 0.930), p = 0.001.
Whilst the amino acid composition of the active site is important, the spatial arrangement of residues within the active site is expected to play an important role in enzyme catalysis.125 To investigate structural conservation of the active site, we aligned the structure of each enzyme in the library with the X-ray crystal structure of Mentha spicata limonene synthase (PDB ID: 2ONG), which we use as a reference structure. Alternative analysis using a linalool synthase gave similar results (Fig. S6 and Table S11). For each specified residue within the reference structure, the algorithm finds the closest corresponding residue for each of the proteins, with a 4.0 Å cut-off (Table S10). This effectively aligns the sequences of each active site, which is shown in sequence logo format in Fig. 4. A more detailed structural analysis was not attempted due to limitations of the structural prediction, which we do not expect to predict metal binding or fully reproduce loop rearrangement, correct side chain packing and H-bonding within the active site (Fig. S2 and S3).
Fig. 4. (A) and (B) Active site sequence conservation in the linear (A) and cyclic (B) members of the mTS library. Each active site sequence was determined using PDB ID: 2ONG as reference and an additional analysis using PDB ID: 5NX5 is shown in Fig. S6. Total letter height at each position indicates occupancy, where heights below 1 reflect enzymes for which no residue was identified within 4.0 Å of the reference position. Enlarged logos with separate occupancy plots are shown in Fig. S7. (C) Difference sequence logo showing residue-frequency differences between the linear and cyclic groups. Positive and negative values indicate greater residue frequency in the linear and cyclic groups, respectively. (D) Alignment of the AlphaFold2 model of F8TWD2 (member of the linear group; orange) with an X-ray crystal structure of a limonene synthase (PDB ID: 2ONG; blue). The all-atom RMSD is 1.08 Å. Conserved residues within the active site are highlighted right. The 2-fluorogeranyl diphosphate (FGPP) substrate analog and three Mn2+ found in the X-ray crystal structure are shown in magenta and purple, respectively. Fig. S2 shows an overlay of all active site residues in each enzyme.

The active site alignments identified established highly conserved residues and motifs across both groups of enzymes, including W324,29 and the metal binding D352/353/356 (part of the DDxxD domain),36 and Y573, which, despite being conserved highly in both groups, it is always conserved in the cyclic group. This conservation has previously been shown to be essential for cyclic monoterpene synthases.39 There are also conserved motifs which have not yet been shown, such as, R315, E430 and S452 which are likely to play important roles in monoterpene catalysis as they are conserved across both groups. In fact, serine can be involved in metal binding, but also has been shown to indirectly impact catalysis of terpene synthases in TSs.31,126 Mutating serine residues in the active site of mTSs has been shown to lead to slight change in product profile,29 or inactivation.127
There are some residues which are more commonly conserved in the linear group of enzymes, most noticeably, G454 (2ONG numbering). There are many residues which are more frequently conserved in the cyclic group of enzymes which suggests they play a role in the formation or stabilisation of the α-terpinyl cation. These include Y63, T349, Y427 (which are typically aromatic residues), E504, R507, K512, D577, and H579 (Fig. 4). Consistent with this, previous work has shown that removal of H579 results in reduced cyclisation capacity, as well as T349 and R507, to a lesser extent.29 Additionally, for the linear group, there also appears to be up to 5 residues missing from the ‘end’ of the active site sequence (frequency < 100% at a given position, with no corresponding residues within a 4.0 Å radius), which is due to some enzymes in this group having fewer residues associated with the active site. These ‘missing’ residues are at the surface of the protein, and in 2ONG are found in the J loop, which forms part of the ‘cap’ that could protect the cations from premature quenching,128 (Fig. S4). While our analysis in Fig. 3 and 4 provides insight into difference in mTS enzymes that produce cyclic vs. linear monoterpenes, we do not identify any discriminatory features (i.e. differences found only in one group). We reasoned that a ML model based on active site characteristics may be able to perform this classification (workflow diagram seen in Fig. S9).
Extreme gradient boosting (XGBoost)112 is a decision tree-based supervised ML algorithm that has been shown to build accurate classification models and can perform well with relatively small datasets.129–131 Additionally, as these models are more interpretable than black-box models, we can understand and learn from which features influence the ability of the model to distinguish between the groups.132 Due to the modest size of the dataset, leave-one-out (LOO) cross-validation was also used.133,134 This method of cross-validation leaves one dataset entry (i.e. enzyme) out of the training dataset for each round, only testing the one missing observation. This is then repeated for each entry within the dataset. This approach is more computationally expensive than other cross-validation methods, but it has been shown to minimize the chance of overfitting135 and has been used successfully in ML models for protein/enzymology before.136–138 Bayesian optimization (BO) has also been shown to be useful for feature selection for high-dimensional data used in training molecular biology XGBoost models.139,140
Whilst the spatial arrangement of active site residues plays an important role in enzyme catalysis, explicitly capturing this typically requires more complex models, such as geometric or graph-based approaches.141 Given the modest size of the dataset used in this study and the use of structure prediction (cf. experimental structures), such approaches are expected to lead to overfitting and may not provide additional value. We therefore opted to use Fpocket-derived active site properties, which captures key characteristics of the active site without requiring high-resolution structural detail or large ML models. Fpocket can generate up to 46 features describing the active site, including features describing the occurrence of each of the 20 individual amino acids. Using all Fpocket features (Table S2) would result in ML models with large search space, so to reduce the size of the feature space, non-amino acid features were grouped based on redundancy between the features. Prior to any feature-space exploration, the dataset was randomly split in a stratified manner, to a ∼80% test/train (69 members; 47 linear and 22 cyclic). All Bayesian optimization and feature selection steps were performed exclusively on the test/train dataset with the ∼20% validation set (18 members, 12 linear and 6 cyclic) reserved for final model evaluation.
To elucidate the best combination of features for ML model training, BO was performed (Table S12). Features were grouped, allowing the reduction of the number of features whilst reducing the testing space needed. Combining BO with the LOO-XGBoost ML model resulted in a model with accuracy of 92.0% (Fig. S10–S12 and Tables S13, S14). The features identified in this model are denoted FS1. Using these features, log loss was calculated using stratified k-fold (5 folds, 10 repeats, Fig. 5). Here, the Bayesian-optimized FS1 model performed with higher mean accuracy (83.5 ± 8.4%) and lower mean validation log-loss (0.379) than the model trained with the full feature set (80.7 ± 8.2% accuracy, log-loss = 0.441). This supports the effectiveness of the feature selection and demonstrates that a reduced, interpretable subset improves overall model performance and generalization. The resulting 12-feature FS1 model, trained using the selected XGBoost parameters, was therefore taken forward as the final model configuration. The resulting 12 feature (FS1) model was then trained on the entire test/train set (69 observations; 80% of the dataset) and evaluated on the independent hold-out ‘validation set’ (18 observations; 20% of the dataset). This yielded an overall accuracy of 94.4%, with only one linear observation misclassified (UniProt ID: A0A6C0M6B5, confusion matrix in Fig. S13). Repeated stratified five-fold cross-validation across the complete modelling dataset (87 observations; 10 repeats) gave a mean accuracy of 90.0% (SD 2.4%), balanced accuracy of 89.2% (SD 2.9%), and MCC of 0.775 (SD 0.054; Fig. S13). As the dataset comprises a single enzyme family, similar sequences and structures occur across the training and validation sets. To assess whether this inflated model performance, leave-one-cluster-out evaluation was performed, giving accuracies of 87.4% for sequence clusters and 93.1% for structural clusters, with balanced accuracies of 86.9% and 93.0%, respectively (Table S18).
Fig. 5. Evaluation of the 12-feature FS1 XGBoost model. (Top) Correlation matrix of the final FS1 features selected using leave-one-out cross-validation, showing low correlation between features. (Bottom) Mean training and validation log-loss curves obtained using repeated stratified five-fold cross-validation of the test/train set (69 observations; 10 repeats). The full feature set was included for comparison (final mean validation log loss: 0.442), with FS1 achieving a lower final mean validation log loss of 0.384. The final 12-feature FS1 XGBoost model was trained on the entire test/train set and evaluated on the independent holdout set (18 observations), achieving an overall accuracy of 94.4%, balanced accuracy of 95.8% and MCC of 0.886. For the linear class, precision was 100%, recall was 91.7%, specificity was 100% and F1 score was 95.7%. For the cyclic class, precision was 85.7%, recall was 100%, specificity was 91.7% and F1 score was 92.3%.

Further, we explore alternative classification methods using both the full feature set and FS1. This analysis included the use of additional methods, including Random Forest and logistic regression, ESM-2, phmmer and Foldseek, which all performed similarly, or worse than XGBoost (Tables S15, S18 and Fig. S15). This demonstrates the predictive power of active site information and indicates that this type of information is important for discriminating between the two groups. FS1 was selected to continue from here, given better overall performance.
Of the 12 features within FS1 (Fig. 6), histidine ranked first by both mean gain and mean absolute SHAP. Cysteine ranked second and non-polar residue count ranked third by mean gain, whereas non-polar and polar residue counts ranked second and third by mean absolute SHAP, respectively. Higher histidine counts favoured cyclic predictions, consistent with the reported role of histidine in α-terpinyl cation formation.29,37 Polarity-related variables (‘polar’ and ‘non-polar’) were also among the most influential features, with polar residue count ranking third by mean absolute SHAP. Higher non-polar residue counts favoured linear predictions, whereas higher polar residue counts generally favoured cyclic predictions, although the directional effect was weaker for polar residue count. Polarity-related active-site variables have been shown to affect cyclisation, however, primarily influencing the type or number of cyclisation events.31,39 Cysteine ranked third by mean gain but ninth by mean absolute SHAP and has been linked to differences in cyclisation type rather than the initial 1,6-ring closure.31,38 The importance and prevalence of these features in the final models may also be due to the fact the hydroxylation of the end products used for the model. It could be interesting to investigate how they contribute to mTS catalysis in future experimental work.
Fig. 6. Feature importance for the FS1 XGBoost model across repeated stratified five-fold cross-validation of the complete modelling dataset. Top left: Mean XGBoost gain across the 50 fitted models. Histidine had a mean gain of 6.626 (SD of 1.181), cysteine had a mean gain of 3.466 (SD of 2.147) and non-polar residue count had a mean gain of 2.637 (SD of 0.560). Top right: Mean absolute out-of-fold SHAP importance across the 50 fitted models. Histidine had a mean absolute SHAP value of 1.761 (SD of 0.304), non-polar residue count had a mean absolute SHAP value of 1.268 (SD of 0.344) and polar residue count had a mean absolute SHAP value of 0.772 (SD of 0.204). Bottom: Directional SHAP values pooled across the out-of-fold predictions. Positive SHAP values favoured the cyclic class and negative SHAP values favoured the linear class. Higher histidine and polar residue counts favoured cyclic predictions, while higher non-polar residue counts favoured linear predictions. Error bars in (top left) and (top right) show the SD across the 50 fitted models.

During preparation of this manuscript, a preprint from the Pluskal group reported a UniProt-mined terpene synthase dataset.49,142 There are 14 enzymes within the preprint dataset (5 of which produce limonene and 9 of which produce linalool) which are not present in our dataset (Table S17). To benchmark our approach, we refit the final XGBoost model (FS1, same hyperparameters/class weights) on the full training and validation data (87 observations) and evaluated it on these additional 14 enzymes, resulting in an accuracy of 92.9% (with only the limonene-producing, UniProt ID: G1JUH4, predicted incorrectly, Fig. S14). We also tested a selection of mTSs that catalyse the production of different linear and (bi)cyclic monoterpene products. Of these, 7/10 of the linear monoterpene-producing enzymes were predicted to produce linalool (i.e. linear), while 8/8 monocyclic and 4/9 bicyclic monoterpene-producing enzymes were predicted to produce limonene (i.e. (mono)cyclic), respectively (Fig. S16 and Table S19). This suggests that the ML model has some transferability to predict linear and monocyclic monoterpene products, but may not generalize to wider terpenes, as exemplified by the poor performance with the bicyclic monoterpene-producing enzymes. It would be interesting to use the new TPS dataset to extend the ML model to explore additional mechanistic steps in mTSs. However, this would require a different model architecture to do multi-class and/or multi-label (for multiple product) classification, which is outside the scope of the current study.
As a preliminary investigation of wider generalisability, the workflow was retrained using 66 sesquiterpene synthases, 39 linear and 27 monocyclic (1,6-cyclising) (Table S20) selected from a previously published dataset.8,114 This dataset covers enzymes producing a broader range of products than the predominantly linalool- and limonene-based mTS dataset. Without further feature selection, hyperparameter optimisation or model improvement, repeated stratified five-fold cross-validation achieved a mean accuracy of 84.4% (SD 3.8%), balanced accuracy of 84.1% (SD 4.0%), and MCC of 0.680 (SD 0.079; Fig. S17).
Although we cannot claim wider generalizability of the existing ML model, the approach taken here addresses a broader class of enzyme engineering challenges that are not unique to monoterpene synthases. In particular, mTSs share key characteristics with enzyme families such as cytochrome P450s,9–12 β-lactamases,13–15 and serine proteases,16,17 where weak sequence–function relationships and strong dependence on local active site chemistry complicate prediction and engineering. In this study, these challenges are addressed by explicitly defining and aligning structurally conserved active site residues and by representing functional differences through the local compositional and physicochemical properties of these regions, rather than relying on whole-protein or sequence-based representations.
Conclusions
Here, we have developed a simple workflow that allows structure-guided sequence and feature comparisons of related classes of enzymes; in this case the limonene and linalool producing mTSs. This exploits structure prediction using AlphaFold2,51 active site identification and characterization using Fpocket,106 nearest-neighbor active site sequence alignment and an XGBoost112 based ML model with good classification performance. Our active site alignment algorithm identified a number of well-known mTS motifs, as well as identifying additional residues that appear to be important for general catalysis due to their conservation. Analysis of the relative importance of features in the ML classifier further identified some known features that aided in distinguishing between limonene and linalool producing mTSs (small, charged, uncharged, aromatic active site residues), with more that do not yet have a known direct effect on monoterpene cyclisation. This integrative approach offers a data-driven path to decoding enzyme specificity and as this method does not require mechanistic information, only requiring identification of active site residues, we anticipate that this can be readily integrated into (semi)-rational engineering approaches for terpene synthases and other ‘difficult-to-engineer’ enzyme families.
Conflicts of interest
There are no conflicts to declare.
Abbreviations
- ML
Machine learning
- mTS
Monoterpene synthase
- TS
Terpene synthase
- AS
Active site
- mTS
Monoterpene synthase
- BO
Bayesian optimization
- LOO
Leave one out (cross-validation)
- XGBoost
Extreme gradient boosting
Supplementary Material
Acknowledgments
This work was partly supported by BBSRC grant: BB/Y008456/1 and EPSRC grants: EP/S01778X/1 and EP/S022856/1. COR was supported by an EPSRC studentship ref. 2602504. The authors would like to acknowledge the assistance given by Research IT and the use of the Computational Shared Facility at The University of Manchester.
Data availability
All Python code, AlphaFold structures, Fpocket output and ML training and evaluation data used in this paper is freely available at https://github.com/cathaloraghallaigh/ATC and at ZENDO (https://doi.org/10.5281/zenodo.22877417). All additional data supporting this study are provided in the supplementary information (SI) accompanying this paper. Supplementary information: enzyme datasets and feature definitions, benchmarking of AlphaFold predictions against crystal structures, active-site residue analyses, and sensitivity analyses of structural model choice and residue mapping. It also provides machine-learning algorithm comparisons, cross-validation results, independent validation, and further testing on monoterpene synthases with other products and on sesquiterpene synthases. See DOI: https://doi.org/10.1039/d6dd00312e.
References
- Lin B. Luo X. Liu Y. Jin X. A Comprehensive Review and Comparison of Existing Computational Methods for Protein Function Prediction. Briefings Bioinf. 2024;25(4):bbae289. doi: 10.1093/bib/bbae289. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Kumar N. Srivastava R. Deep Learning in Structural Bioinformatics: Current Applications and Future Perspectives. Briefings Bioinf. 2024;25(3):bbae042. doi: 10.1093/bib/bbae042. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Jänes J. Beltrao P. Deep Learning for Protein Structure Prediction and Design—Progress and Applications. Mol. Syst. Biol. 2024;20(3):162–169. doi: 10.1038/s44320-024-00016-x. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Röttig M. Rausch C. Kohlbacher O. Combining Structure and Sequence Information Allows Automated Prediction of Substrate Specificities within Enzyme Families. PLoS Comput. Biol. 2010;6(1):e1000636. doi: 10.1371/journal.pcbi.1000636. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Roy A. Yang J. Zhang Y. COFACTOR: An Accurate Comparative Algorithm for Structure-Based Protein Function Annotation. Nucleic Acids Res. 2012;40(W1):W471–W477. doi: 10.1093/nar/gks372. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Chevrette M. G. Aicheler F. Kohlbacher O. Currie C. R. Medema M. H. SANDPUMA: Ensemble Predictions of Nonribosomal Peptide Chemistry Reveal Biosynthetic Diversity across Actinobacteria. Bioinformatics. 2017;33(20):3202–3210. doi: 10.1093/bioinformatics/btx400. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Mou Z. Eakes J. Cooper C. J. Foster C. M. Standaert R. F. Podar M. Doktycz M. J. Parks J. M. Machine Learning-Based Prediction of Enzyme Substrate Scope: Application to Bacterial Nitrilases. Proteins: Struct., Funct., Bioinf. 2021;89(3):336–347. doi: 10.1002/prot.26019. [DOI] [PubMed] [Google Scholar]
- Durairaj J. Melillo E. Bouwmeester H. J. Beekwilder J. Ridder D. d. Dijk A. D. J. v. Integrating Structure-Based Machine Learning and Co-Evolution to Investigate Specificity in Plant Sesquiterpene Synthases. PLoS Comput. Biol. 2021;17(3):e1008197. doi: 10.1371/journal.pcbi.1008197. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Klenk J. M. Dubiel P. Sharma M. Grogan G. Hauer B. Characterization and Structure-Guided Engineering of the Novel Versatile Terpene Monooxygenase CYP109Q5 from Chondromyces apiculatus DSM436. Microb. Biotechnol. 2019;12(2):377–391. doi: 10.1111/1751-7915.13354. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Hansen C. C. Nelson D. R. Møller B. L. Werck-Reichhart D. Plant Cytochrome P450 Plasticity and Evolution. Mol. Plant. 2021;14(8):1244–1265. doi: 10.1016/j.molp.2021.06.028. [DOI] [PubMed] [Google Scholar]
- Yang L. Pu Z. Wu J. Liu X. Wang Z. Yu H. Wang L. Meng Y. Xu G. Yang L. Zheng W. Regulating the N-Oxidation Selectivity of P450BM3 Monooxygenases for N-Heterocycles through Computer-Assisted Structure-Guided Design. Nat. Commun. 2025;16(1):6494. doi: 10.1038/s41467-025-61773-3. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Wu Z. Kan S. B. J. Lewis R. D. Wittmann B. J. Arnold F. H. Machine Learning-Assisted Directed Protein Evolution with Combinatorial Libraries. Proc. Natl. Acad. Sci. U. S. A. 2019;116(18):8852–8858. doi: 10.1073/pnas.1901979116. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Judge A. Sankaran B. Hu L. Palaniappan M. Birgy A. Prasad B. V. V. Palzkill T. Network of Epistatic Interactions in an Enzyme Active Site Revealed by Large-Scale Deep Mutational Scanning. Proc. Natl. Acad. Sci. U. S. A. 2024;121(12):e2313513121. doi: 10.1073/pnas.2313513121. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Lu S. Montoya M. Hu L. Neetu N. Sankaran B. Prasad B. V. V. Palzkill T. Mutagenesis and Structural Analysis Reveal the CTX-M β-Lactamase Active Site Is Optimized for Cephalosporin Catalysis and Drug Resistance. J. Biol. Chem. 2023;299(5):104630. doi: 10.1016/j.jbc.2023.104630. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Chen J. Z. Bisardi M. Lee D. Cotogno S. Zamponi F. Weigt M. Tokuriki N. Understanding Epistatic Networks in the B1 β-Lactamases through Coevolutionary Statistical Modeling and Deep Mutational Scanning. Nat. Commun. 2024;15(1):8441. doi: 10.1038/s41467-024-52614-w. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Fraser B. J. Wilson R. P. Ferková S. Ilyassov O. Lac J. Dong A. Li Y.-Y. Seitova A. Li Y. Hejazi Z. Kenney T. M. G. Penn L. Z. Edwards A. Leduc R. Boudreault P.-L. Morin G. B. Bénard F. Arrowsmith C. H. Structural Basis of TMPRSS11D Specificity and Autocleavage Activation. Nat. Commun. 2025;16(1):4351. doi: 10.1038/s41467-025-59677-3. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Klaushofer R. Bloch K. Eder L. S. Marzaro S. Schubert M. Böttcher-Friebertshäuser E. Brandstetter H. Dahms S. O. Structural Insights into Proprotein Convertase Activation Facilitate the Engineering of Highly Specific Furin Inhibitors. Nat. Commun. 2025;16(1):8206. doi: 10.1038/s41467-025-63479-y. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Bryant D. H. Moll M. Chen B. Y. Fofanov V. Y. Kavraki L. E. Analysis of Substructural Variation in Families of Enzymatic Proteins with Applications to Protein Function Prediction. BMC Bioinf. 2010;11(1):242. doi: 10.1186/1471-2105-11-242. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Heid E. Probst D. Green W. H. Madsen G. K. H. EnzymeMap: Curation, Validation and Data-Driven Prediction of Enzymatic Reactions. Chem. Sci. 2023;14(48):14229–14242. doi: 10.1039/D3SC02048G. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Huang J. Gao Q. Tang Y. Wu Y. Zhang H. Qin Z. A Deep Learning Model for Type II Polyketide Natural Product Prediction without Sequence Alignment. Digital Discovery. 2023;2(5):1484–1493. doi: 10.1039/D3DD00107E. [DOI] [Google Scholar]
- Dhibar S. Basak S. Jana B. Prediction of Enzyme Function Using an Interpretable Optimized Ensemble Learning Framework. Chem. Sci. 2025;16(39):18438–18449. doi: 10.1039/D5SC04513D. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Varadi M. Bertoni D. Magana P. Paramval U. Pidruchna I. Radhakrishnan M. Tsenkov M. Nair S. Mirdita M. Yeo J. Kovalevskiy O. Tunyasuvunakool K. Laydon A. Žídek A. Tomlinson H. Hariharan D. Abrahamson J. Green T. Jumper J. Birney E. Steinegger M. Hassabis D. Velankar S. AlphaFold Protein Structure Database in 2024: Providing Structure Coverage for over 214 Million Protein Sequences. Nucleic Acids Res. 2024;52(D1):D368–D375. doi: 10.1093/nar/gkad1011. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Hennen S. G. Bomble Y. J. Urbanowicz B. R. Bharadwaj V. S. Decoding Substrate Specificity Determining Factors in Glycosyltransferase-B Enzymes – Insights from Machine Learning Models. Digital Discovery. 2025;4(8):2214–2228. doi: 10.1039/D4DD00338A. [DOI] [Google Scholar]
- Clark J. D. Mi X. Mitchell D. A. Shukla D. Substrate Prediction for RiPP Biosynthetic Enzymes via Masked Language Modeling and Transfer Learning. Digital Discovery. 2025;4(2):343–354. doi: 10.1039/D4DD00170B. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Fruscia L. D. Weber J. M. Leveraging Large Language Models for Enzymatic Reaction Prediction and Characterization. Digital Discovery. 2025;4(12):3588–3609. doi: 10.1039/D5DD00187K. [DOI] [Google Scholar]
- Leferink N. G. H. Ranaghan K. E. Karuppiah V. Currin A. Kamp M. W. V. D. Mulholland A. J. Scrutton N. S. Experiment and Simulation Reveal How Mutations in Functional Plasticity Regions Guide Plant Monoterpene Synthase Product Outcome. ACS Catal. 2019;8(5):3780–3791. doi: 10.1021/ACSCATAL.8B00692. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Lei D. Qiu Z. Qiao J. Zhao G. R. Plasticity Engineering of Plant Monoterpene Synthases and Application for Microbial Production of Monoterpenoids. Biotechnol. Biofuels. 2021;14(1):1–15. doi: 10.1186/S13068-021-01998-8. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Leferink N. G. H. Jervis A. J. Zebec Z. Toogood H. S. Hay S. Takano E. Scrutton N. S. A “Plug and Play” Platform for the Production of Diverse Monoterpene Hydrocarbon Scaffolds in Escherichia coli. ChemistrySelect. 2016;1(9):1893–1896. doi: 10.1002/SLCT.201600563. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Srividya N. Davis E. M. Croteau R. B. Lange B. M. Functional Analysis of (4S)-Limonene Synthase Mutants Reveals Determinants of Catalytic Outcome in a Model Monoterpene Synthase. Proc. Natl. Acad. Sci. U. S. A. 2015;112(11):3332–3337. doi: 10.1073/PNAS.1501203112/-/DCSUPPLEMENTAL. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Styczynski M. P. Fischer C. R. Stephanopoulos G. N. The Intelligent Design of Evolution. Mol. Syst. Biol. 2006;2:2006.0020. doi: 10.1038/msb4100065. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Xu J. Ai Y. Wang J. Xu J. Zhang Y. Yang D. Converting S-Limonene Synthase to Pinene or Phellandrene Synthases Reveals the Plasticity of the Active Site. Phytochemistry. 2017;137:34–41. doi: 10.1016/J.PHYTOCHEM.2017.02.017. [DOI] [PubMed] [Google Scholar]
- Mendez-Perez D. Alonso-Gutierrez J. Hu Q. Molinas M. Baidoo E. E. K. Wang G. Chan L. J. G. Adams P. D. Petzold C. J. Keasling J. D. Lee T. S. Production of Jet Fuel Precursor Monoterpenoids from Engineered Escherichia coli. Biotechnol. Bioeng. 2017;114(8):1703–1712. doi: 10.1002/bit.26296. [DOI] [PubMed] [Google Scholar]
- Zhang C. Chen X. Lee R. T. C. Rehka T. Maurer-Stroh S. Rühl M. Bioinformatics-Aided Identification, Characterization and Applications of Mushroom Linalool Synthases. Commun. Biol. 2021;4(1):1–11. doi: 10.1038/s42003-021-01715-z. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Zhang F. An T. Tang X. Zi J. Luo H. B. Wu R. Enzyme Promiscuity versus Fidelity in Two Sesquiterpene Cyclases (TEAS versus ATAS) ACS Catal. 2020;10(2):1470–1484. doi: 10.1021/ACSCATAL.9B05051/SUPPL_FILE/CS9B05051_SI_011.PDB. [DOI] [Google Scholar]
- Whitehead J. N. Leferink N. G. H. Johannissen L. O. Hay S. Scrutton N. S. Decoding Catalysis by Terpene Synthases. ACS Catal. 2023;13(19):12774–12802. doi: 10.1021/acscatal.3c03047. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Christianson D. W. Structural and Chemical Biology of Terpenoid Cyclases. Chem. Rev. 2017;117(17):11570. doi: 10.1021/ACS.CHEMREV.7B00287. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Rajaonarivony J. I. M. Gershenzon J. Miyazaki J. Croteau R. Evidence for an Essential Histidine Residue in 4S-Limonene Synthase and Other Terpene Cyclases. Arch. Biochem. Biophys. 1992;299(1):77–82. doi: 10.1016/0003-9861(92)90246-S. [DOI] [PubMed] [Google Scholar]
- Hyatt D. C. Croteau R. Mutational Analysis of a Monoterpene Synthase Reaction: Altered Catalysis through Directed Mutagenesis of (−)-Pinene Synthase from Abies grandis. Arch. Biochem. Biophys. 2005;439(2):222–233. doi: 10.1016/j.abb.2005.05.017. [DOI] [PubMed] [Google Scholar]
- Xu J. Xu J. Ai Y. Farid R. A. Tong L. Yang D. Mutational Analysis and Dynamic Simulation of S-Limonene Synthase Reveal the Importance of Y573: Insight into the Cyclization Mechanism in Monoterpene Synthases. Arch. Biochem. Biophys. 2018;638:27–34. doi: 10.1016/j.abb.2017.12.007. [DOI] [PubMed] [Google Scholar]
- Major D. T. Electrostatic Control of Chemistry in Terpene Cyclases. ACS Catal. 2017;7(8):5461–5465. doi: 10.1021/acscatal.7b01328. [DOI] [Google Scholar]
- Leferink N. G. H. Ranaghan K. E. Battye J. Johannissen L. O. Hay S. Kamp M. W. v. d. Mulholland A. J. Scrutton N. S. Taming the Reactivity of Monoterpene Synthases to Guide Regioselective Product Hydroxylation. ChemBioChem. 2020;21(7):985–990. doi: 10.1002/CBIC.201900672. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Leferink N. G. H. Escorcia A. M. Ouwersloot B. R. Johanissen L. O. Hay S. Kamp M. W. v. d. Scrutton N. S. Molecular Determinants of Carbocation Cyclisation in Bacterial Monoterpene Synthases. ChemBioChem. 2022;23(5):e202100688. doi: 10.1002/CBIC.202100688. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Alan Johnson L. Konrad Allemann R. Engineering Terpene Synthases and Their Substrates for the Biocatalytic Production of Terpene Natural Products and Analogues. Chem. Commun. 2025;61(12):2468–2483. doi: 10.1039/D4CC05785F. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Liang J. Wang D. Li X. Huang W. Xie C. Fu M. Zhang H. Meng Q. In Silico Genome-Wide Mining and Analysis of Terpene Synthase Gene Family in Hevea brasiliensis. Biochem. Genet. 2023;61(3):1185–1209. doi: 10.1007/s10528-022-10311-7. [DOI] [PubMed] [Google Scholar]
- Tararina M. A. Yee D. A. Tang Y. Christianson D. W. Structure of the Repurposed Fungal Terpene Cyclase FlvF Implicated in the C–N Bond-Forming Reaction of Flavunoidine Biosynthesis. Biochemistry. 2022;61(18):2014–2024. doi: 10.1021/acs.biochem.2c00335. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Priya P. Yadav A. Chand J. Yadav G. Terzyme: A Tool for Identification and Analysis of the Plant Terpenome. Plant Methods. 2018;14(1):4. doi: 10.1186/s13007-017-0269-0. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Tao H. Lauterbach L. Bian G. Chen R. Hou A. Mori T. Cheng S. Hu B. Lu L. Mu X. Li M. Adachi N. Kawasaki M. Moriya T. Senda T. Wang X. Deng Z. Abe I. Dickschat J. S. Liu T. Discovery of Non-Squalene Triterpenes. Nature. 2022;606(7913):414–419. doi: 10.1038/s41586-022-04773-3. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Abe T. Shiratori H. Kashiwazaki K. Hiasa K. Ueda D. Taniguchi T. Sato H. Abe T. Sato T. Structural-Model-Based Genome Mining Can Efficiently Discover Novel Non-Canonical Terpene Synthases Hidden in Genomes of Diverse Species. Chem. Sci. 2024;15(27):10402–10407. doi: 10.1039/D4SC01381F. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Samusevich R., Hebra T., Bushuiev R., Engst M., Kulhánek J., Bushuiev A., Smith J. D., Čalounová T., Smrčková H., Molineris M., Schwartz R., Tajovská A., Perković M., Chatpatanasiri R., Kampranis S. C., Major D. T., Sivic J. and Pluskal T., Structure-Enabled Enzyme Function Prediction Unveils Elusive Terpenoid Biosynthesis in Archaea, bioRxiv, 2025, preprint, 10.1101/2024.01.29.577750 [DOI]
- Salmon M. Laurendon C. Vardakou M. Cheema J. Defernez M. Green S. Faraldos J. A. O'Maille P. E. Emergence of Terpene Cyclization in Artemisia annua. Nat. Commun. 2015;6(1):1–10. doi: 10.1038/ncomms7143. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Jumper J. Evans R. Pritzel A. Green T. Figurnov M. Ronneberger O. Tunyasuvunakool K. Bates R. Žídek A. Potapenko A. Bridgland A. Meyer C. Kohl S. A. A. Ballard A. J. Cowie A. Romera-Paredes B. Nikolov S. Jain R. Adler J. Back T. Petersen S. Reiman D. Clancy E. Zielinski M. Steinegger M. Pacholska M. Berghammer T. Bodenstein S. Silver D. Vinyals O. Senior A. W. Kavukcuoglu K. Kohli P. Hassabis D. Highly Accurate Protein Structure Prediction with AlphaFold. Nature. 2021;596(7873):583–589. doi: 10.1038/s41586-021-03819-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Karuppiah V. Ranaghan K. E. Leferink N. G. H. Johannissen L. O. Shanmugam M. Cheallaigh A. N. Bennett N. J. Kearsey L. J. Takano E. Gardiner J. M. Kamp M. W. V. D. Hay S. Mulholland A. J. Leys D. Scrutton N. S. Structural Basis of Catalysis in the Bacterial Monoterpene Synthases Linalool Synthase and 1,8-Cineole Synthase. ACS Catal. 2017;7(9):6268–6282. doi: 10.1021/ACSCATAL.7B01924. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Aharoni A. Giri A. P. Verstappen F. W. A. Bertea C. M. Sevenier R. Sun Z. Jongsma M. A. Schwab W. Bouwmeester H. J. Gain and Loss of Fruit Flavor Compounds Produced by Wild and Cultivated Strawberry Species. Plant Cell. 2004;16(11):3110–3131. doi: 10.1105/tpc.104.023895. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Bohlmann J. Phillips M. Ramachandiran V. Katoh S. Croteau R. cDNA Cloning, Characterization, and Functional Expression of Four New Monoterpene Synthase Members of the Tpsd Gene Family from Grand Fir (Abies grandis) Arch. Biochem. Biophys. 1999;368(2):232–243. doi: 10.1006/abbi.1999.1332. [DOI] [PubMed] [Google Scholar]
- Byun-McKay A. Godard K.-A. Toudefallah M. Martin D. M. Alfaro R. King J. Bohlmann J. Plant A. L. Wound-Induced Terpene Synthase Gene Expression in Sitka Spruce That Exhibit Resistance or Susceptibility to Attack by the White Pine Weevil. Plant Physiol. 2006;140(3):1009–1021. doi: 10.1104/pp.105.071803. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Chen F. Tholl D. D'Auria J. C. Farooq A. Pichersky E. Gershenzon J. Biosynthesis and Emission of Terpenoid Volatiles from Arabidopsis Flowers. Plant Cell. 2003;15(2):481–494. doi: 10.1105/tpc.007989. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Chen Z. Vining K. J. Qi X. Yu X. Zheng Y. Liu Z. Fang H. Li L. Bai Y. Liang C. Li W. Lange B. M. Genome-Wide Analysis of Terpene Synthase Gene Family in Mentha longifolia and Catalytic Activity Analysis of a Single Terpene Synthase. Genes. 2021;12(4):518. doi: 10.3390/genes12040518. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Colby S. M. Alonso W. R. Katahira E. J. McGarvey D. J. Croteau R. 4S-Limonene Synthase from the Oil Glands of Spearmint (Mentha spicata). cDNA Isolation, Characterization, and Bacterial Expression of the Catalytically Active Monoterpene Cyclase. J. Biol. Chem. 1993;268(31):23016–23024. doi: 10.1016/S0021-9258(19)49419-2. [DOI] [PubMed] [Google Scholar]
- Crowell A. L. Williams D. C. Davis E. M. Wildung M. R. Croteau R. Molecular Cloning and Characterization of a New Linalool Synthase. Arch. Biochem. Biophys. 2002;405(1):112–121. doi: 10.1016/S0003-9861(02)00348-X. [DOI] [PubMed] [Google Scholar]
- Danner H. Boeckler G. A. Irmisch S. Yuan J. S. Chen F. Gershenzon J. Unsicker S. B. Köllner T. G. Four Terpene Synthases Produce Major Compounds of the Gypsy Moth Feeding-Induced Volatile Blend of Populus trichocarpa. Phytochemistry. 2011;72(9):897–908. doi: 10.1016/j.phytochem.2011.03.014. [DOI] [PubMed] [Google Scholar]
- Del Terra L. Lonzarich V. Asquini E. Navarini L. Graziosi G. Suggi Liverani F. Pallavicini A. Functional Characterization of Three Coffea arabica L. Monoterpene Synthases: Insights into the Enzymatic Machinery of Coffee Aroma. Phytochemistry. 2013;89:6–14. doi: 10.1016/j.phytochem.2013.01.005. [DOI] [PubMed] [Google Scholar]
- Falara V. Akhtar T. A. Nguyen T. T. H. Spyropoulou E. A. Bleeker P. M. Schauvinhold I. Matsuba Y. Bonini M. E. Schilmiller A. L. Last R. L. Schuurink R. C. Pichersky E. The Tomato Terpene Synthase Gene Family. Plant Physiol. 2011;157(2):770–789. doi: 10.1104/pp.111.179648. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Galata M. Sarker L. S. Mahmoud S. S. Transcriptome Profiling, and Cloning and Characterization of the Main Monoterpene Synthases of Coriandrum sativum L. Phytochemistry. 2014;102:64–73. doi: 10.1016/j.phytochem.2014.02.016. [DOI] [PubMed] [Google Scholar]
- Günnewich N. Page J. E. Köllner T. G. Degenhardt J. Kutchan T. M. Functional Expression and Characterization of Trichome-Specific (-)-Limonene Synthase and (+)-α-Pinene Synthase from Cannabis sativa. Nat. Prod. Commun. 2007;2(3):1934578X0700200301. doi: 10.1177/1934578X0700200301. [DOI] [Google Scholar]
- Hurd M. C. Kwon M. Ro D.-K. Functional Identification of a Lippia Dulcis Bornyl Diphosphate Synthase That Contains a Duplicated, Inhibitory Arginine-Rich Motif. Biochem. Biophys. Res. Commun. 2017;490(3):963–968. doi: 10.1016/j.bbrc.2017.06.147. [DOI] [PubMed] [Google Scholar]
- Iijima Y. Davidovich-Rikanati R. Fridman E. Gang D. R. Bar E. Lewinsohn E. Pichersky E. The Biochemical and Molecular Basis for the Divergent Patterns in the Biosynthesis of Terpenes and Phenylpropenes in the Peltate Glands of Three Cultivars of Basil. Plant Physiol. 2004;136(3):3724–3736. doi: 10.1104/pp.104.051318. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Jia J.-W. Crock J. Lu S. Croteau R. Chen X.-Y. (3R)-Linalool Synthase from Artemisia annua L.: cDNA Isolation, Characterization, and Wound Induction. Arch. Biochem. Biophys. 1999;372(1):143–149. doi: 10.1006/abbi.1999.1466. [DOI] [PubMed] [Google Scholar]
- Keeling C. I. Weisshaar S. Ralph S. G. Jancsik S. Hamberger B. Dullat H. K. Bohlmann J. Transcriptome Mining, Functional Characterization, and Phylogeny of a Large Terpene Synthase Gene Family in Spruce (Piceaspp.) BMC Plant Biol. 2011;11(1):43. doi: 10.1186/1471-2229-11-43. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Landmann C. Fink B. Festner M. Dregus M. Engel K.-H. Schwab W. Cloning and Functional Characterization of Three Terpene Synthases from Lavender (Lavandula angustifolia) Arch. Biochem. Biophys. 2007;465(2):417–429. doi: 10.1016/j.abb.2007.06.011. [DOI] [PubMed] [Google Scholar]
- Lin Y.-L. Lee Y.-R. Huang W.-K. Chang S.-T. Chu F.-H. Characterization of S-(+)-Linalool Synthase from Several Provenances of Cinnamomum osmophloeum. Tree Genet. Genomes. 2014;10(1):75–86. doi: 10.1007/s11295-013-0665-1. [DOI] [Google Scholar]
- Liu G. Yang M. Yang X. Ma X. Fu J. Five TPSs Are Responsible for Volatile Terpenoid Biosynthesis in Albizia julibrissin. J. Plant Physiol. 2021;258–259:153358. doi: 10.1016/J.JPLPH.2020.153358. [DOI] [PubMed] [Google Scholar]
- Lücker J. El Tamer M. K. Schwab W. Verstappen F. W. A. van der Plas L. H. W. Bouwmeester H. J. Verhoeven H. A. Monoterpene Biosynthesis in Lemon (Citrus limon) Eur. J. Biochem. 2002;269(13):3160–3171. doi: 10.1046/j.1432-1033.2002.02985.x. [DOI] [PubMed] [Google Scholar]
- Martin D. M. Fäldt J. Bohlmann J. Functional Characterization of Nine Norway Spruce TPS Genes and Evolution of Gymnosperm Terpene Synthases of the TPS-d Subfamily. Plant Physiol. 2004;135(4):1908–1927. doi: 10.1104/pp.104.042028. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Maruyama T. Ito M. Kiuchi F. Honda G. Molecular Cloning, Functional Expression and Characterization of d-Limonene Synthase from Schizonepeta tenuifolia. Biol. Pharm. Bull. 2001;24(4):373–377. doi: 10.1248/bpb.24.373. [DOI] [PubMed] [Google Scholar]
- Masumoto N. Korin M. Ito M. Geraniol and Linalool Synthases from Wild Species of Perilla. Phytochemistry. 2010;71(10):1068–1075. doi: 10.1016/j.phytochem.2010.04.006. [DOI] [PubMed] [Google Scholar]
- Morehouse B. R. Kumar R. P. Matos J. O. Olsen S. N. Entova S. Oprian D. D. Functional and Structural Characterization of a (+)-Limonene Synthase from Citrus sinensis. Biochemistry. 2017;56(12):1706–1715. doi: 10.1021/acs.biochem.7b00143. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Nagegowda D. A. Gutensohn M. Wilkerson C. G. Dudareva N. Two Nearly Identical Terpene Synthases Catalyze the Formation of Nerolidol and Linalool in Snapdragon Flowers. Plant J. 2008;55(2):224–239. doi: 10.1111/j.1365-313X.2008.03496.x. [DOI] [PubMed] [Google Scholar]
- Nawade B. Yahyaa M. Reuveny H. Shaltiel-Harpaz L. Eisenbach O. Faigenboim A. Bar-Yaakov I. Holland D. Ibdah M. Profiling of Volatile Terpenes from Almond (Prunus dulcis) Young Fruits and Characterization of Seven Terpene Synthase Genes. Plant Sci. 2019;287:110187. doi: 10.1016/j.plantsci.2019.110187. [DOI] [PubMed] [Google Scholar]
- Rajaonarivony J. I. M. Gershenzon J. Croteau R. Characterization and Mechanism of (4S)-Limonene Synthase, A Monoterpene Cyclase from the Glandular Trichomes of Peppermint (Mentha x Piperita) Arch. Biochem. Biophys. 1992;296(1):49–57. doi: 10.1016/0003-9861(92)90543-6. [DOI] [PubMed] [Google Scholar]
- Shimada T. Endo T. Fujii H. Omura M. Isolation and Characterization of a New d-Limonene Synthase Gene with a Different Expression Pattern in Citrus unshiu Marc. Sci. Hortic. 2005;105(4):507–512. doi: 10.1016/j.scienta.2005.02.009. [DOI] [Google Scholar]
- Sugiura M. Ito S. Saito Y. Niwa Y. Koltunw A. M. Sugimoto O. Sakai H. Molecular Cloning and Characterization of a Linalool Synthase from Lemon Myrtle. Biosci., Biotechnol., Biochem. 2011;75(7):1245–1248. doi: 10.1271/bbb.100922. [DOI] [PubMed] [Google Scholar]
- Yuan J. S. Köllner T. G. Wiggins G. Grant J. Degenhardt J. Chen F. Molecular and Genomic Basis of Volatile-Mediated Indirect Defense against Insects in Rice. Plant J. Cell Mol. Biol. 2008;55(3):491–503. doi: 10.1111/j.1365-313X.2008.03524.x. [DOI] [PubMed] [Google Scholar]
- Yuba A. Yazaki K. Tabata M. Honda G. Croteau R. cDNA Cloning, Characterization, and Functional Expression of 4S-(−)-Limonene Synthase from Perilla frutescens. Arch. Biochem. Biophys. 1996;332(2):280–287. doi: 10.1006/abbi.1996.0343. [DOI] [PubMed] [Google Scholar]
- Yue Y. Yu R. Fan Y. Characterization of Two Monoterpene Synthases Involved in Floral Scent Formation in Hedychium coronarium. Planta. 2014;240(4):745–762. doi: 10.1007/s00425-014-2127-x. [DOI] [PubMed] [Google Scholar]
- Dudareva N. Cseke L. Blanc V. M. Pichersky E. Evolution of Floral Scent in Clarkia: Novel Patterns of S-Linalool Synthase Gene Expression in the C. breweri Flower. Plant Cell. 1996;8(7):1137–1148. doi: 10.1105/tpc.8.7.1137. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Adal A. M. Sarker L. S. Malli R. P. N. Liang P. Mahmoud S. S. RNA-Seq in the Discovery of a Sparsely Expressed Scent-Determining Monoterpene Synthase in Lavender (Lavandula) Planta. 2019;249(1):271–290. doi: 10.1007/s00425-018-2935-5. [DOI] [PubMed] [Google Scholar]
- Huang K.-F. Wen C.-H. Lee Y.-R. Chu F.-H. Cloning and Characterization of Terpene Synthase Genes from Taiwan Cherry. Tree Genet. Genomes. 2019;15(4):51. doi: 10.1007/s11295-019-1355-4. [DOI] [Google Scholar]
- Huang K.-F. Lee Y.-R. Tseng Y.-H. Wang S.-Y. Chu F.-H. Cloning and Functional Characterization of a Monoterpene Synthase Gene from Eleutherococcus trifoliatus. Holzforschung. 2015;69(2):163–171. doi: 10.1515/hf-2014-0078. [DOI] [Google Scholar]
- Jongedijk E. Cankar K. Ranzijn J. van der Krol S. Bouwmeester H. Beekwilder J. Capturing of the Monoterpene Olefin Limonene Produced in Saccharomyces cerevisiae. Yeast. 2015;32(1):159–171. doi: 10.1002/yea.3038. [DOI] [PubMed] [Google Scholar]
- Hattan J. Shindo K. Ito T. Shibuya Y. Watanabe A. Tagaki C. Ohno F. Sasaki T. Ishii J. Kondo A. Misawa N. Identification of a Novel Hedycaryol Synthase Gene Isolated from Camellia brevistyla Flowers and Floral Scent of Camellia cultivars. Planta. 2016;243(4):959–972. doi: 10.1007/s00425-015-2454-6. [DOI] [PubMed] [Google Scholar]
- Yang Z. Li Y. Gao F. Jin W. Li S. Kimani S. Yang S. Bao T. Gao X. Wang L. MYB21 Interacts with MYC2 to Control the Expression of Terpene Synthase Genes in Flowers of Freesia hybrida and Arabidopsis thaliana. J. Exp. Bot. 2020;71(14):4140–4158. doi: 10.1093/jxb/eraa184. [DOI] [PubMed] [Google Scholar]
- Magnard J.-L. Bony A. R. Bettini F. Campanaro A. Blerot B. Baudino S. Jullien F. Linalool and Linalool Nerolidol Synthases in Roses, Several Genes for Little Scent. Plant Physiol. Biochem. 2018;127:74–87. doi: 10.1016/j.plaphy.2018.03.009. [DOI] [PubMed] [Google Scholar]
- Hattan J. Shindo K. Sasaki T. Misawa N. Isolation and Functional Characterization of New Terpene Synthase Genes from Traditional Edible Plants. J. Oleo Sci. 2018;67(10):1235–1246. doi: 10.5650/jos.ess18163. [DOI] [PubMed] [Google Scholar]
- Liu M. Liu X. Hu J. Xue Y. Zhao X. Genetic Diversity of Limonene Synthase Genes in Rongan Kumquat (Fortunella crassifolia) Funct. Plant Biol. 2020;47(5):425–439. doi: 10.1071/FP19051. [DOI] [PubMed] [Google Scholar]
- Zager J. J. Lange I. Srividya N. Smith A. Lange B. M. Gene Networks Underlying Cannabinoid and Terpenoid Accumulation in Cannabis. Plant Physiol. 2019;180(4):1877–1897. doi: 10.1104/pp.18.01506. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Zhou Y. Deng R. Xu X. Yang Z. Enzyme Catalytic Efficiencies and Relative Gene Expression Levels of (R)-Linalool Synthase and (S)-Linalool Synthase Determine the Proportion of Linalool Enantiomers in Camellia sinensis var. sinensis. J. Agric. Food Chem. 2020;68(37):10109–10117. doi: 10.1021/acs.jafc.0c04381. [DOI] [PubMed] [Google Scholar]
- Nawade B. Shaltiel-Harpaz L. Yahyaa M. Kabaha A. Kedoshim R. Bosamia T. C. Ibdah M. Characterization of Terpene Synthase Genes Potentially Involved in Black Fig Fly (Silba adipata) Interactions with Ficus carica. Plant Sci. 2020;298:110549. doi: 10.1016/j.plantsci.2020.110549. [DOI] [PubMed] [Google Scholar]
- Yang P. Zhao H.-Y. Wei J.-S. Zhao Y.-Y. Lin X.-J. Su J. Li F.-P. Li M. Ma D.-M. Tan X.-K. Liang H.-L. Sun Y.-W. Zhan R.-T. He G.-Z. Zhou X.-F. Yang J.-F. Chromosome-Level Genome Assembly and Functional Characterization of Terpene Synthases Provide Insights into the Volatile Terpenoid Biosynthesis of Wurfbainia villosa. Plant J. 2022;112(3):630–645. doi: 10.1111/tpj.15968. [DOI] [PubMed] [Google Scholar]
- Chen R. Hu J. Pan Y. Song H. Xian X. Isolation and Expression Analysis of OfLIS Gene in Osmanthus fragrans. Int. J. Agric. Biol. 2020;24(6):1551–1557. doi: 10.17957/IJAB/15.1594. [DOI] [Google Scholar]
- Chen X. Yauk Y.-K. Nieuwenhuizen N. J. Matich A. J. Wang M. Y. Perez R. L. Atkinson R. G. Beuning L. L. Characterisation of an (S)-Linalool Synthase from Kiwifruit (Actinidia arguta) That Catalyses the First Committed Step in the Production of Floral Lilac Compounds. Funct. Plant Biol. 2010;37(3):232–243. doi: 10.1071/FP09179. [DOI] [Google Scholar]
- Cao Y. Hu S. Dai Q. Liu Y. Tomato Terpene Synthases TPS5 and TPS39 Account for a Monoterpene Linalool Production in Tomato Fruits. Biotechnol. Lett. 2014;36(8):1717–1725. doi: 10.1007/s10529-014-1533-2. [DOI] [PubMed] [Google Scholar]
- Feng L. Chen C. Li T. Wang M. Tao J. Zhao D. Sheng L. Flowery Odor Formation Revealed by Differential Expression of Monoterpene Biosynthetic Genes and Monoterpene Accumulation in Rose (Rosa rugosa Thunb.) Plant Physiol. Biochem. 2014;75:80–88. doi: 10.1016/j.plaphy.2013.12.006. [DOI] [PubMed] [Google Scholar]
- Zhu B.-Q. Cai J. Wang Z.-Q. Xu X.-Q. Duan C.-Q. Pan Q.-H. Identification of a Plastid-Localized Bifunctional Nerolidol/Linalool Synthase in Relation to Linalool Biosynthesis in Young Grape Berries. Int. J. Mol. Sci. 2014;15(12):21992–22010. doi: 10.3390/ijms151221992. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Ito M. Kiuchi F. Yang L. L. Honda G. Perilla citriodora from Taiwan and Its Phytochemical Characteristics. Biol. Pharm. Bull. 2000;23(3):359–362. doi: 10.1248/bpb.23.359. [DOI] [PubMed] [Google Scholar]
- Lu X. Bai J. Tian Z. Li C. Ahmed N. Liu X. Cheng J. Lu L. Cai J. Jiang H. Wang W. Cyclization Mechanism of Monoterpenes Catalyzed by Monoterpene Synthases in Dipterocarpaceae. Synth. Syst. Biotechnol. 2024;9(1):11–18. doi: 10.1016/j.synbio.2023.11.009. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Le Guilloux V. Schmidtke P. Tuffery P. Fpocket: An Open Source Platform for Ligand Pocket Detection. BMC Bioinf. 2009;10(1):168. doi: 10.1186/1471-2105-10-168. [DOI] [PMC free article] [PubMed] [Google Scholar]
- The UniProt Consortium UniProt: The Universal Protein Knowledgebase in 2025. Nucleic Acids Res. 2025;53(D1):D609–D617. doi: 10.1093/nar/gkae1010. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Livingstone C. D. Barton G. J. Protein Sequence Alignments: A Strategy for the Hierarchical Analysis of Residue Conservation. CABIOS, Comput. Appl. Biosci. 1993;9(6):745–756. doi: 10.1093/BIOINFORMATICS/9.6.745. [DOI] [PubMed] [Google Scholar]
- Oken A. C. Turcu A. L. Tzortzini E. Georgiou K. Nagel J. Westermann F. G. Barniol-Xicota M. Seidler J. Kim G.-R. Lee S.-D. Nicke A. Kim Y.-C. Müller C. E. Kolocouris A. Vázquez S. Mansoor S. E. A Polycyclic Scaffold Identified by Structure-Based Drug Design Effectively Inhibits the Human P2X7 Receptor. Nat. Commun. 2025;16(1):8283. doi: 10.1038/s41467-025-62643-8. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Arp G. Jiang A. K. Dufault-Thompson K. Levy S. Zhong A. Wassan J. T. Grant M. R. Li Y. Hall B. Jiang X. Identification of Gut Bacteria Reductases That Biotransform Steroid Hormones. Nat. Commun. 2025;16(1):6285. doi: 10.1038/s41467-025-61425-6. [DOI] [PMC free article] [PubMed] [Google Scholar]
- He J. Chai X. Zhang Q. Wang Y. Wang Y. Yang X. Wu J. Feng B. Sun J. Rui W. Ze S. Fu Y. Zhao Y. Zhang Y. Zhang Y. Liu M. Liu C. She M. Hu X. Ma X. Yang H. Li D. Zhao S. Li G. Zhang Z. Tian Z. Ma Y. Cao L. Yi B. Li D. Nussinov R. Eng C. Chan T. A. Ruppin E. Gutkind J. S. Cheng F. Liu M. Lu W. The Lactate Receptor HCAR1 Drives the Recruitment of Immunosuppressive PMN-MDSCs in Colorectal Cancer. Nat. Immunol. 2025;26(3):391–403. doi: 10.1038/s41590-024-02068-5. [DOI] [PubMed] [Google Scholar]
- Chen T. and Guestrin C., XGBoost: A Scalable Tree Boosting System, in Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (KDD '16), Association for Computing Machinery, New York, NY, USA, 2016, pp. 785–794. 10.1145/2939672.2939785 [DOI] [Google Scholar]
- Lundberg S. M. and Lee S.-I., A Unified Approach to Interpreting Model Predictions, in Advances in Neural Information Processing Systems, Curran Associates, Inc., 2017, vol. 30 [Google Scholar]
- Durairaj J. Di Girolamo A. Bouwmeester H. J. de Ridder D. Beekwilder J. van Dijk A. D. J. An Analysis of Characterized Plant Sesquiterpene Synthases. Phytochemistry. 2019;158:157–165. doi: 10.1016/j.phytochem.2018.10.020. [DOI] [PubMed] [Google Scholar]
- Camacho C. Coulouris G. Avagyan V. Ma N. Papadopoulos J. Bealer K. Madden T. L. BLAST+: Architecture and Applications. BMC Bioinf. 2009;10(1):421. doi: 10.1186/1471-2105-10-421. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Eddy S. R. Accelerated Profile HMM Searches. PLoS Comput. Biol. 2011;7(10):e1002195. doi: 10.1371/journal.pcbi.1002195. [DOI] [PMC free article] [PubMed] [Google Scholar]
- van Kempen M. Kim S. S. Tumescheit C. Mirdita M. Lee J. Gilchrist C. L. M. Söding J. Steinegger M. Fast and Accurate Protein Structure Search with Foldseek. Nat. Biotechnol. 2024;42(2):243–246. doi: 10.1038/s41587-023-01773-0. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Sievers F. Wilm A. Dineen D. Gibson T. J. Karplus K. Li W. Lopez R. McWilliam H. Remmert M. Söding J. Thompson J. D. Higgins D. G. Fast, Scalable Generation of High-quality Protein Multiple Sequence Alignments Using Clustal Omega. Mol. Syst. Biol. 2011;7(1):539. doi: 10.1038/msb.2011.75. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Field C. R., ColabAlign, 2024, 10.5281/zenodo.14169502 [DOI]
- Letunic I. Bork P. Interactive Tree of Life (iTOL) v6: Recent Updates to the Phylogenetic Tree Display and Annotation Tool. Nucleic Acids Res. 2024;52(W1):W78–W82. doi: 10.1093/nar/gkae268. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Aldas-Bulos V. D. Plisson F. Benchmarking Protein Structure Predictors to Assist Machine Learning-Guided Peptide Discovery. Digital Discovery. 2023;2(4):981–993. doi: 10.1039/D3DD00045A. [DOI] [Google Scholar]
- Saldaño T. Escobedo N. Marchetti J. Zea D. J. Mac Donagh J. Velez Rueda A. J. Gonik E. García Melani A. Novomisky Nechcoff J. Salas M. N. Peters T. Demitroff N. Fernandez Alberti S. Palopoli N. Fornasari M. S. Parisi G. Impact of Protein Conformational Diversity on AlphaFold Predictions. Bioinformatics. 2022;38(10):2742–2748. doi: 10.1093/bioinformatics/btac202. [DOI] [PubMed] [Google Scholar]
- Loizzi M. González V. Miller D. J. Allemann R. K. Nucleophilic Water Capture or Proton Loss: Single Amino Acid Switch Converts δ-Cadinene Synthase into Germacradien-4-Ol Synthase. ChemBioChem. 2018;19(1):100–105. doi: 10.1002/cbic.201700531. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Monera O. D. Sereda T. J. Zhou N. E. Kay C. M. Hodges R. S. Relationship of Sidechain Hydrophobicity and α-Helical Propensity on the Stability of the Single-Stranded Amphipathic α-Helix. J. Pept. Sci. 1995;1(5):319–329. doi: 10.1002/psc.310010507. [DOI] [PubMed] [Google Scholar]
- López-Gallego F. Wawrzyn G. T. Schmidt-Dannert C. Selectivity of Fungal Sesquiterpene Synthases: Role of the Active Site's H-1α Loop in Catalysis. Appl. Environ. Microbiol. 2010;76(23):7723–7733. doi: 10.1128/AEM.01811-10. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Whitehead J. N. Leferink N. G. H. Komati Reddy G. Levy C. W. Hay S. Takano E. Scrutton N. S. How a 10-Epi-Cubebol Synthase Avoids Premature Reaction Quenching to Form a Tricyclic Product at High Purity. ACS Catal. 2022;12(19):12123–12131. doi: 10.1021/acscatal.2c03155. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Leferink N. G. H. Dunstan M. S. Hollywood K. A. Swainston N. Currin A. Jervis A. J. Takano E. Scrutton N. S. An Automated Pipeline for the Screening of Diverse Monoterpene Synthase Libraries. Sci. Rep. 2019;9(1):11936. doi: 10.1038/s41598-019-48452-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Kumar R. P. Morehouse B. R. Matos J. O. Malik K. Lin H. Krauss I. J. Oprian D. D. Structural Characterization of Early Michaelis Complexes in the Reaction Catalyzed by (+)-Limonene Synthase from Citrus Sinensis Using Fluorinated Substrate Analogues. Biochemistry. 2017;56(12):1716–1725. doi: 10.1021/acs.biochem.7b00144. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Chen C. Zhang Q. Yu B. Yu Z. Lawrence P. J. Ma Q. Zhang Y. Improving Protein-Protein Interactions Prediction Accuracy Using XGBoost Feature Selection and Stacked Ensemble Classifier. Comput. Biol. Med. 2020;123:103899. doi: 10.1016/j.compbiomed.2020.103899. [DOI] [PubMed] [Google Scholar]
- Zhong J. Sun Y. Peng W. Xie M. Yang J. Tang X. XGBFEMF: An XGBoost-Based Framework for Essential Protein Prediction. IEEE Trans. NanoBiosci. 2018;17(3):243–250. doi: 10.1109/TNB.2018.2842219. [DOI] [PubMed] [Google Scholar]
- Sena J. Johannissen L. O. Blaker J. J. Hay S. A Machine Learning Model for the Prediction of Water Contact Angles on Solid Polymers. J. Phys. Chem. B. 2025;129(10):2739–2745. doi: 10.1021/acs.jpcb.4c06608. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Rudin C. Stop Explaining Black Box Machine Learning Models for High Stakes Decisions and Use Interpretable Models Instead. Nature Machine Intelligence. 2019;1(5):206–215. doi: 10.1038/s42256-019-0048-x. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Luntz A. Brailovsky V. On Estimation of Characters Obtained in Statistical Procedure of Recognition. Tekhnicheskaya Kibernetika. 1969:6–12. [Google Scholar]
- Lachenbruch P. A. Mickey M. R. Estimation of Error Rates in Discriminant Analysis. Technometrics. 1968;10(1):1–11. doi: 10.2307/1266219. [DOI] [Google Scholar]
- Raschka S., Model Evaluation, Model Selection, and Algorithm Selection in Machine Learning, arXiv 2020, preprint, arXiv:1811.12808, 10.48550/arXiv.1811.12808. [DOI]
- Swan A. L. Mobasheri A. Allaway D. Liddell S. Bacardit J. Application of Machine Learning to Proteomics Data: Classification and Biomarker Identification in Postgenomics Biology. OMICS: J. Integr. Biol. 2013;17(12):595–610. doi: 10.1089/omi.2013.0017. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Goldman S. Das R. Yang K. K. Coley C. W. Machine Learning Modeling of Family Wide Enzyme-Substrate Specificity Screens. PLoS Comput. Biol. 2022;18(2):e1009853. doi: 10.1371/journal.pcbi.1009853. [DOI] [PMC free article] [PubMed] [Google Scholar]
- De Ferrari L. Mitchell J. B. From Sequence to Enzyme Mechanism Using Multi-Label Machine Learning. BMC Bioinf. 2014;15(1):150. doi: 10.1186/1471-2105-15-150. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Liang W. Zheng S. Shu Y. Huang J. Bayesian Optimization-Assisted Engineering of Formate Dehydrogenase Encapsulation in Multivariate Zeolitic Imidazolate Framework. Chem. Mater. 2025;37(1):429–440. doi: 10.1021/acs.chemmater.4c02816. [DOI] [Google Scholar]
- Yang K. Liu L. Wen Y. The Impact of Bayesian Optimization on Feature Selection. Sci. Rep. 2024;14(1):3948. doi: 10.1038/s41598-024-54515-w. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Shen X. Zhang S. Long J. Chen C. Wang M. Cui Z. Chen B. Tan T. A Highly Sensitive Model Based on Graph Neural Networks for Enzyme Key Catalytic Residue Prediction. J. Chem. Inf. Model. 2023;63(14):4277–4290. doi: 10.1021/acs.jcim.3c00273. [DOI] [PubMed] [Google Scholar]
- Engst M. Brokeš M. Čalounová T. Tajovská A. Perković M. Samusevich R. Chatpatanasiri R. Pluskal T. MARTS-DB: A Database of Mechanisms and Reactions of Terpene Synthases. BMC Bioinf. 2025;27(1):10. doi: 10.1186/s12859-025-06341-8. [DOI] [PMC free article] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Data Availability Statement
All Python code, AlphaFold structures, Fpocket output and ML training and evaluation data used in this paper is freely available at https://github.com/cathaloraghallaigh/ATC and at ZENDO (https://doi.org/10.5281/zenodo.22877417). All additional data supporting this study are provided in the supplementary information (SI) accompanying this paper. Supplementary information: enzyme datasets and feature definitions, benchmarking of AlphaFold predictions against crystal structures, active-site residue analyses, and sensitivity analyses of structural model choice and residue mapping. It also provides machine-learning algorithm comparisons, cross-validation results, independent validation, and further testing on monoterpene synthases with other products and on sesquiterpene synthases. See DOI: https://doi.org/10.1039/d6dd00312e.
