Abstract
We introduce Finenzyme, a Protein Language Model (PLM) that employs a multifaceted learning strategy based on transfer learning from a decoder-based Transformer, conditional learning using specific functional keywords, and fine-tuning for the in silico modeling of enzymes. Our experiments show that Finenzyme significantly enhances generalist PLMs like ProGen for the in silico prediction and generation of enzymes belonging to specific Enzyme Commission (EC) categories. Our in silico experiments demonstrate that Finenzyme generated sequences can diverge from natural ones, while retaining similar predicted tertiary structure, predicted functions and the active sites of their natural counterparts. We show that embedded representations of the generated sequences obtained from the embeddings computed by both Finenzyme and ESMFold closely resemble those of natural ones, thus making them suitable for downstream tasks, including e.g. EC classification. Clustering analysis based on the primary and predicted tertiary structure of sequences reveals that the generated enzymes form clusters that largely overlap with those of natural enzymes. These overall in silico validation experiments indicate that Finenzyme effectively captures the structural and functional properties of target enzymes, and can in perspective support targeted enzyme engineering tasks.
Keywords: Large Language Models, Protein Language Models, Fine-tuning of Large Language Models, Conditional Transformers, In silico enzyme design and modeling
Graphical abstract

Highlights
-
•
A new PLM model, Finenzyme, is proposed, based on transfer learning, conditional learning and fine-tuning.
-
•
Fine-tuning boosts prediction and generation of specific Enzyme Commission (EC) categories.
-
•
Fine-tuned models generate sequences close to natural ones, using only keywords associated with each EC category.
-
•
The predicted tertiary structure and functions of the generated Finenzyme sequences are close to that of natural enzymes.
-
•
We show some examples of application of Finenzyme for supporting the design of novel, functionally characterized enzymes.
1. Introduction
Large Language Models (LLMs) have revolutionized biomedical domains, ranging from clinical decision support, to differential diagnosis of diseases, drug design and protein modeling [1], [2], [3], [4]. Protein Language Models (PLMs) leverage Transformer architectures [5] to capture the “language” of proteins [6] by learning complex distributions of amino acid sequences from massive datasets such as UniParc and UniProt [7], [8], [9]. These models enable applications ranging from secondary structure prediction to long range contacts and mutational effects, functional annotation and many others [10], [11], [12], [13], [14].
Conditional learning further extends PLM capabilities by guiding sequence generation through the addition of keywords that denote specific functional attributes. This strategy enables the design of enzyme sequences with tailored properties [14].
While many PLM models are employed in zero-shot scenarios—where the model generates or embeds sequences based solely on its pre-existing knowledge—certain applications require fine-tuning to enhance the performance of pre-trained autoregressive PLMs for specific tasks. Fine-tuning primarily involves adapting the model's pre-trained weights through an additional training phase on data tailored to a particular task. In some cases, architectural modifications are made before retraining, such as adding new neural layers or reducing the size of existing ones [15], [16], [17].
Several studies have demonstrated the effectiveness of fine-tuning in the context of PLMs. For instance, parameter-efficient fine-tuning of PLMs improved protein-protein interactions [18]. Analogously, fine-tuning of pre-trained antibody language models boosted the binding antigen specificity predictions to SARS-CoV2 spike protein and influenza hemagglutinin [19]. Fine-tuning PLMs with deep mutational scanning improved variant effect prediction [20]. In a different context, fine-tuning the ProtBert-BFD and Prot5-XL-Uniref50 Transformers [21] led to a significant improvement in Gene Ontology and Enzyme Commission (EC) number predictions [22], close to the top-model results of the CAFA3 challenge [23]. Fine-tuning ESM2 and ProstT5 [24] models boosted predictions in various tasks, ranging from disorder and mutation effects prediction to sub-cellular location prediction. The combination of a pre-trained encoder with data expansion enabled geometric learning-based models to improve the prediction of protein stability changes upon mutations [25]. Our recent results on the fine-tuning of the ProGen model on the lysozyme family of proteins achieved a statistically significant improvement in terms of accuracy and perplexity [26], confirming previous results on a similar task [14].
However, fine-tuning does not always yield improvements. In the context of biomedical NLP, it has been shown that fine-tuning stability and performance depends on the characteristics of the pre-trained model, whose performance cannot always be improved [27]. In the context of PLMs, no statistically significant improvement has been achieved by fine-tuning PLMs in several tasks, such as secondary structure prediction [28].
Given such contradictory results, in this work we aim at investigating the conditions under which fine-tuning improves the performance of pre-trained PLMs applied to enzyme prediction and generation. Enzymes, accelerating chemical reactions by orders of magnitude, play a key role in cellular biology and span fundamental medical and industrial applications, ranging from pharmacology to the reduction of environmental pollution [29], [30].
To explore the impact of fine-tuning in this context we build on ProGen [14], a PLM trained on 280 million protein sequences, and we propose a dual learning strategy: a) PLM conditional learning and b) fine-tuning of general models pre-trained on a large corpus of proteins, not limited to enzymes only. From this standpoint, our approach follows a different learning strategy than the recently proposed specialized pre-training of conditional Transformers on natural enzymes downloaded from the BRENDA database [31]. Indeed, our proposed Finenzyme models combine conditional Transformers pre-trained on large corpora of proteins with well-focused fine-tuning on specific enzymatic categories to generate sequences whose predicted properties on the one hand are close to that of natural enzymes, and on the other hand explore avenues not probed by natural evolution. Besides these results, we analyze and motivate the conditions under which fine-tuning by conditional transfer-learning is not effective. Fig. 1 provides an overview of Finenzyme computational experiments and of the computational procedures and tools we used to evaluate and characterize the generated enzyme sequences.
Fig. 1.
High-level scheme of Finenzyme and summary of the main computational tasks performed in this work. Finenzyme is fine-tuned on EC categories using a conditional learning approach. Fine-tuned models predictions are evaluated on separated test sets (Enzyme sequence prediction). Finenzyme embedded representations of enzymes are visualized through Umap projections and used to classify EC categories. Fine tuned models generate new enzyme sequences: their functions are in silico predicted through CLEAN, and their ESMfold 3D structures are compared and clustered with natural ones using Foldseek. The primary structure of generated enzymes is also compared with natural ones, and their kinetic parameters are computationally predicted. Finenzyme can be also applied to the in silico selection of candidate sequences, to support the design of functionally characterized enzymes.
Our main contributions can be summarized as follows:
-
•
We propose Finenzyme, a protein language model based on a multi-faceted strategy spanning transfer learning, conditional learning and fine-tuning.
-
•
The conditions under which fine-tuning can improve the prediction and generation of enzyme sequences are analyzed, showing that fine-tuning boosts prediction and generation of specific EC categories.
-
•
Fine-tuned models generate sequences close to natural ones, using only keywords associated with each EC category.
-
•
The predicted tertiary structure and functions of the generated Finenzyme sequences are close to that of natural enzymes and active sites of natural enzymes are highly conserved in Finenzyme generated sequences. Additionally, we show some examples of application of Finenzyme for supporting the design of novel, functionally characterized enzymes.
-
•
Source code is available to reproduce all experiments, scripts, and tutorials to allow users to fine-tune Finenzyme on any EC category or groups of functionally related EC categories.
2. Materials and methods
We first introduce the enzyme datasets used in our experiments (Section 2.1), then we outline the main characteristics of conditional Transformers for protein modeling (Section 2.2). Section 2.3 details the fine-tuning and conditional learning techniques of Finenzyme, and Section 2.4 its evaluation procedures. Then we describe in which way we performed the generation of novel sequences using the fine-tuned models (Section 2.5) and the procedures to evaluate their quality (Section 2.6). Section 2.7 provides the in silico EC classification of Finenzyme generated sequences and Section 2.8 the prediction of their kinetic parameters. Finally the procedures for the visualization of the embedded representation of the natural and Finenzyme generated sequences obtained with Finenzyme and ESMFold are summarized in Section 2.9 and the clustering procedures to compare natural and generated sequences are described in Section 2.10.
Fig. 1 provides a schematic summary of the main computational tasks performed in this work.
2.1. Enzyme datasets
The Enzyme Commission (EC) taxonomy[32] is a classification system for enzymes, based on the chemical reactions they catalyze. This system categorizes enzymes into four numerical levels, forming a unique EC number for each type of enzyme-catalyzed reaction. For instance, the enzyme class 3.2.1.4 refers to a cellulase, i.e. enzymes that catalyze the reaction of endohydrolysis of (1-4)-beta-D-glucosidic linkages in cellulose, lichenin, and cereal beta-D-glucans. The first number, 3. refers to the general class of the catalyzed reactions: in this case hydrolases that catalyze the hydrolytic cleavage of C-O, C-N, C-C, and some other bonds. The second number (3.2) refers to enzymes that hydrolyze glycosyl compounds. The third number (3.2.1, glycosidases) focuses on specific glycosil compounds on which the enzyme acts, i.e. hydrolyzing O- and S-glycosyl compounds and, finally, 3.2.1.4 refers to cellulases.
The enzyme datasets used in our prediction experiments include enzyme sequence datasets downloaded from UniProt [9]. The considered sequences were filtered to include only those with lengths between 10 and 500 nucleotides (sequences shorter than 10 nucleotides were considered partial reads, and the upper limit of 500 was chosen based on the model's maximum input size). Duplicate sequences were removed to avoid redundancy, and the remaining sequences were split according to their EC classes. In more detail, we considered both top-level general EC classes and low-level specific EC classes. The top-level classes represent broad categories of enzyme functions, such as oxidoreductases (EC1), transferases (EC2), hydrolases (EC 3), lyases (EC 4), isomerases (EC 5), ligases (EC 6), and translocases (EC7). Each top-level class is further divided into subclasses and sub-subclasses, leading to specific EC numbers that denote particular enzyme-catalyzed reactions. We considered various low-level EC classes, namely, alcohol dehydrogenase (EC 1.1.1.1), DNA methyltransferase (EC 2.1.1.37), hydrolases acting on c-halide compounds (EC 3.8.1), cellulase (EC 3.2.1.4), ribulose-bisphosphate carboxylase (EC 4.1.1.39), chorismate mutase (EC 5.4.99.5), biotin ligase (EC 6.3.4.15), and proton-translocating transhydrogenase (EC 7.2.1.1).
The sequences in top-level classes were further filtered to maintain reviewed-only sequences, i.e., high-quality, manually annotated and non-redundant protein sequences from the UniProtKB/Swiss-Prot database (for details about non-redundant sequences see https://www.uniprot.org/help/redundancy). This filter was not applied to the most specific EC categories, due to their small cardinality. In other words, the sequences for the most specific EC classes were retrieved from the UniProtKB/TrEMBL (unreviewed) sequence database that contains protein sequences associated with computationally generated annotation and large-scale functional characterization. Supplementary Section S1 and Table S1 in Supplementary Information summarize the main characteristics of the EC categories and of the enzyme datasets included in our experiments.
2.2. Conditional Transformers for protein learning and generation
In this work we build upon a pre-trained conditional model, ProGen [14]; it was obtained by adapting the CTRL conditional Transformer architecture [33], originally presented to generate text, to the generation of functionally characterized amino acid sequences. While CTRL used pre-pended tags to guide the generation of specific text (Fig. S3 in Supplementary Information), ProGen uses functional tags to guide the autoregressive generation of amino acid sequences. These functional tags denote a protein family, a Gene Ontology term, or other properties of the protein (more details about the tags used by ProGen are reported in Section 2.3).
More precisely, given a training instance, i.e. a protein sequence and its related keywords t, where , represent the amino acids, the model learns through back-propagation the conditioned probability , that can be factorized using the chain rule of probability:
| (1) |
where denotes the conditional probability of given all preceding elements and the functional tags t.
This formulation breaks down protein language modeling into a next-amino acid prediction task. Consequently, the pre-trained model with parameters θ can be trained to minimize the negative log-likelihood over a dataset of sequences :
| (2) |
Thus, by acquiring knowledge about the conditional probability distribution, the protein language model can generate new sequences of length m by sequentially sampling its components: , , …, .
The ProGen architecture features an internal embedding dimension of , and an inner dimension of , 36 layers, and 16 attention heads per layer. Each layer includes dropout with a probability of 0.1 following the residual connections, and token embeddings are tied to the final output layer embeddings. Fig. S3 (Supplementary Information) provides a high level scheme of the conditional Transformer we used in the experiments.
2.3. Fine-tuning of conditional Transformers
The goal of the fine-tuning strategy we applied is to leverage the knowledge acquired from millions of protein sequences and transfer the “learned knowledge” encapsulated in the weights of the pre-trained ProGen model into a new model, fine-tuned to learn the “language” of a specific EC category of enzymes.
To achieve this, we fine-tuned the ProGen model using conditional transfer learning to guide the model in learning the “language” necessary for generating enzymes within the considered EC classes.
Instead of applying architectural modifications [15] - which could introduce challenges related to weight initialization and require computationally expensive training, our conditional transfer learning strategy was implemented through a strategic redesign of the vocabulary tags used by ProGen, followed by a few epochs of conditional fine-tuning. In more detail, the ProGen tag-encoding vocabulary consists of 129,406 tags, encoded as numerical values (see Table 1, central column):
-
1.Functional category tags (range: [0–129,380]) consist of:
-
2.
Amino acid tags [129,381–129,405]: The next 24 tags represent the starting amino acids, using IUPAC encoding [36].
-
3.
Padding token: The final tag encodes a special PAD token used to pad sequences to a uniform length.
Table 1.
Vocabulary changes in the encoding for tags and amino acids; Finenzyme avoids using codes from k + 1 to 129,380.
| Encoding | ProGen Tag | Finenzyme Tag |
|---|---|---|
| 0 to k − 1 | UniProt tags | k EC classes |
| k | stop token | |
| k + 1 to 1,164 | not used | |
| 1,165 to 129,380 | Taxonomy NCBI IDs | not used |
| 129,381 to 129,405 | Amino acids IUPAC | Amino acids IUPAC |
| 129,406 | Pad Token | Pad Token |
We modified this vocabulary to represent the EC classes that the LLM must learn (see Table 1, right column). Given k EC classes, we reassign the first k numbers (0 to ) to these EC classes for the specific generation task. Additionally, we introduce a stop token (encoded as k) to mark the end of a sequence and allow generating sequences shorter than 512 amino acids (see section “Generation of new enzyme sequences”). The remaining UniProt and NCBI Taxonomy tags are no longer used, while the last 24 amino acid tags and the PAD token are retained.
By employing the strategy of tag-vocabulary redesign and following the experimental setup described below, we first fine-tuned seven models to specialize in generating enzymes belonging to the general EC classes.
After confirming the effectiveness of our conditional learning strategy, we aimed to understand under which conditions fine-tuning a protein LLM is most effective. To this end, we updated the tag vocabulary and fine-tuned ProGen to generate enzymes from seven additional, more specific EC classes (see Section 2.1).
2.4. Model fine-tuning and evaluation - experimental setup
For each EC category, we randomly split the available sequences into 90% training and 10% test sets. The test data was further processed to generate two distinct test sets: 1) A full test set, which includes all available test data; 2) A filtered test set, consisting of sequences with at most 70% BLAST sequence identity to any sequence in the training set. This 70% threshold was chosen to assess the model's ability to generalize to unseen sequences while ensuring a sufficient number of samples in the test set (at least 100 per category). As shown in Supplementary Table S2, this choice balances the need for sequence diversity with the practical requirement of having enough data for evaluation1
For what regards the fine tuning process, we first considered proteins invariant to the direction of sequence generation. Thus, during training, we reversed the amino acid sequence with a probability of 0.2.
To further encourage the models to learn and generate sequences even when an initial keyword is absent, we randomly omitted the initial tag during training with a probability of 0.2 (this is the same strategy and parameter value used by the ProGen model).
All models were fine-tuned on all weights for four epochs.2 Note that we also run preliminary experiments were we froze the first blocks of weights and we fine-tuned only the weights in the upper blocks. This however lead to poor results (data not-shown) and suggested that all the weights are needed to improve results.
The learning rate was set to 0.0001, and the mini-batch size to 2, representing the number of training examples in one iteration. We opted for a small batch size to mitigate overfitting, given the nature of our dataset and model. Indeed, smaller batch sizes can improve generalization by introducing greater stochasticity in the weight updates during training. Additionally, a warm-up period of 1000 iterations was implemented, during which the learning rate was gradually increased. To prevent the “exploding gradient” issue common in deep neural networks, gradient norm clipping with a norm of 0.25 was applied. The Adam (Adaptive Moment Estimation) algorithm [37] was utilized to compute adaptive learning rates for each parameter, ensuring efficient optimization.
Evaluation of fine-tuned models was performed by using two prediction strategies:
a) Teacher forcing (TF) testing: next amino-acid prediction is performed by “forcing” the LLM to predict the next amino acid using the correct and not the currently predicted previous sequence, where represents the enzyme sequence that precedes the amino acid , and the predicted sequence before the amino acid .
b) Prefixed testing (PF): prediction using a prefix string for the amino acids without teacher forcing. This time the model predicts the next amino acid using , but starts enjoying the correct prefix of amino acids. In other words, while PF uses the previous predicted sequence to predict the next amino acid, TF uses the correct “true” sequence instead.
During the model assessment phase, sequence prediction was performed by using top-k sampling with , i.e. the most probable next amino acid is selected at each prediction step. The metrics computed during the prediction tasks are mean hard accuracy per-token, mean soft accuracy per-token based on BLOSUM62 [38] amino acid substitution matrix, and perplexity, that is the exponent of the cross-entropy loss calculated over each token in the dataset. Lower perplexity indicates a higher-quality model.
For training and testing the models, we used two multi-processor servers equipped with 128 GB of RAM and NVIDIA A100 GPU accelerators (Table S3 reports the average empirical time for training required by the EC classes we tested).
2.5. Generation of new enzyme sequences
Once Fine-tuned, the Finenzyme models can be used to generate new sequences by using one of the following two strategies (detailed in the following): (a) sequence generation by scratch using keywords only; (b) sequence generation from a prefixed sequence of amino acids and keywords.
(a) Sequence generation with keywords only. We started to generate new sequences using a prompt with keywords only. We applied top-p filtering techniques [39] for choosing the generated amino acid, with and we fixed a repetition penalty parameter equal to 1.23 to address the well-known problem of repetitions of the last predicted tokens. Top-p filtering (i.e., nucleus sampling) dynamically restricts the candidate tokens (the nucleus from which to sample) to those whose cumulative probability is at most p. The next token is then randomly chosen from this subset. This approach balances the exploration of diverse amino acid combinations with the need for plausible and coherent sequence generation. When the token distribution is highly peaked, the nucleus contains only a few tokens with significantly higher probabilities, leading to more deterministic outcomes. In contrast, flatter distributions yield a larger nucleus, allowing for greater diversity in the generated sequences.
After generating the full sequence, we employed the learned stop token (see section 2.3) to identify the most probable endpoint, allowing the sequence to terminate before reaching the maximum length of 512 amino acids. To determine the optimal end-position, we adopted the “global optimal stop-choosing technique”, i.e., for each position in the sequence, we recorded the predicted probability of the stop token and then selected the position with the highest probability as the final sequence length.
(b) Sequence generation from a keyword plus prefixed sequence of amino acids. Besides the keyword, the second approach for novel sequence generation employed a prefixed amino acid sequence as a prompt. The input sequence length was gradually increased from 1 to 250 amino acids, therefore generating 250 sequences. This process was carried out for each natural enzyme used as a prefixed sequence as well as for their reversed (flipped) versions, thereby generating sequences in total per natural enzyme. Top-p filtering was applied with , and a repetition penalty of 1.2 was enforced.
2.6. Evaluation of the quality of Finenzyme sequences
For each specific EC number, we generated 1000 sequences using top-p filtering with , resulting in a total of 14000 sequences generated from scratch using only the functional tag as a prompt. After filtering for duplicate sequences (Supplementary Table S4 presents the distribution of duplicated sequences), we computationally analyzed the resulting 6885 unique identified sequences. To this aim we used computational tools to compare the primary and predicted tertiary structure between natural and generated sequences (results reported in section “Predicted structure and function of Finenzyme sequences are similar to natural enzymes”).
For the primary sequence, we evaluated the BLASTp max identity score of Finenzyme generated sequences versus natural enzymes available both in the EC training and test set and in PDB; BLASTp was applied by using default parameter values and by considering the full enzyme length. We also compared the length distribution of the generated sequences with respect to natural ones.
To evaluate the 3D quality of the generated sequences, their structures were predicted with ESMFold [10] and then compared with those of natural proteins available in the PDB database. The structural similarity was evaluated using the Template Similarity score (TM) [40].
To identify the natural protein most similar to each generated sequence we used the Foldseek top-hit for each generated enzyme vs. the PDB database, by using a “structural bit” score, computed as the product of the bit score from the Smith-Waterman algorithm and the geometric mean of the average Local Distance Difference Test (LDDT) and TM scores [41].
2.7. EC classification of the generated enzymes
As a further assessment of our fine-tuning process, we investigated whether the sequences generated by a fine-tuned Finenzyme model exhibited functional properties consistent with the EC class targeted during fine-tuning. To this aim, we employed CLEAN [42], a state-of-the-art tool for enzyme function prediction in terms of EC classification. Specifically, we used CLEAN to predict the EC class of sequences generated by Finenzyme, evaluating whether the assigned classifications aligned with the intended natural enzyme functions. CLEAN leverages embeddings from the ESM1b model [43] to create detailed vector representations of amino acid sequences, which effectively capture the functional similarities between enzymes. These embeddings are processed using a neural network trained with a contrastive loss function, designed to cluster embeddings of sequences with the same EC number while pushing sequences belonging to different EC classes apart.
During the prediction phase, the query enzyme embedding is compared to the centroids of EC number clusters previously established during the training phase. EC numbers are assigned based on the significant proximity of the query's embedding to these centroids. Two assignment methods are used: the maximum separation method, which selects EC numbers whose centroids are most distinctly separate from others, and the p-value-based method, which identifies EC numbers with statistical significance compared to a background distribution. We experimented with both strategies and obtained similar predictions; for this reason in section “Predicted functions of Finenzyme sequences and natural enzymes are similar” and Fig. 3c we show only the results using the maximum separation method.
Fig. 3.
Analysis of the predicted tertiary structure and function of Finenzyme generated sequences. (a) Correlation between structural similarity (TM-score) and sequence identity (BLAST Max ID) between generated and natural enzymes retrieved from PDB across all low-level EC classes. (b) Correlation between ESMFold prediction confidence (pLDDT) and structural similarity to known proteins in the PDB (TM-score). Blue scatterplots (a, b) refer to top-p = 0.5 nucleus filtering. (c) Results of CLEAN EC category predictions. Bar plots show the F1 scores for EC number predictions for the generated sequences, with blue bars representing results of sequences generated with p = 0.50 and orange bars with p = 0.75. “Global” represents the weighted average F1 score when considering all EC numbers combined. (d) Distribution of TM-scores, and (e) pLDDT values for each low-level EC category. TM-scores have been computed through the most similar pair of natural and generated sequences detected through FoldSeek.
2.8. Prediction of the kinetic parameters of the generated enzymes
To assess the predicted catalytic quality of Finenzyme sequences and substrates, in section “A Finenzyme-driven selection procedure to generate sequences close to a specific natural enzyme” we show the results obtained when predicting the Michaelis-Menten kinetic parameters , , and on sequences generated by a Finenzyme model targeting three sequences of significant scientific interest. Kinetic parameters were predicted by using the state of the art UniKP [44] tool, which leverages machine learning models and protein representation techniques. It employs the ProtT5 [21] model to encode enzyme sequences and a SMILES Transformer [45] to represent substrate structures, creating detailed vector representations of both. These representations are then processed by an Extra Tree ensemble model [46], which predicts the kinetic parameters.
2.9. Embedded representations on natural and Finenzyme generated sequence
Embeddings extracted from the final layers of PLMs have been shown to capture relevant features related to protein structure and function [21], [43]. Accordingly, we obtained embedded representations from the last hidden layer of both Finenzyme and ESM2 [10] for natural and Finenzyme generated sequences and compared their distributions.
To compute the embeddings, we applied an average pooling operation over the last hidden layer of each model. Given that the final hidden layer of Finenzyme has dimensions and that of ESM2 is , the pooling operation resulted in two fixed-size vectors of 1280 and 5120 dimensions, respectively. Padding tokens were masked during this process to ensure that only relevant sequence information contributed to the embeddings.
Then, the embeddings were used to learn a 2D latent space utilizing UMAP (Uniform Manifold Approximation and Projection for Dimension Reduction) [47], i.e. the embedded representations of the enzymes were projected into a bidimensional latent space.4
The embeddings of the full dataset were also used to train a random forest classifier for predicting EC classes.5 K-fold cross-validation with was applied to evaluate the generalization performance.
Using the same experimental setting we trained a logistic regression classifier6 with l-BFGS optimizer [48]. Considering the size of the embeddings compared to the low cardinality of the data, an “L2 penalty” was applied.
2.10. Clustering of natural and Finenzyme generated sequences
To evaluate the relationship between the structure of the natural and Finenzyme generated sequences, we clustered them using Single-Linkage Hierarchical Clustering.7 Prior to clustering, we filtered out sequences from natural and generated ones that did not have any subfamily keyword assigned (that ism those tagged only with 3.8.1). To compute the similarities among all sequences, we utilized two algorithms. Firstly, BLASTp [49], which accounts only for sequence similarity; secondly, Foldseek [41], which considers both primary sequence and tertiary structure. The algorithms were both run in an “all-vs-all” manner. Finally, the E-value for both algorithms was used to build the distance matrix similarly to the large scale hierarchical clustering algorithm [50]. All sequence pairs whose E-value was worse than 0.05 or with no hits were assumed to be unrelated and their distance was set to 1. The linkage matrix computed by hierarchical clustering is reordered to minimize the distance between successive leaves [51]. This reordering is also applied to the distance matrix and results in a reordered adjacency matrix , whose heatmap provides a more intuitive and visually coherent representation of the clustering structure. The adjacency matrix is derived from the distance matrix as follows:
where the self-similarities (on the diagonal) were set to twice the maximum of non-diagonal entries to have large enough values, similar to the approach proposed by [50].
3. Results
In this section we first show that Finenzyme significantly improves ProGen sequence prediction of the most specific EC categories. Then we highlight that the primary and the predicted tertiary structures of the enzymes generated by Finenzyme are close to natural ones, using ESMFold [10], FoldSeek [41], and their predicted functions are preserved using CLEAN [42].
We analyze the embedded representations of the natural and Finenzyme sequences, revealing that they are very close, independently if obtained from the last layers of Finenzyme or ESMFold, and the same embeddings can be used to accurately classify EC categories. We also show that clusters of natural and generated enzymes, constructed from adjacency matrices computed through BLAST and Foldseek largely superpose. The diversity of the generated Finenzyme sequences has been assessed and compared with natural ones using normalized Shannon entropy, showing also that parameter p of top-p sampling can be used to tune the grade of diversity of the generated sequences. A thorough analysis of the active sites in Finenzyme generated and natural EC 3.8.1 sequences reveals that they are highly conserved in generated enzymes, confirming the functional analysis results obtained with CLEAN.
Finally, we present some application examples to show how to use Finenzyme to support the design of functionally characterized enzymes. Fig. 1 provides a schematic summary of the main computational tasks performed in this work.
3.1. Fine tuning is effective only for low-level EC classes
High level (more general) classes encompass a broader range of enzymes available for fine-tuning. However, these enzymes are more heterogeneous than those in low level (more specific) classes, raising the question whether the effectiveness of fine-tuning may depend on the hierarchy of EC categories.
To answer this question, we fine-tuned seven models on general EC classes and seven on specific EC classes, and we tested their performance at predicting the next amino acid in a sequence. The general EC classes represent broad categories of enzyme functions, such as oxidoreductases (EC1), transferases (EC2), hydrolases (EC 3), lyases (EC 4), isomerases (EC 5), ligases (EC 6), and translocases (EC7). Specific EC classes denote particular enzyme catalyzed reactions and in our experiments we used alcohol dehydrogenase (EC 1.1.1.1), DNA methyltransferase (EC 2.1.1.37), cellulase (EC 3.2.1.4), ribulose-bisphosphate carboxylase (EC 4.1.1.39), chorismate mutase (EC 5.4.99.5), biotin ligase (EC 6.3.4.15), and proton-translocating transhydrogenase (EC 7.2.1.1) (see Materials and Methods for details).
We recall that, for each EC category, we randomly split the data into 90% training and 10% test sets. To highlight the model's generalization capabilities, we also considered a “filtered test set” by selecting test examples having at most 70% sequence identity with any sequence in the training set (see Section “Fine tuning and sequence generation testing - experimental setup”). We then used both TF and PF to test the fine-tuned models and evaluated the results by computing the mean accuracy per-token, the mean soft accuracy per-token based on BLOSUM62 [38] amino acid substitution matrix, and the perplexity (see Section 2.3 for details).
Fig. 2a compares the TF results on the test set obtained between Finenzyme and ProGen models, summarizing the distribution of accuracy, soft-accuracy, and perplexity for general and specific EC classes. Our results highlight that fine-tuned models greatly outperform ProGen when the EC classes are specific (i.e. low in the EC hierarchy), while results are comparable when the EC classes are general. Importantly, the difference in accuracy, soft accuracy and perplexity remains significant also when using the filtered test set (Wilcoxon rank-sum test, ), highlighting that Finenzyme generalizes better than ProGen also for test sets that largely differ from those used in the training set.
Fig. 2.
Comparison of ProGen and Finenzyme on general and specific EC classes. (a) accuracy (top), soft accuracy (center), perplexity (bottom) compared between general and specific EC classes. “Filtered” denotes whether the test dataset was filtered using BLAST against the training set to retain only sequences with less than 70% identity. (b) soft accuracy of the specific EC classes. “TF” denotes teacher forcing, “PF” testing without teacher forcing with a prefixed chain of 20 amino acids. (c) soft accuracy comparison of Finenzyme and ProGen on general (left) and specific (right) EC classes on the full test set. The x-axis reports the position of the amino acid and the y-axis the corresponding average soft-accuracy in ProGen (red line) and Finenzyme (green). Shadows represent the standard deviation.
Fig. 2b shows a comparison between TF and PF results using soft-accuracy for the more specific EC classes, confirming that fine-tuned models achieve significantly better results in both scenarios – similar results are also obtained when measuring accuracy and perplexity (Fig. S5, Supplementary Information).
Fig. 2c compares the soft accuracy calculated on general and specific EC categories at each position in the sequence showing that Finenzyme consistently outperforms ProGen over the entire sequence length, independently on the amino acid position. Figs. S6–S13 in the Supplementary Information show the same comparison using accuracy, soft accuracy, and perplexity in all test configurations.
3.2. Predicted structure and function of Finenzyme sequences are similar to natural enzymes
We used Finenzyme to generate sequences for the more specific classes and we then carried out in silico experiments to computationally assess their biological plausibility. For each low-level EC category, we generated 2000 sequences using only the EC keywords. To accomplish this, during fine-tuning we applied top-p filtering techniques [39] for choosing the next generated amino acid, with , or and repetition penalty parameter of 1.2. Then we adopted the “global optimal stop-choosing technique” to choose the length of the generated enzyme (Section 2.5). In the following sections, we present our analysis comparing the sequence, predicted structure, and predicted function of Finenzyme generated sequences with those of natural enzymes.
3.2.1. In depth predicted tertiary and primary structure analysis of the generated Finenzyme sequences
To assess the tertiary structure similarity between the generated and the natural sequences, we used ESMFold [10] to predict structures for Finenzyme sequences and applied Foldseek [41] to select the most structurally similar pairs of generated and natural enzymes retrieved from PDB according to the Foldseek structural bit scores (see “Evaluation of the quality of Finenzyme sequences” in Materials and Methods for details).
Fig. 3a shows the relationship between structural similarity (measured by the TM-score [40]) and sequence identity between top-hit Foldseek pairs of generated and natural proteins available in PDB [41]. The TM-score distribution (top x-axis – Fig. 3a) is centered around 0.9, indicating that the predicted structures of the generated and natural proteins are very similar. The distribution of sequence similarity (left y-axis – Fig. 3a) spans a wide range of values, with a mode between 0.3 and 0.4, and has a low correlation with TM-score (Pearson correlation ). Finenzyme thus generates primary sequences that may differ from those of natural enzymes, but can preserve the tertiary structure of the enzymes, thereby probably retaining their original enzymatic functions.
Finally, we checked whether the enzyme structural similarities were correlated with the confidence of the ESMFold predictions. Fig. 3b shows that there is indeed a high correlation between the pLDDT values (quantifying the confidence of ESMFold predictions) and TM scores (Pearson correlation ). We observe that a high correlation between pLDDT and TM scores is maintained for sequences with maximum BLAST sequence identity below 80%, and even below 60% or 40%. This also applies when assessing the correlation of the pLDDT and the sequence-similarity. (Fig. S19 and S20). A noticeable decline in pLDDT and TM scores correlation is observed only when sequence similarity falls below 20%. This suggests that our model can reliably preserve the predicted tertiary structure with high confidence, even when the primary structure of the generated sequences significantly differs from natural ones. Details of the TM-score and pLDDT values distribution for each low-level EC category are displayed in Fig. 3d, e. Supplementary Fig. S17 and S18 present the relationships between structural and sequential similarity and ESMFold prediction confidence for each EC class.
To analyze the diversity of the Finenzyme generated sequences, we compared the normalized Shannon entropy of the di-grams computed on both natural and Finenzyme generated sequences across 7 specific EC categories. As shown in Fig. 4, the normalized Shannon entropy is comparable between natural and generated sequences, with values approaching 1. As expected, higher values of p in top-p sampling yield increased diversity. By tuning p, we can adjust the level of diversity in the generated sequences.
Fig. 4.
Analysis of the diversity in natural (green) and generated (blue, p = 0.50) and (orange, p = 0.75) sequences in 7 specific EC categories. Y axis represents the normalized Shannon entropy of the sequence di-grams.
3.2.2. The predicted structure of Finenzyme generated sequences is more preserved than the primary
We compared predicted structures from Finenzyme sequences and their closest match found in the PDB database by Foldseek. Fig. 5 depicts some specific cases of Finenzyme-generated sequences having low primary structure similarity but high 3D structural similarity with natural enzymes. More precisely Fig. 5 shows the closest match between the Finenzyme sequences (in green) and natural ones (in yellow), considering three different enzymes. Note that, in all cases, binding sites are never mutated (highlighted in red). Furthermore, when the ligand is available, it fits well within the binding pockets.
Fig. 5.
Comparison of the tertiary structure of Finenzyme-generated sequences (green) and natural enzymes (yellow). Enzyme generated from (a) EC family 6.3.4.15 (biotin ligase). The target found is “2e41”, a biotin protein ligase (UniProt accession O57883) from Pyrococcus horikoshii: sequence similarity =34.9%, TM-score =0.94, pLDDT (ESMFold) =0.95; the biotin binding sites of the enzyme are predicted with high confidence (highlighted in red) and provide adequate space for the ligand (highlighted in violet). (b) EC 3.2.1.4 (cellulase). The PDB target found is “8ihw”, an endoglucanase from Eisenia fetida: similarity =39.6%, TM-score =0.95, pLDDT (ESMFold) =0.92. (c) EC 5.4.99.5 (chorismate mutase). The PDB target found is “2pv7”, a prephenate dehydrogenase from Haemophilus influenzae: similarity =49.6%, TM-score =0.9, pLDDT (ESMFold) =0.9; the binding sites of the enzyme are predicted with high confidence highlighted in red) and provide adequate space for the ligand (highlighted in violet). The binding sites (highlighted in red) are highly conserved in Finenzyme-generated enzymes.
3.2.3. Predicted functions of Finenzyme sequences and natural enzymes are similar
We predicted the functional properties of Finenzyme sequences using CLEAN, a state-of-the-art enzyme function prediction tool [42], that assigns EC numbers to protein sequences (see details in section “EC classification of the generated enzymes” in Materials and Methods). The F1 scores for sequences generated at top-p, with are displayed in Fig. 3c (the recall scores are reported in Supplementary Fig. S16). The lower scores observed for (especially for EC 1.1.1.1) indicate that classifying sequences with reduced sequence identity to natural enzymes are challenging for CLEAN. However, for both generation settings, CLEAN achieves robust predictions with minimal variations across the classes, showing that the predicted functional characteristics are well preserved even when primary sequences diverge from those of natural enzymes.
3.3. An in depth computational analysis of hydrolases acting on C-halide compounds (EC 3.8.1)
To demonstrate the possible applications of Finenzyme, we turned our attention on an industrially important family of enzymes (EC 3.8.1). Organohalogen compounds are extensively employed across various sectors including agriculture, healthcare, manufacturing, and more, due to their versatility in applications ranging from herbicides and insecticides to refrigerants and synthetic precursors. Despite their broad utility, they represent a significant environmental pollutant [52], permeating marine and terrestrial ecosystems alike [53]. Their chemical stability and resultant persistence in the biosphere poses ongoing environmental and health risks due to the difficulty in degrading these compounds and their toxic effects even at low concentrations [54], [55]. Enzymes such as hydrolases acting on halide bonds in C-halide compounds have gained considerable interest for their potential in bioremediation and industrial applications, where they transform organohalogens into less harmful compounds [56], [57], [58]. Thus, the study and enhancement of dehalogenase activity are pivotal in advancing our capabilities to manage and mitigate the impact of halogenated pollutants globally.
We trained Finenzyme with EC 3.8.1 sequences retrieved from Uniprot [9]. Fig. S25 shows the number of enzyme sequences belonging to each dehalogenase subclass. The model was trained by prepending each sequence with the more general 3.8.1 token. Given the low number of sequences for some subclasses, we only prepended more specific tokens (3.8.1 + 3.8.1.2 / 3.8.1 + 3.8.1.5 / 3.8.1 + 3.8.1.3) when more than 1000 instances were available for that specific subclass. We observe that fine-tuning on a set of functionally related EC categories allows for transfer learning and a substantial enlargement of training examples. For instance, with EC 3.8.1.3 we can increase the number of available examples for training our model from about 1300 to about 20000, thanks to the addition of the related training examples from EC 3.8.1.2 and 3.8.1.5.
The prediction results for each subclass are depicted in Fig. 6a. Compared to ProGen, Finenzyme achieved substantial improvement in all the evaluation measures and across all the EC subclasses (Wilcoxon signed rank tests, p ). Although the improvement is more evident on the most represented subclasses, our results show that this experimental setting boosts performance across all EC subcategories.
Fig. 6.
Prediction results and analysis of the embedded representations of natural and Finenzyme generated proteins for hydrolases acting on C-halide compounds. (a) Hard accuracy, soft accuracy and perplexity performance of Finenzyme model against Progen for the different EC 3.8.1 subclasses on the test set. (b) UMAP projections of the embedded Finenzyme representations of (left) natural enzymes, and (right) Finenzyme generated sequences. Considering the disparity in cardinality of certain subclasses, we normalized the distribution of the generated sequences through kernel density estimation maps. (c) UMAP projections obtained from the embeddings of ESM2 comparing natural and generated sequences. (d) Confusion matrices of EC classification tasks performed by Random Forest (blue) and Logistic Regression (red), using the embeddings retrieved from Finenzyme and ESM2.
To assess the effectiveness of tokens in influencing model generation toward the correct subclass, we generated 1000 sequences for each subclass using nucleus filtering with top-. The lengths of the generated sequences in each experiment, compared to those of the training set, are shown in Fig. S26. Duplicates were removed, and a summary of the generated and natural sequences is provided in Supplementary Table S7. Fig. S27 shows the distribution of the max-ID BLAST sequence identity scores between the generated sequences and natural enzymes across 3.8.1, 3.8.1.2, 3.8.1.3, and 3.8.1.5 EC categories.
3.3.1. Embedded representations of natural and Finenzyme generated sequences are similar
After using Finenzyme to generate sequences with subfamily keywords 3.8.1.2, 3.8.1.5, 3.8.1.3, we used the last hidden layer of Finenzyme and ESM2 [10] to embed both natural and Finenzyme generated sequences (details in Section 2.9). The embedded representations were then projected into 2D representations through UMAP [47]. Fig. 6b, c presents the UMAP projections obtained from the embeddings obtained from Finenzyme and ESM2, respectively (details in section “Embedded representations on natural and Finenzyme generated sequences” - Material and Methods). Independently of the method used to generate the embeddings, the UMAP projections of natural and Finenzyme embedded representations are very close (Fig. 6b, c, right), indicating that our fine-tuned model captures the main features of the natural enzymes.
The embeddings were then used to train random forest and logistic regression classifiers. Confusion matrices obtained through cross-validation show that both classifiers correctly predict EC subclasses of 3.8.1 using independently Finenzyme or ESM2 embeddings (Fig. 6d).
3.3.2. Clustering of the 3D representation of natural and generated sequences is close
To evaluate the relationship between the structure of the natural and Finenzyme generated sequences, we clustered them using Single-Linkage Hierarchical Clustering. The metric used to construct the distance matrix was the E-value retrieved from two experiments that include both the natural and Finenzyme sequences: 1) Protein BLAST all-vs-all, which accounts for primary structure similarity; 2) Foldseek all-vs-all, which accounts for both primary and tertiary structure similarity (see details in Section 2.10).
In order to provide a more comprehensive database for structural analysis comparison, we removed the constraint for PDB experimentally validated structures, and we predicted the structures for the entire natural dataset using ESMFold. The heatmaps of the adjacency matrices reordered according to the clustering algorithm proposed in [51] are shown in Fig. 7. At the side of each heatmap, kernel-density estimates allow to visually compare the relative distributions of Finenzyme sequences and natural proteins. We can observe that each subclass contains sub-clusters, and, importantly, clusters of natural and Finenzyme sequences largely overlap.
Fig. 7.
Heatmaps of the clustered natural enzymes and Finenzyme generated (“artificial”) sequences. Different colors highlight different EC categories. Heatmaps are obtained from the adjacency matrix computed through BLAST (left) and Foldseek (right) E-values. On the sides: kernel density estimation of the distribution of natural and artificial, i.e. Finenzyme generated sequences.
3.3.3. Finenzyme shows a significantly larger correlation than ProGen between the confidence of its predictions and structural or sequence similarity with natural enzymes
We compared the confidence in the prediction of Finenzyme with ProGen by evaluating the log-likelihood of their generated sequences, which shows how confident a model is in the prediction of an enzyme sequence. Fig. 8a, b shows that Finenzyme log-likelihood strongly correlates with pLDDT, TM-score, and sequence identity, and the correlation is significantly larger for Finenzyme with respect to ProGen, underscoring that Finenzyme is able to predict with higher confidence sequences that are more biologically plausible.
Fig. 8.
Correlation between the log-likelihood of the predicted sequences with TM-scores, pLDDT, sequence identity and E-value. (a) Pearson correlation between the log-likelihood scores of Finenzyme (blue) and ProGen (orange) against pLDDT, TM-score, sequence identity, and E-value. The metrics are computed either through ESMFold or Foldseek (b) Log-likelihood of the predictions of Finenzyme (left) and ProGen (right) as a function of TM-score, pLDDT, E-value and sequence identity. (c) Comparison of the predicted ESMFold 3D structure of the natural enzyme (fluoroacetate dehalogenase, in gold) and the corresponding Finenzyme sequence (green) generated from EC subcategory 3.8.1.3 (fluoroacetate dehalogenases). Important residues belonging to the catalytic triad (blue) or favoring the reaction (magenta) are highlighted. (d) The same Finenzyme predicted sequence superimposed onto the corresponding natural one retrieved from the AlphaFoldDB. Residues belonging to the catalytic triad are highlighted in blue and others facing the active site in magenta.
Moreover, the fine-tuned model shows a decrease in sequence identity correlation compared to the TM-score, which evaluates the tertiary structure similarity. These results show that Finenzyme can better learn the 3D structural plausibility of a sequence rather than its mere primary sequence.
Driven by the ability of our model to preserve structure, we compared the predicted structure of one of the sequences generated by our model using the keyword 3.8.1.3 (fluoroacetate dehalogenases), i.e. the least represented subfamily, with its most similar PDB counterpart. The TM-score between the two structures is 0.91, despite a sequence identity of 0.39 (Fig. 8c). Furthermore, driven by these results, we conducted a literature investigation of the PDB structure. The PDB reference structure (PDB code = 3R3U) corresponds to the fluoroacetate dehalogenase RPA1163, which has been thoroughly studied due to its potential industrial applications [59], [60], [61].
By examining our predicted structure and comparing it with the real one, we can observe that the catalytic triad has been maintained (Fig. 8c, blue), and several important amino acids necessary for the enzymatic reaction are also well conserved (Fig. 8c, magenta). However, some regions of the proteins differ. We hypothesized that it might be due to the generated enzyme resembling the structure of a natural protein for which no PDB structure has been experimentally solved yet. Interestingly, the training set sequence with the highest sequence identity has a predicted structure present in the AlphaFoldDB [62]. We, therefore, compared our enzyme with the latter. In this case, the sequence identity reaches 0.76, but the TM-score skyrockets to 0.97 (Fig. 8d). Furthermore, all amino acid residues facing the active site are extremely conserved.
3.4. Analysis of active site conservation
To analyze the capability of Finenzyme to preserve active sites, we generated 1,000 sequences for three subfamilies of hydrolases acting on C-halide compounds (EC 3.8.1). After removing duplicates, we retrieved the following numbers of sequences per subclass: 487 for EC 3.8.1.2, 491 for 3.8.1.3 and 502 for 3.8.1.5. Then, we aligned the generated sequences with their respective most similar natural sequences in the training set to characterize their differences. More precisely, we used EsmFold [10] to predict the tertiary structures for both generated and training set sequences, and we utilized Foldseek [41] to find the most similar structure pairs and aligned them.
From the aligned sequences, we found that 191 contained active site annotations in UniProt [63] (with a few annotations for binding sites as well), and we examined whether our model preserves such functional sites.
To this end we considered the following null-hypothesis :
and the alternative hypothesis :
The average mutation rate (considering gaps) of the aligned 191 sequences is 0.18. The number of matches for these sites is 559 out of 560. The only mutation found occurred between biochemically similar amino acids (from aspartic acid to glutamic acid). Assuming each site is conserved independently, the number of conserved sites out of 560 can be modeled as a binomial distribution where the probability of conservation is and the total number of sites is . Therefore, the expected number of conserved sites is:
The variance is:
Therefore, the standard deviation is .
Consider that n is pretty high and p is not too close to 0 or 1, we can approximate to a Gaussian distribution. The approximation allows us to calculate the -score:
A -score of 10.98 corresponds to a
Hence, we reject the null hypothesis, concluding that the observed conservation of important annotated sites is not due to random chance and shows a significant deviation. This suggests that the model is capable of generating diverse sequences that still maintain amino acids in positions annotated as being of critical importance for enzyme function.
Interestingly, for three aligned generated sequences, their natural counterparts have PDB structures available. Each PDB corresponds to a different hydrolase subfamily: 3.8.1.5, 3.8.1.2 and 3.8.1.3. Fig. 10 shows the superposition of the predicted structures of Finenzyme generated sequences and the most similar natural found in the PDB database. Although the degree of sequence identity varies across the three sequences, the three-dimensional similarity with the PDB structures is well-maintained. In Fig. 10, the mutated residues in the target are shown in red, and the important amino acids annotated in UniProt are shown in green. For all structures, the essential residues are never mutated and align well with their counterparts.
Fig. 10.
Superposition of the predicted structures of artificial sequences with their most similar structures from the PDB database. The mutations in the artificial sequences are shown in red, and the important binding and active sites, which are never mutated, are highlighted in green. (a) Haloalkane Dehalogenase from Rhodococcus (EC: 3.8.1.5); PDB: 1BNS, Seq-ID =0.986, TM-score =1 (b) Structure of L-2-Haloacid Dehalogenase from Xanthobacter autotrophicus (EC: 3.8.1.2); PDB: 1AQ6, Seq-ID =0.849, TM-score =0.934 (c) Structure of Fluoroacetate Dehalogenase from Burkholderia (EC: 3.8.1.3); PDB: 1Y37, Seq-ID =0.570, TM-score =0.988. Seq-ID stands for sequence identity.
These results further confirm that the model is able to generate artificial proteins with varying degrees of sequence identity to their natural counterparts while retaining both the overall structure and essential residues necessary for catalyzing their enzymatic reactions.
3.5. A Finenzyme-driven selection procedure to generate sequences close to a specific natural enzyme
As an example of Finenzyme application, we show a technique to generate sequences close to a specific target natural enzyme. We demonstrate that the predicted tertiary structure of the generated sequences is very close to that of the target enzyme, and also that their predicted kinematic parameters are very similar.
We focused on three sequences of significant scientific interest: P11766 (alcohol dehydrogenase - EC 1.1.1.1) is characterized by its high activity towards formaldehyde [64]; P05102 (EC 2.1.1.37) prevents the incorporation of damaged bases into DNA, maintaining genomic stability and integrity [65]; Q9SL92 (EC 6.3.4.15) is a potential target for genetic manipulation to improve crop yield and stress tolerance through enhanced metabolic efficiency [66]. The generation process, detailed in section “Generation of new Finenzyme sequences” (Material and Methods), used both keywords and amino acid prefix of the targets, yielding a total of 3000 sequences. First, duplicate sequences were filtered out (Supplementary Table S5 provides the count of unique sequences per target). Next, predicted structures were obtained using ESMfold and were then evaluated versus the target enzymes. Structures with a pLDDT greater than 0.7, being the latter a good confidence score, as defined by ESM authors [43], were retained, and those with a TM-score to the target greater than 0.7 were selected. Note that these are parameters of the selection procedure, and as such they can be tuned according to specific generation objectives (see Supplementary Section S5 for details).
We also matched the predicted tertiary structure of the generated sequences against the Protein Data Bank (PDB) using Foldseek. Supplementary Table S5 shows that for each target enzyme we obtained a relatively large set of novel sequences that satisfy the above criteria. Subsequently, we clustered the filtered sequences using their Foldseek representation and the MMseqs suite for fast and deep clustering adapted to the “3Di” alphabet of Foldseek [67]. From each cluster we extracted a representative sequence and performed multiple sequence alignments based on the Foldseek enzyme representations. The alignment results relative to P05102 are visualized using a circular tree (Fig. 9a), that includes the natural enzyme targets and the corresponding sequences selected through our proposed technique. Analogous circular tree figures for P11766 and Q9SL92 are available in the Supplementary Information (Fig. S22, and S23).
Fig. 9.
Finenzyme driven selection of sequences and evaluation of the predicted kinetic parameters. (a) Circular tree representing the Finenzyme sequences generated from the enzyme P05102 (PDB structure identifier 1fjx) (b–g) Kinetic parameter km prediction using UniKP. Each scatter plot displays the predicted km between the natural sequence and the substrate as a red line; the blue points are predictions of Finenzyme sequences generated through the selection procedure. The enzyme-substrate pairs are (b) P11766 and 20-HETE, (c) P05102 and DNA, (d) P05102 and S-adenosyl-L-methionine, (e) Q9SL92 and biotin, (f) Q9SL92 and ATP, (g) Q9SL92 and methylcrotonyl-CoA carboxylase. Values are displayed in μM.
Finally, Michaelis-Menten kinetic parameters of the natural enzyme target and of the selected sequenced were predicted using UniKP ([44] (see “Prediction of the kinetic parameters of the generated sequences” in Material and Methods). The results of these predictions are depicted in Fig. 9b–g, highlighting the predicted functional capabilities of both natural and Finenzyme selected sequences. Supplementary Table S6 describes the substrate/protein interaction data available in UniProt for the chosen targets, and the predicted values for the true targets.
4. Discussion
Several studies demonstrated that fine-tuning can improve PLM performance in various tasks and settings [18], [19], [21]; however, other works highlighted that this technique does not necessarily lead to better results [27], [68].
In the context of enzymes, we showed under which conditions fine-tuning a conditional Transformer can lead to significantly better results compared to a “general” PLM. Fine-tuning specific EC categories and exploiting specific keywords in the conditional learning process boosts performance. On the other hand, for more general categories, e.g., the top high-level EC categories, this strategy does not yield better results (Fig. 2c). This can be explained by the fact that general models like ProGen can generalize on overly represented EC categories, by leveraging their knowledge about the large corpus of data on which they have been trained. Conversely, the general model cannot leverage specific knowledge to accurately predict underrepresented and functionally characterized EC categories, while fine-tuned models can learn and focus on specific and well-characterized EC categories.
Analysis of Finenzyme sequences using ESMFold, Foldseek, and other state-of-the-art software tools, reveals that the predicted tertiary structure is highly conserved compared to natural enzymes (average TM-score ≃0.9), while the sequence identity is less preserved (Fig. 3). This indicates that Finenzyme sequences can diverge from natural ones but tend to preserve their 3D structure. The in silico prediction of EC categories of the generated and natural enzymes shows that Finenzyme generated enzymes maintain the same function of the natural ones (Fig. 3c). Moreover, the analysis of the conservation of active sites confirms that Finenzyme generated sequences preserve the functions of natural enzymes (Fig. 10).
For overly underrepresented EC categories, we propose transfer-learning techniques by jointly fine-tuning the conditional Transformer on functionally related EC categories while maintaining their specificity using their corresponding keywords. Results with EC 3.8.1 low-level categories show that Finenzyme can simultaneously learn multiple functionally related EC categories, using multiple keywords and improving the cardinality of the training examples.
Embedded representations of natural and generated sequences obtained from the last hidden layer of both Finenzyme and ESMFold models span a similar distribution (Fig. 6b, c), and the embedded representations themselves can be used to correctly classify the EC categories (Fig. 6d).
Hierarchical clustering of the 3D structures of both Finenzyme generated sequences and the corresponding natural ones yields largely superposed clusters, confirming that our model is able to generate sequences that mimic natural ones.
We also observe that every EC category includes sub-clusters (Fig. 7). This is expected since EC numbers do not specify enzymes but rather enzyme-catalyzed reactions. Namely, if different enzymes catalyze the same reaction, they receive the same EC number, regardless of their structural similarity [69]. Moreover, through convergent evolution, completely different protein folds can catalyze an identical reaction and therefore would be assigned the same EC number [70]. However, the smoothness of the Foldseek heatmap (Fig. 7) shows that the predicted tertiary structures are more maintained among clusters than primary sequences, confirming the relationships we showed between TM-score and sequence identity scores (Fig. 3).
The intra-cluster structure depicted in Fig. 7 suggests to apply fine-tuning to generate specific protein domains, as proposed by [71]. Another solution could be to cluster the dataset before training and assign additional sub-keywords to each sub-cluster. This can be accomplished by: a) fine-tuning Finenzyme on the discovered sub-clusters; b) by adding a new tag for each sub-cluster; c) by using multiple keywords of other functionally similar enzyme categories to overcome the under-representation of small enzyme subcategories.
Nevertheless, by looking at the kernel density distributions on the sides of the heatmaps, we found that Finenzyme sequences span all the main natural clusters fairly well (Fig. 7).
The log-likelihood of the Finenzyme predictions are significantly more correlated than ProGen with the TM-score, pLDDT, E-value and sequence identity, showing that fine-tuning can significantly boost the reliability and the confidence of the predicted sequences (Fig. 8a, b). Preservation of the overall tertiary structure and of active and binding sites outlines the model's capabilities to generate novel enzymes that retain fundamental functional properties (Fig. 8c, d; Fig. 10a, b, c).
Finally, on the basis of the reliability and biological plausibility of Finenzyme sequences, we propose a “Finenzyme driven selection” process that can be applied to guide the in silico synthesis of enzymes with structure and function similar to a specific target (Fig. 9).
Note that while planning our work, we considered the high computational costs typically associated with training LLMs. To mitigate this, we selected a decoder-only conditional Transformer with a relatively low number of parameters. Our preliminary experiments employed a more efficient fine-tuning strategy by freezing the lower layers and fine-tuning only the upper ones. However, under these settings, performance was poor, suggesting that freezing the lower layers hindered the model's ability to properly adapt to the novel pre-appended tags. These findings align with studies in parameter-efficient fine-tuning, which indicate that while freezing model parameters can provide efficiency benefits, it allows for only limited trainable adaptations [16]. Although updating all weights is less parameter-efficient than other fine-tuning strategies, such as LoRA and its variants [15], our results reinforced the decision to pursue full fine-tuning. Consequently, we opted for standard fine-tuning—updating all layers—while limiting training to four epochs. This approach maintained relatively low computational demands, with fine-tuning for even the largest enzyme classes being completed in a reasonable timeframe on a single NVIDIA A100 GPU (see Supplementary Table S3).
We must, however, acknowledge that despite the good results we achieved, fine-tuning pre-trained LLMs carries the risk of data leakage due to various factors [72]. These include biased experimental settings, which is not a concern in our case, as well as an incomplete understanding of the datasets used during pre-training.
In particular, LLMs (including PLMs) are often trained on numerous open datasets, yet the full and detailed list of these datasets is rarely disclosed. As a result, even when unbiased training and testing procedures are employed, there is no absolute guarantee that a fine-tuned PLM has not previously encountered portions of the test data during its pre-training phase. In other words, it remains difficult to determine whether strong performance arises from genuine generalization or simply from memorization of pre-existing sequences.
A shortcoming of this work is that only in silico validations have been performed. For instance, the limited coverage of tertiary structures in the PDB database might have constrained the comparison of our generated sequences with natural ones. To address this issue and corroborate each validation strategy, we performed several experiments—including assessments of sequence, structure, and active site conservation, clustering with natural proteins, and predictions of function and kinetic parameters (with CLEAN and UniKP, respectively). Finally, we demonstrated that the model itself can score its generated sequences (using log-likelihood) to filter out non-functional ones arising from stochastic sampling strategies. Therefore, despite shortcomings in each metric, the consensus among our assessments strongly supports the strength of our strategy. Nonetheless, we plan to conduct wet-laboratory experiments in future work as a final characterization of our generated enzymes.
Another limitation of our approach is that Transformers, despite being state-of-the-art for sequence generation, are data-hungry models, and thus require large amounts of training data. Fortunately, thanks to recent developments in high-throughput sequencing technologies [73], the number of available protein sequences has been exponentially increasing. In the future, such a massive amount of sequences will also allow to train even larger and more complex models. This provides a distinct advantage of purely sequence-based approaches over those that require experimentally determined protein structures, given that structural databases remain orders of magnitude smaller in coverage. Moreover, we have shown that training on functionally related proteins can help overcome data scarcity for underrepresented categories (e.g., EC: 3.8.1.3), partially mitigating coverage gaps in public databases. However, as more sequence data become available, the corresponding annotations often rely on automated pipelines (e.g., BLAST-based similarity searches) or machine learning tools (such as CLEAN), which both depend on data with manual annotations. Thus, human biases can propagate through such annotations. Therefore, while large, semi-automated datasets are valuable, continued expert curation and the integration of complementary experimental evidence will be essential to reduce erroneous annotations and maintain high-quality training data.
Concluding, our experimental results and analyses demonstrate that Finenzyme fine-tuning of conditional Transformers on a dataset of functionally related sequences enables the generation of novel sequences spanning a broad range of enzyme classes. This is particularly promising given that enzymes play a central role in a wide array of biological processes, can be produced at scale using advanced bioprocessing techniques, and have broad applications in fields such as drug development, food production, and environmental remediation. Remarkably, the generated sequences display significant divergence in primary structure while preserving both tertiary structure and biological function, as witnessed by the conservation of active sites and by the computational prediction of their function, indicating that our approach effectively explores the protein fitness landscape. We anticipate that Finenzyme fine-tuning can be harnessed to create highly diverse synthetic enzyme libraries, thereby enhancing in vitro protein design frameworks—such as enzyme-directed evolution—by providing access to unexplored regions of the fitness landscape that remain inaccessible via natural evolutionary processes.
CRediT authorship contribution statement
Marco Nicolini: Writing – review & editing, Validation, Software, Methodology, Data curation, Conceptualization. Emanuele Saitto: Writing – review & editing, Software, Resources, Methodology, Investigation, Data curation, Conceptualization. Ruben Emilio Jimenez Franco: Writing – original draft, Validation, Software, Investigation, Data curation. Emanuele Cavalleri: Writing – review & editing, Data curation. Aldo Javier Galeano Alfonso: Writing – review & editing, Software, Data curation. Dario Malchiodi: Writing – review & editing, Validation, Formal analysis. Alberto Paccanaro: Writing – review & editing, Writing – original draft, Supervision, Methodology, Formal analysis, Conceptualization. Peter N. Robinson: Writing – review & editing, Validation, Supervision, Methodology, Investigation, Conceptualization. Elena Casiraghi: Writing – review & editing, Writing – original draft, Supervision, Project administration, Methodology, Investigation, Formal analysis, Conceptualization. Giorgio Valentini: Writing – review & editing, Writing – original draft, Supervision, Project administration, Methodology, Investigation, Funding acquisition, Formal analysis, Conceptualization.
Declaration of Competing Interest
The authors declare that they have no known competing financial interests or personal relationships that could have appeared to influence the work reported in this paper.
Acknowledgements
This work was supported by National Center for Gene Therapy and Drugs Based on RNA Technology—MUR (Project no. CN_00000041) funded by NextGeneration EU program, and by FAIR (Future Artificial Intelligence Research) project, funded by the NextGenerationEU program within the PNRR-PE-AI scheme (M4C2, Investment 1.3, Line on Artificial Intelligence) - AIDH – FAIR - PE0000013.
Supplementary material related to this article can be found online at https://doi.org/10.1016/j.csbj.2025.03.037.
The resulting average sequence identity of the reduced test set for several EC classes is below 50% (Fig. S1 and S2 in the Supplementary Information).
The number of fine-tuning epochs has been selected according to early-stopping strategies.
The selection of the p and penalty parameters has been performed based on preliminary results in terms of duplicate sequences on lysozymes fine-tuning.
The implementation of the algorithm was retrieved from the Python package umap. The distance metric used was the cosine similarity, with parameters min_dist =1, and n_neighbors =100. These settings were kept equal for the embeddings retrieved from both models.
We applied the RandomForestClassifier from the sklearn Python package. The number of trees was set to 2000 while all the other hyperparameters were set to their default values.
We applied the LogisticRegression implementation from the Python package sklearn.
Single-linkage hierarchical clustering was performed by using hierarchy.linkage implementation from the scipy library.
Appendix A. Supplementary material
The following is the Supplementary material related to this article.
Additional text and figures relative to enzyme datasets used in the experiments, model architecture, training time and results on the test set. Three specific sections are also included: a) Only-keywords Finenzyme generation, CLEAN results, and per-class analysis of ESMFold predictions; b) An example of the application of Finenzyme to generate sequences close to a specific natural enzyme; c) Finenzyme generation of hydrolases acting on C-halide compounds (EC 3.8.1).
Code and data availability
The Finenzyme code, the scripts to reproduce the experiments and tutorials to fine-tune Finenzyme on any EC category are available from https://github.com/AnacletoLAB/Finenzyme.
The datasets used to train Finenzyme were downloaded from UniProt. Details are available from the notebook tutorial in the above GitHub repository.
References
- 1.Moor M., Banerjee O., Shakeri Z., Krumholz H., Leskovec J., Topol E., et al. Foundation models for generalist medical artificial intelligence. Nature. 2023;616:259–265. doi: 10.1038/s41586-023-05881-4. [DOI] [PubMed] [Google Scholar]
- 2.Wang B., Xie Q., Pei J., Chen Z., Tiwari P., Li Z., et al. Pre-trained language models in biomedical domain: a systematic survey. ACM Comput Surv. oct 2024;56(3) doi: 10.1145/3611651. [DOI] [Google Scholar]
- 3.van Tilborg D., Grisoni F. Traversing chemical space with active deep learning for low-data drug discovery. Nat Comput Sci. 2024;4(7407):786–796. doi: 10.1038/s43588-024-00697-2. [DOI] [PubMed] [Google Scholar]
- 4.Bessadok A., Grisoni F. An explainable foundation model for drug repurposing. Nat Med. 2024 doi: 10.1038/s41591-024-03333-8. [DOI] [PubMed] [Google Scholar]
- 5.Vaswani A., Shazeer N., Parmar N., Uszkoreit J., Jones L., Gomez A.N., et al. Attention is all you need. Proceedings of the 31st international conference on neural information processing systems; NIPS'17; Red Hook, NY, USA: Curran Associates Inc.; 2017. pp. 6000–6010. [Google Scholar]
- 6.Ofer D., Brandes N., Linial M. The language of proteins: nlp, machine learning & protein sequences. Comput Struct Biotechnol J. 2021;19:1750–1758. doi: 10.1016/j.csbj.2021.03.022. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 7.Ferruz N., Höcker B. Controllable protein design with language models. Nat Mach Intell. 2022;4(6):521–532. [Google Scholar]
- 8.Valentini G., Malchiodi D., Gliozzo J., Mesiti M., Soto-Gomez M., Cabri A., et al. The promises of large language models for protein design and modeling. Front Bioinform. 2023;3 doi: 10.3389/fbinf.2023.1304099. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 9.The UniProt Consortium UniProt: the universal protein knowledgebase in 2023. Nucleic Acids Res. 2023;51(D1):D523–D531. doi: 10.1093/nar/gkac1052. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 10.Lin Z., Akin H., Rao R., Hie B., Zhu Z., Lu W., et al. Evolutionary-scale prediction of atomic-level protein structure with a language model. Science. 2023;379(6637):1123–1130. doi: 10.1126/science.ade2574. [DOI] [PubMed] [Google Scholar]
- 11.Ferruz N., Schmidt S., Höcker B. Protgpt2 is a deep unsupervised language model for protein design. Nat Commun. 2022;13(4348) doi: 10.1038/s41467-022-32007-7. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 12.Brandes N., Ofer D., Peleg Y., Rappoport N., Linial M. ProteinBERT: a universal deep-learning model of protein sequence and function. Bioinformatics. 2022;38(8):2102–2110. doi: 10.1093/bioinformatics/btac020. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 13.Ruffolo J.A., Madani A. Designing proteins with language models. Nat Biotechnol. 2024;42(2):200–202. doi: 10.1038/s41587-024-02123-4. [DOI] [PubMed] [Google Scholar]
- 14.Madani A., Krause B., Greene E., Subramanian S., et al. Large language models generate functional protein sequences across diverse families. Nat Biotechnol. 2023;26 doi: 10.1038/s41587-022-01618-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 15.Hu E.J., Shen Y., Wallis P., Allen-Zhu Z., Li Y., Wang S., et al. Proceedings of ICLR, vol. 1. 2022. Lora: low-rank adaptation of large language models; p. 3. [Google Scholar]
- 16.Ding N., Qin Y., Yang G., Wei F., Yang Z., Su Y., et al. Parameter-efficient fine-tuning of large-scale pre-trained language models. Nat Mach Intell. 2023;5(3):220–235. [Google Scholar]
- 17.Chung H.W., Hou L., Longpre S., Zoph B., Tay Y., Fedus W., et al. Scaling instruction-finetuned language models. J Mach Learn Res. 2024;25(70):1–53. http://jmlr.org/papers/v25/23-0870.html [Google Scholar]
- 18.Sledzieski S., Kshirsagar M., Berger B., Dodhia R., Lavista Ferres J.M. Parameter-efficient fine-tuning of protein language models improves prediction of protein-protein interactions. Machine learning for structural biology workshop; NeurIPS; 2023. pp. 1–10. [Google Scholar]
- 19.Wang M., Patsenker J., Li H., Kluger Y., Kleinstein S. Supervised fine-tuning of pre-trained antibody language models improves antigen specificity prediction. May 2024. https://doi.org/10.1101/2024.05.13.593807 bioRxiv. [DOI] [PMC free article] [PubMed]
- 20.Lafita A., Gonzalez F., Hossam M., Smyth P., Deasy J., Allyn-Feuer A., et al. Fine-tuning protein language models with deep mutational scanning improves variant effect prediction. 2024. arXiv:2405.06729
- 21.Elnaggar A., Heinzinger M., Dallago C., Rehawi G., Wang Y., Jones L., et al. Prottrans: toward understanding the language of life through self-supervised learning. IEEE Trans Pattern Anal Mach Intell. 2022;44(10):7112–7127. doi: 10.1109/TPAMI.2021.3095381. [DOI] [PubMed] [Google Scholar]
- 22.Wenzel M., Grüner E., Strodthoff N. Insights into the inner workings of transformer models for protein function prediction. Bioinformatics. 2024;40(3) doi: 10.1093/bioinformatics/btae031. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 23.Zhou N., Jiang Y., Bergquist T.R., et al. The CAFA challenge reports improved protein function prediction and new functional annotations for hundreds of genes through experimental screens. Genome Biol. 2019;20(244) doi: 10.1186/s13059-019-1835-8. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 24.Heinzinger M., Weissenow K., Sanchez J.G., Henkel A., Mirdita M., Steinegger M., et al. Bilingual language model for protein sequence and structure. 2024. https://doi.org/10.1101/2023.07.23.550085 bioRxiv. [DOI] [PMC free article] [PubMed]
- 25.Xu Y., Liu D., Gong H. Improving the prediction of protein stability changes upon mutations by geometric learning and a pre-training strategy. Nat Comput Sci. 2024;4:840–850. doi: 10.1038/s43588-024-00716-2. [DOI] [PubMed] [Google Scholar]
- 26.Nicolini M., Malchiodi D., Cabri A., Cavalleri E., Mesiti M., Paccanaro A., et al. Proceedings of the 17th international joint conference on biomedical engineering systems and technologies - BIOSTEC. 2024. Fine-tuning of conditional transformers improves the generation of functionally characterized proteins; pp. 561–568. [Google Scholar]
- 27.Tinn R., Cheng H., Gu Y., Usuyama N., Liu X., Naumann T., et al. Fine-tuning large neural language models for biomedical natural language processing. Patterns. 2023;4(4) doi: 10.1016/j.patter.2023.100729. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 28.Schmirler R., Heinzinger M., Rost B. Fine-tuning protein language models boosts predictions across diverse tasks. Nat Commun. 2024;15(7407) doi: 10.1038/s41467-024-51844-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 29.Walsh G., Walsh E. Biopharmaceutical benchmarks 2022. Nat Biotechnol. 2022;40:1722–1760. doi: 10.1038/s41587-022-01582-x. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 30.Yeh A., Norn C., Kipnis Y., et al. De novo design of luciferases using deep learning. Nature. 2023;614:774–780. doi: 10.1038/s41586-023-05696-3. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 31.Munsamy G., Illanes-Vicioso R., Funcillo S., Nakou I.T., Lindner S., Ayres G., et al. Conditional language models enable the efficient design of proficient enzymes. 2024. https://doi.org/10.1101/2024.05.03.592223 bioRxiv.
- 32.Webb E.C., et al. Academic Press; 1992. Enzyme nomenclature 1992. Recommendations of the nomenclature committee of the international union of biochemistry and molecular biology on the nomenclature and classification of enzymes. [Google Scholar]
- 33.Keskar N., McCann B., Varshney L., Xiong C., Socher R. CTRL: a conditional transformer language model for controllable generation. 2019. arXiv:1909.05858 arXiv preprint.
- 34.The UniProt Consortium Uniprot: the universal protein knowledgebase in 2023. Nucleic Acids Res. 2023;51(D1):D523–D531. doi: 10.1093/nar/gkac1052. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 35.Federhen S. The ncbi taxonomy database. Nucleic Acids Res. 2012;40(D1):D136–D143. doi: 10.1093/nar/gkr1178. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 36.Pettit L.D., Powell K. The iupac stability constants database. Chem Int. 2006;56:14–15. [Google Scholar]
- 37.Kingma D., Ba J. Adam: a method for stochastic optimization. 2014. arXiv:1412.6980https://api.semanticscholar.org/CorpusID:6628106 CoRR.
- 38.Henikoff S., Henikoff J.G. Amino acid substitution matrices from protein blocks. Proc Natl Acad Sci. 1992;89(22):10915–10919. doi: 10.1073/pnas.89.22.10915. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 39.Holtzman A., Buys J., Du L., Forbes M., Choi Y. The curious case of neural text degeneration. 2020. arXiv:1904.09751https://doi.org/10.48550/arXiv.1904.09751
- 40.Zhang Y., Skolnick J. Scoring function for automated assessment of protein structure template quality. Proteins, Struct Funct Bioinform. 2004;57(4):702–710. doi: 10.1002/prot.20264. [DOI] [PubMed] [Google Scholar]
- 41.Van Kempen M., Kim S.S., Tumescheit C., Mirdita M., Lee J., Gilchrist C.L., et al. Fast and accurate protein structure search with foldseek. Nat Biotechnol. 2024;42(2):243–246. doi: 10.1038/s41587-023-01773-0. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 42.Yu T., Cui H., Li J.C., Luo Y., Jiang G., Zhao H. Enzyme function prediction using contrastive learning. Science. 2023;379(6639):1358–1363. doi: 10.1126/science.adf2465. [DOI] [PubMed] [Google Scholar]
- 43.Rives A., Meier J., Sercu T., Goyal S., Lin Z., Liu J., et al. Biological structure and function emerge from scaling unsupervised learning to 250 million protein sequences. Proc Natl Acad Sci. 2021;118(15) doi: 10.1073/pnas.2016239118. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 44.Yu H., Deng H., He J., Keasling J.D., Luo X. Unikp: a unified framework for the prediction of enzyme kinetic parameters. Nat Commun. 2023;14(1):8211. doi: 10.1038/s41467-023-44113-1. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 45.Honda S., Shi S., Ueda H.R. Smiles transformer: pre-trained molecular fingerprint for low data drug discovery. 2019. arXiv:1911.04738 arXiv preprint.
- 46.Geurts P., Ernst D., Wehenkel L. Extremely randomized trees. Mach Learn. 2006;63:3–42. doi: 10.1007/s10994-006-6226-1. [DOI] [Google Scholar]
- 47.McInnes, et al. UMAP: uniform manifold approximation and projection. J Open Source Softw. 2018;3(29):861. doi: 10.21105/joss.00861. [DOI] [Google Scholar]
- 48.Liu D.C., Nocedal J. On the limited memory bfgs method for large scale optimization. Math Program. 1989;45(1):503–528. [Google Scholar]
- 49.Altschul S.F., Madden T.L., Schäffer A.A., Zhang J., Zhang Z., Miller W., et al. Gapped blast and psi-blast: a new generation of protein database search programs. Nucleic Acids Res. 1997;25(17):3389–3402. doi: 10.1093/nar/25.17.3389. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 50.Krause A., Stoye J., Vingron M. Large scale hierarchical clustering of protein sequences. BMC Bioinform. 2005;6:1–12. doi: 10.1186/1471-2105-6-15. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 51.Bar-Joseph Z., Gifford D.K., Jaakkola T.S. Fast optimal leaf ordering for hierarchical clustering. Bioinformatics. 2001;17(suppl_1):S22–S29. doi: 10.1093/bioinformatics/17.suppl_1.s22. [DOI] [PubMed] [Google Scholar]
- 52.Häggblom M.M., Bossert I.D. Springer; 2003. Microbial processes and environmental applications. [Google Scholar]
- 53.Gribble G.W., Gribble G. Springer; 1996. Naturally occurring organohalogen compounds—a comprehensive survey. [PubMed] [Google Scholar]
- 54.Liu Y., Jiang L., Wang H., Wang H., Jiao W., Chen G., et al. A brief review for fluorinated carbon: synthesis, properties and applications. Nanotechnol Rev. 2019;8(1):573–586. [Google Scholar]
- 55.Shaw S. Halogenated flame retardants: do the fire safety benefits justify the risks? Rev Environ Health. 2010;25(4):261–306. doi: 10.1515/reveh.2010.25.4.261. [DOI] [PubMed] [Google Scholar]
- 56.Janssen D.B., Dinkla I.J., Poelarends G.J., Terpstra P. Bacterial degradation of xenobiotic compounds: evolution and distribution of novel enzyme activities. Environ Microbiol. 2005;7(12):1868–1882. doi: 10.1111/j.1462-2920.2005.00966.x. [DOI] [PubMed] [Google Scholar]
- 57.Kurihara T., Esaki N. Bacterial hydrolytic dehalogenases and related enzymes: occurrences, reaction mechanisms, and applications. Chem Rec. 2008;8(2):67–74. doi: 10.1002/tcr.20141. [DOI] [PubMed] [Google Scholar]
- 58.Swanson P.E. Dehalogenases applied to industrial-scale biocatalysis. Curr Opin Biotechnol. 1999;10(4):365–369. doi: 10.1016/S0958-1669(99)80066-4. [DOI] [PubMed] [Google Scholar]
- 59.Chan P.W., Yakunin A.F., Edwards E.A., Pai E.F. Mapping the reaction coordinates of enzymatic defluorination. J Am Chem Soc. 2011;133(19):7461–7468. doi: 10.1021/ja200277d. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 60.Kim T.H., Mehrabi P., Ren Z., Sljoka A., Ing C., Bezginov A., et al. The role of dimer asymmetry and protomer dynamics in enzyme catalysis. Science. 2017;355(6322) doi: 10.1126/science.aag2355. [DOI] [PubMed] [Google Scholar]
- 61.Yue Y., Fan J., Xin G., Huang Q., Wang J.-b., Li Y., et al. Comprehensive understanding of fluoroacetate dehalogenase-catalyzed degradation of fluorocarboxylic acids: a qm/mm approach. Environ Sci Technol. 2021;55(14):9817–9825. doi: 10.1021/acs.est.0c08811. [DOI] [PubMed] [Google Scholar]
- 62.Varadi M., Bertoni D., Magana P., Paramval U., Pidruchna I., Radhakrishnan M., et al. Alphafold protein structure database in 2024: providing structure coverage for over 214 million protein sequences. Nucleic Acids Res. 2024;52(D1):D368–D375. doi: 10.1093/nar/gkad1011. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 63.Bateman A., Martin M., Orchard S., Magrane M., Ahmad S., Alpi E., et al. Uniprot: the universal 780 protein knowledgebase in 2023. Nucleic Acids Res. 2022 doi: 10.1093/nar/gkac1052. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 64.Engeland K., Höög J., Holmquist B., Estonius M., Jörnvall H., Vallee B.L. Mutation of arg-115 of human class iii alcohol dehydrogenase: a binding site required for formaldehyde dehydrogenase activity and fatty acid activation. Proc Natl Acad Sci. 1993;90(6):2491–2494. doi: 10.1073/pnas.90.6.2491. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 65.Mi S., Alonso D., Roberts R.J. Functional analysis of gln-237 mutants of hha l methyltransferase. Nucleic Acids Res. 1995;23(4):620–627. doi: 10.1093/nar/23.4.620. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 66.Puyaubert J., Denis L., Alban C. Dual targeting of arabidopsis holocarboxylase synthetase1: a small upstream open reading frame regulates translation initiation and protein targeting. Plant Physiol. 2008;146(2):478. doi: 10.1104/pp.107.111534. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 67.Hauser M., Steinegger M., Söding J. Mmseqs software suite for fast and deep clustering and searching of large protein sequence sets. Bioinformatics. 2016;32(9):1323–1330. doi: 10.1093/bioinformatics/btw006. [DOI] [PubMed] [Google Scholar]
- 68.Schmirler R., Heinzinger M., Rost B. Fine-tuning protein language models boosts predictions across diverse tasks. 2024. https://doi.org/10.1101/2023.12.13.571462 bioRxiv. [DOI] [PMC free article] [PubMed]
- 69.Bansal P., Morgat A., Axelsen K.B., Muthukrishnan V., Coudert E., Aimo L., et al. Rhea, the reaction knowledgebase in 2022. Nucleic Acids Res. 2022;50(D1):D693–D700. doi: 10.1093/nar/gkab1016. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 70.Omelchenko M.V., Galperin M.Y., Wolf Y.I., Koonin E.V. Non-homologous isofunctional enzymes: a systematic analysis of alternative solutions in enzyme evolution. Biol Direct. 2010;5:1–20. doi: 10.1186/1745-6150-5-31. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 71.Nijkamp E., Ruffolo J.A., Weinstein E.N., Naik N., Madani A. Progen2: exploring the boundaries of protein language models. Cell Syst. 2023;14(11):968–978. doi: 10.1016/j.cels.2023.10.002. [DOI] [PubMed] [Google Scholar]
- 72.Bernett J., Blumenthal D.B., Grimm D.G., Haselbeck F., Joeres R., Kalinina O.V., et al. Guiding questions to avoid data leakage in biological machine learning applications. Nat Methods. 2024;21(8):1444–1453. doi: 10.1038/s41592-024-02362-y. [DOI] [PubMed] [Google Scholar]
- 73.Lee J.-Y. The principles and applications of high-throughput sequencing technologies. Dev Reprod. 2023;27(1):9. doi: 10.12717/DR.2023.27.1.9. [DOI] [PMC free article] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Additional text and figures relative to enzyme datasets used in the experiments, model architecture, training time and results on the test set. Three specific sections are also included: a) Only-keywords Finenzyme generation, CLEAN results, and per-class analysis of ESMFold predictions; b) An example of the application of Finenzyme to generate sequences close to a specific natural enzyme; c) Finenzyme generation of hydrolases acting on C-halide compounds (EC 3.8.1).
Data Availability Statement
The Finenzyme code, the scripts to reproduce the experiments and tutorials to fine-tune Finenzyme on any EC category are available from https://github.com/AnacletoLAB/Finenzyme.
The datasets used to train Finenzyme were downloaded from UniProt. Details are available from the notebook tutorial in the above GitHub repository.











