Abstract
Thermal stability is one of the most important properties of enzymes, which sustains life and determines the potential for the industrial application of biocatalysts. Although traditional methods such as directed evolution and classical rational design contribute greatly to this field, the enormous sequence space of proteins implies costly and arduous experiments. The developm ent of enzyme engineering focuses on automated and efficient strategies because of the breakthrough of high-throughput DNA sequencing and machine learning models. In this review, we propose a data-driven architecture for enzyme thermostability engineering and summarize some widely adopted datasets, as well as machine learning-driven approaches for designing the thermal stability of enzymes. In addition, we present a series of existing challenges while applying machine learning in enzyme thermostability design, such as the data dilemma, model training, and use of the proposed models. Additionally, a few promising directions for enhancing the performance of the models are discussed. We anticipate that the efficient incorporation of machine learning can provide more insights and solutions for the design of enzyme thermostability in the coming years.
Keywords: enzyme design, data-driven, machine learning, thermal stability
Introduction
Enzymes that have evolved over 3.5 billion years accelerate chemical reactions for the maintenance of life, adapting to a range of approximately ‒20°C to 120°C [ 1– 4] . The variation in catalytic activity exhibited by different temperature-adapted enzymes demonstrates the complex and critical temperature dependence of biocatalysts [5]. Hence, thermal stability serves as one of the most important factors for the efficient catalytic function of enzymes. Enzymes are considered highly efficient, versatile biocatalysts widely involved in a variety of industrial applications, including food, beverages, biorefineries, pharmaceuticals, and the degradation of toxic environmental pollutants [ 6– 10] . Nevertheless, in practical applications, a number of enzyme catalysis reactions must be performed at extreme temperatures [ 11– 15] . Thus, the thermal stability of enzymes is the main bottleneck in the industrial application of biocatalysts.
Enzyme thermal stability refers to the range of temperatures at which the enzymes can remain thermodynamically stable for catalytic function. For example, the rigid backbone of the 3D structure of hyperthermophilic enzymes can maintain the active and stable conformation to resist irreversible denaturation at extremely high temperatures [16]. Moreover, both thermal stability and thermodynamic stability should be incorporated into the workflow of enzyme design, avoiding misfolding or destabilized conformation of designed target proteins [ 17– 19] .
For decades, early research by enzyme engineering groups focused on improving activity or designing new functions using directed evolution and classical rational design methods [ 20– 26] . However, costly experiments and complicated factors related to protein thermal stability have led to difficulties and limitations in enzyme engineering [ 27– 32] . In 1973, Anfinsen proposed that the protein primary sequence determines the tertiary structure, greatly inspiring the field of protein design [33]. However, the precise and efficient design of enzymes is challenging due to the complex discontinuous mapping relationships among protein sequence, structure, and function. In terms of sequence complexity, for example, a protein of 100 amino acids in length has a sequence space of 20 100 (~10 130), which is well beyond the order of magnitude of proteins that have been sampled in nature (~10 12) [34]. Such an immense protein sequence space has far exceeded the workload that can be covered by directed evolution or traditionally rational design methods [35]. Therefore, significantly reducing the scale of protein sequence space is highly important for efficient and accurate enzyme engineering.
The development of next-generation sequencing has had a significant impact on data-intensive studies in biology. In the data-driven research paradigm, machine learning-based methods for enzyme design have received much attention in recent years [ 36– 38] . Here, we propose an architecture for the thermal stability design of enzymes based on supervised machine learning models ( Figure 1). Compared with traditional enzyme engineering methods, intelligent design strategies can notably improve the efficiency and reduce the experimental cost for enzyme engineering, which has become a promising trend in the context of the upcoming fourth industrial technology revolution [ 39– 42] .
Figure 1 .
Architecture of thermal stability design for enzymes based on machine learning
(A) The model can be trained on diverse datasets, including multiple sequence alignments, protein structures, and enzyme thermal stability labels. Furthermore, data splitting is a typical technical strategy to develop and evaluate machine learning models to improve the robustness and accuracy of the models. Specifically, the training dataset is used to fit the parameters to the machine learning models. Hyperparameters are combined and fine-tuned based on the validation set to improve the generalization performance of the model. Eventually, the test set is responsible for the unbiased evaluation of the machine learning models. (B) Two types of supervised machine learning algorithms provide potentially viable solutions for designing enzyme thermal stability. A large number of existing machine learning models effectively incorporate discriminative models and generative adversarial networks to guide protein engineering [ 43– 47] . Energy-based and machine learning–based methods help filter thermodynamically stable protein sequences. (C) As shown on the left, a variety of metrics are commonly leveraged to evaluate the performance of discriminative models. However, the validation results of biological analysis and experiments govern the validity of the models in enzyme engineering. Apart from the aforementioned methods, the preliminary screening and verification of candidates can be performed through molecular dynamics simulation technology [48]. C is modified from ref. [49].
Overview of Machine Learning
Machine learning relies on analyzing and learning patterns from datasets to update the model parameters for predicting new samples. Additionally, the methods can often provide a reusable and efficient solution when the volume of data or multidimensional complex features is too large to analyze manually or when an automated analysis is desired [50]. In this review, we introduce some basic concepts about traditional machine learning and deep neural networks.
Traditional machine learning models
As an important branch of artificial intelligence (AI), traditional machine learning methods are widely used because of the high interpretability of algorithms and low training costs. Thus, these models are still the primary choice when dealing with emerging disciplines and small-volume datasets ( Figure 2).
Figure 2 .
Machine learning-related terms and methods
(A) Relationship between artificial intelligence, machine learning, and deep learning is illustrated with typical examples. (B) Regression models build a function describing a mathematical relationship between an independent variable (observed temperature) and a dependent variable (predicted temperature) [51]. (C) Classification models, such as support vector machines (SVMs), transform two groups of data in a distinct way as much as possible [52]. (D) Clustering model groups objects with similar characteristics into different clusters. For example, the different optimal temperatures of enzymes can be grouped using multiple sequence alignment (MSA) features [53]. B, C and D are modified from ref. [50].
Deep neural networks
In the 1940s, McCulloch and Pitts first proposed a computational model of neurons, opening up the field of algorithmic research in deep learning [54]. In 1989, the universal approximation theorem proposed by Hornik et al. [55] provided a mathematical proof of deep learning and became one of the important theoretical bases for deep learning algorithms. In the last few years, a variety of different deep learning models have made outstanding contributions to life science [ 56– 59] ( Figure 3). Notably, AlphaFold2 uses MSA Transformer architecture to achieve an unprecedented level of accuracy in the end-to-end prediction of protein 3D structure [60].
Figure 3 .
Neural network methods
(A) A multilayer perceptron (MLP) is a class of fully connected neural network models including the input layer, the hidden layer, and the output layer [61]. In MLP, the neurons (nodes) process information by non-linear activation function [62]. (B) A recurrent neural network (RNN) takes input sequence elements and saves them as a hidden state that contains information related to the previous content. Data such as protein sequences have a sequential order that needs to be followed strictly to carry genetic information. Thus, RNNs are powerful models for processing biological sequence data [63]. (C) Convolutional neural networks (CNNs) use convolutional kernels to perform mathematical operations called convolution, which learn the spatial hierarchy of training data [64]. (D) Graph neural networks (GNNs) are a useful algorithm for leveraging graph-structured knowledge. Biological information, such as protein structure, can be described as a topological relationship among atoms and bonds [65]. (E) As an encoder–decoder architecture, the key components of Transformer consist of a multi-headed self-attention layer, followed by a feedforward layer [66]. Additionally, the performance of Transformer is heavily dependent on the first layer, which is the multi-headed self-attention layer [67]. Vector K, vector Q, and vector V are three different representations for inputs, which are calculated from the corresponding weight matrix to predict the importance of other segments to the current segment. Recently, Transformer has also been found to be potent when dealing with protein sequences and learning the relationships of residues such as interactions within proteins [ 68– 70] .
A large number of open-source toolkits and frameworks are available to facilitate the use of AI technologies [ 71– 82] . The advantage of deep learning over traditional machine learning algorithms is that it automates the learning of features from datasets and eliminates the need for manual feature engineering. Therefore, the size and quality of the datasets intrinsically become some of the dominant factors for deep neural networks.
Datasets Related to the Thermal Stability of Enzymes
Using machine learning to guide the design of the thermal stability of enzymes is inherently complex and challenging. The data-driven approaches to solving this problem, in turn, are limited by high-quality and large-scale datasets. As a result, multiple groups have worked extensively on dataset collection and refinement ( Table 1). We classified them into two categories for different methods of data collection: (1) large-scale data predicted by predictive models and (2) high-quality data collected from the published literature, released datasets, and high-throughput experiments. Some pros and cons are also presented in Table 1.
Table 1 Enzyme thermal stability datasets
|
Dataset description |
Collection method |
Dataset scale |
Advantage |
Disadvantage |
Availability |
Ref. |
|
Tome: Optimal temperature of enzyme |
Predicted by linear models, Bayesian ridge, and support vector regression |
4447 enzyme families, 6,500,000 sequences |
Large-scale protein sequences and family diversity facilitating training of machine learning models; easy-to-access public dataset |
Insufficient accuracy of prediction methods; low accuracy of data on the optimum temperature of enzymes in the extreme temperature range |
https://zenodo.org/record/2539114 |
|
|
The optimal growth temperature of bacteria and the optimal temperature of enzyme |
Correlation analysis between enzyme temperature optima and the organism growth temperatures |
21,498 OGT of bacteria |
UniProt sequences covered by 43%; easy-to-access public dataset; covering a variety of temperature-adapted bacteria and archaea. |
Stricter quality control needed; not actively maintained |
https://doi.org/10.5281/zenodo.1175608 |
|
|
BRENDA: The optimal temperature of enzymes |
Published papers in PubMed |
32,000,000 sequences with 41,000 optimal temperature labels |
High-quality data collected from the published literature; actively maintained; convenient web interface |
Relatively small proportion of temperature-labeled sequences in protein families; no obviously experimental conditions |
https://www.brenda-enzymes.org/ |
|
|
ThermoMutDB: Environment, ΔΔ G, Δ T m |
Manually collected from published papers |
14,669 mutations across 588 proteins |
High-quality data collected from the published literature; convenient web interface; a variety of protein stability parameters available; continuously maintained and updated |
Limited native protein properties; few sequence data from archaea; small sequence coverage |
http://biosig.unimelb.edu.au/thermomutdb |
|
|
ProThermDB: Mutant thermal stability data |
High-throughput experiments |
More than 32,000 proteins and 120,000 thermal stability data |
Relatively greater variety of protein sequences; extensive high-quality data from experiments; convenient web interface; continuously maintained |
Limited coverage of organisms |
https://web.iitm.ac.in/bioinfo2/prothermdb/index.html |
|
|
FireProt DB: Mutant thermal stability data |
Manually collected from published papers |
237 proteins, 13,274 entries |
Convenient web interface; high-quality data; multiple sources of annotations; |
Limited wild sequence diversity |
https://loschmidt.chemi.muni.cz/fireprotdb |
|
|
Single-site mutations thermal stability data: T m, Δ T m, and Δ H |
Manually collected from published papers for experimentally measured results |
90 wild sequences, 1626 mutant sequences |
Various and high-quality thermal stability data; comprehensive quality control |
Not easily reusable dataset; not actively maintained; limited number of wild sequences |
The Appendix of the article. |
Large-scale datasets from predictive models
First, Li et al. [51] predicted the optimal temperature of 6.5 million enzyme sequences in 4447 enzyme families by constructing several machine learning models, such as linear models, Bayesian ridge, and support vector regression. These models were trained on the growth temperature of microorganisms and the optimal temperature of enzyme sequences in the existing dataset.
Second, Engqvist used the optimal growth temperature (OGT) of 21,498 microorganisms, which contained both bacteria and archaea, spanning a wide range of microorganisms from cryophiles, thermophiles, and hyperthermophiles for predicting the optimum temperature of enzymes [83]. Based on regression models, this dataset enabled the large-scale annotation of protein sequences, which covered approximately 43% of the protein sequences released in UniProt.
High-quality datasets from published papers and datasets
The prediction of thermal stability data using different algorithms reduces the cost of data collection. However, the quality of the datasets obtained by algorithmic prediction is still not as good as the quality of the manually collected datasets. Therefore, some other researchers constructed the thermal stability dataset of proteins by manually collecting experimental results from the published literature or using high-throughput techniques.
The BRENDA database, a widely used and accessed enzyme annotation database, provides a large amount of hand-curated data on the function and properties of enzymes, including more than 32 million sequences, optimal temperature values for approximately 41,000 enzymes, and temperature stability parameters for approximately 26,000 enzyme sequences [84]. However, it is limited by the high cost of experiments for measuring the thermal stability of enzymes, resulting in a relatively small proportion of labelled sequences in protein families.
The ThermoMutDB missense mutant database contains 14,669 wild-type and mutant thermodynamic data for more than 588 proteins manually collected from the published literature, such as melting temperature ( T m) and Gibbs free energy [85]. It facilitates downstream tasks by providing database retrieval and usage services through a web interface. By leveraging high-throughput experimental techniques, ProThermDB has more than 32,000 data points on protein thermal stability, such as T m and Δ T m. The database contains more than 120,000 thermodynamic labels of wild types and mutants [86]. FireProt DB contains the thermal stability data for 740 proteins with more than 25,000 single-point mutants. These data are derived from released datasets and published papers [87].
Pucci et al. [88] manually constructed a dataset describing the thermal stability of 31 proteins with single-point mutations, including experimentally determined melting temperature ( T m), Δ T m, Δ G, and Δ H. The principle of data selection was as follows: (1) only single-point mutant sequences, (2) sequences containing protein structures with a resolution of 2.5 Å or less, (3) only data where monomeric proteins were experimentally determined and characterized, (4) protein sequences where the folded and unfolded states of the wild type and mutants were clearly described in the literature, and (5) mutant sequences that caused significant structural changes with temperature ranges beyond 20°C were not included.
The quality of the datasets, which were manually filtered and constructed, was indeed relatively superior. However, the current goal of wild-type–mutant dataset construction led to limited family classes of proteins, hampering the analysis based on evolutionary adaptation. Nonetheless, this problem will be improved with the development of high-throughput technology in the future [89].
Machine Learning-based Methods and Strategies for Thermal Stability Design of Enzymes
Combining artificial intelligence with high-throughput assays [ 90, 91] , a large number of studies have made important contributions to promote the design of enzyme thermostability. We review three aspects of data-driven strategies for designing enzyme thermal stability, consider recent advances in enzyme thermal stability design, and discuss methods, input data, method reuse, and limitations.
Prediction of the thermal stability of enzymes
As described in the dataset section, machine learning plays a significant role in dataset augmentation. Furthermore, the ability to accurately predict enzyme thermostability paves the way toward efficient enzyme engineering. In 2009, Ku et al. [92] used a statistical method trained on 35 different protein sequences to predict the melting temperature classification directly from protein sequence analysis. The predictor was tested on 75 genomes, including approximately 150,000 proteins, to predict the percentage of high-melting-temperature proteins. The results suggested a correlation between the dipeptide constitution and the melting temperature index. From the perspective of the scale of the training dataset, the very limited amount of data can easily lead to overfitting of the model, which severely limits the generalization ability of the model.
Ten years later, Li et al. [51] leveraged the BRENDA database to predict the temperature optima for enzymes based on regression models, but the optimal temperature of enzymes in the training dataset followed a normal distribution. Hence, the optimal temperature data above 85°C was less than 5%, limiting the prediction for thermostable enzymes. Therefore, Gado et al. [93] used the ensemble learning and resampling strategy to decrease the mean squared error and increase the overall R 2 to expand the application range of the predictor. Importantly, in this work, a well-defined accuracy function was applied to narrow the prediction error. The trained model and the code are available on GitHub to facilitate model reuse. The advantages of TOMER are that it is more robust and accurate due to the large training dataset and well-designed resampling strategies.
Recently, a pretrained deep neural network model, DeepET, provided a method for enzyme thermal stability representation [94]. The pretrained model was developed on a dataset of protein sequences with more than 3 million labels of OGT. As a result, the transfer learning models achieved an R 2 of 0.73 on the test dataset. The attractive advantage of DeepET is that the pretrained model can be reused for the design of downstream tasks and precise prediction with minimum training cost. In addition, the pretrained model can also be used to explore the sequence features that determine the thermal stability factors.
Apart from the regression models for thermal stability prediction, recently, several classification models have shown promising applications for identifying thermostability enzymes, such as iThermo [61] and TMPpred [95]. Shahraki et al. [52] presented a sequence-based machine learning model, TaXyl, to classify xylanases from GH10 and GH11 families with different thermal adaptations. This classifier identified three hyperthermophilic xylanases, with maximum activities of 57%–90% at 100°C and 20 min of incubation. Additionally, multilayer perceptron, support vector machine (SVM), and deep learning models were employed to identify thermally stable cellulases [96] and chitinase [97]. Wand et al. [53] analyzed protein sequences with kernel principal component analysis and SVM models to identify thermophilic proteins. One of the greatest advantages of these models is that they are user-friendly. Simply taking the protein sequences as input yields a corresponding prediction of thermal stability. On the other hand, these models, trained on specific datasets, are almost only accurate in predicting similar protein sequences. Therefore, sequence families and species need to be considered when using these models. In the high-throughput annotation platform of omics data, these methods can be combined to explore more valuable proteins.
Prediction of the changes in the stability of mutant proteins
Whether it is functional improvement or design for new enzymes, possessing a thermodynamically stable structure of the enzyme is a basic goal of enzyme engineering [98]. Additionally, the prediction of thermal stability changes upon mutation advances the engineering of enzyme thermal stability design. Giollo et al. [99] provided NeEMO for stability changes (ΔΔ G) of mutants, which was trained on a large dataset of multiple sequence alignments and protein structures. MAESTROweb is freely accessible on the web server for structure-based protein stability prediction [100]. Another structure-based model, HoTMuSiC, used artificial neural networks with particular activation functions to predict thermal stability changes upon point mutations [101]. Based on a double-checked and refined training dataset, Yang et al. [102] trained PON-tstab to predict protein variant stability with 1106 collected features. In this study, several errors and issues with ProTherm [103] entries were proposed, including sequence and structural differences and some errors in stability data. Cao et al. [104] developed DeepDDG, which achieved a Pearson correlation coefficient of 0.48–0.56 for three independent test sets based on neural networks. Apart from these tools, several webserver-based tools were used to predict stability changes of mutant proteins, including iStable 2.0 [105], MPTherm-pred [106], and SAAFEC-SEQ [107]. The advantage of the structure-based model, over models trained on sequence-only datasets, improved the model performance by directly embedding structure information corresponding to the function and integrating more expert knowledge into the model.
Certainly, machine learning models provide an efficient way to guide enzyme stability design. Therefore, with the increase in the number of existing prediction models, some studies summarized and assessed these models using benchmark tests [ 108– 111] . At present, these methods have made great breakthroughs and still have great potential for improvement in generalization ability and accuracy.
Applications of machine learning-driven methods for designing thermostable enzymes
The ability of machine learning to reduce dimensionality enables enzyme engineering to find the sequence composition of targets in a smaller sequence space. Romero et al. [112] used Bayesian decision theory, Gaussian process, and the structure-based model for guiding the P450 enzyme thermostability design, which identified sequences that improved the thermal stability by sampling the high-dimensional space of protein sequences, resulting in an average increase of 5.1°C in T 50. For the details of the work, protein structures with at least 50% sequence identity were represented as pairwise interactions between amino acids. However, the disadvantage of this approach is that training and prediction on large datasets is expensive due to the high computational complexity of the algorithm.
Lu et al. [64] proposed a self-supervised CNN, MutCompute, to design thermally stable PETase (enzymes that catalyze the hydrolysis of polyethylene terephthalate) by predicting the amino acid positions and types in enzyme sequences that can potentially improve thermal stability. Among all mutants predicted using the model, FAST-PETase with five mutation sites showed the greatest improvement in thermal stability and catalytic efficiency. Compared with wild-type ThermoPETase, the degradation activity of FAST-PETase is increased by 2.4-fold at 40°C and 38-fold at 50°C. The training dataset used by MutCompute contained more than 19,000 protein structures in the Protein Data Bank (PDB) database. The detailed steps of the model are as follows. First, the model creates a microenvironment around the arbitrarily selected central residue within the enzyme and masks all atoms consisting of the central residue, allowing the neural network to predict the type of central residue. Second, the environment contains embedded representation using seven different channels. Third, the trained multilayer CNN network is used to calculate the discrete probability distribution of 20 amino acids at each position in the protein structure. Eventually, if the amino acid types in the corresponding positions are identical, it is considered to be favorable. Otherwise, it is assumed that they need to be optimized. The model architecture that abstracts the complete protein structure into the atomic composition and residue microenvironment provides an enlightening idea for embedding the association between protein structure and function.
Some hybrid methods that combine energy calculation and sequence evolution analysis with machine learning methods have shown strong potential in enzyme thermal stability design [ 113, 114] . The principal disadvantage of hybrid methods, over one predictor, is complicated compiling and environmental setting up for strategy implementation in silico. Conversely, the hybrid strategy of combining multiple methods is more universal and feasible in practical applications. The GRAPE strategy screens potentially stable sequences by constructing a single-point mutation library of enzyme molecules. Collaborating with the clustering and greedy algorithm, the T m value of DuraPETase is increased by 31°C, and enzymatic PET degradation is increased by more than 300- fold at 37°C [115]. Barber-Zucker et al. [116] obtained four mutants of versatile peroxidases, exhibiting better resistance in a high-temperature environment relative to the wild type, in which 43 multiple mutations were designed using AlphaFold2 [60], trRosetta [ 117, 118] and PROSS [119]. In this case, the collaboration of evolutionary analysis, protein design models, and accurate structure prediction models was used to proceed with efficient enzyme design.
Pinney et al. [120] used a logistic regression model to analyze the embedded enzyme sequences of 1005 enzyme families collected from 5864 bacterial species, resulting in more than 150,000 residues related to the optimal temperature of enzymes. The advantage of this study is that it used a large number of homologous enzyme sequences collected from different protein families and species to illustrate the relationship between temperature adaptability and sequence composition. The output results can be further applied to enzyme design or analyzing the evolutionary strategies of thermal stability for other enzymes. Inspiringly, Singer et al. [121] developed a neural network for novel protein sequence generation and protein stability prediction by giving the target secondary structure backbone. The generated thermally stable proteins retained their unfolded and highly thermostable conformations even at temperatures over 99°C. Thus, this approach expands the sequence space for designing thermostable enzymes in a large-scale and efficient manner.
Although machine learning-driven methods of enzyme thermostability design are limited by data quality and model innovation, current methods and strategies have demonstrated efficient and large-scale design capabilities. It is worth mentioning that the machine learning models are trained on a biased dataset; thus, the applicability of the existing models needs to be noted [122]. Finally, integrating expert knowledge into the machine learning model greatly improves the accuracy and generalization ability [123], especially the computable factors that affect the thermal stability of enzymes and the basic principles of biology [ 18, 32, 124, 125] . Figure 4 summarizes these advanced methods and strategies in a timeline. Most enzyme thermostability prediction methods take only sequence-derived information as input, while thermodynamic stability prediction methods add additional structural information to the models. Moreover, an increasing number of methods are available in the form of online web interfaces for related services. This is a preferable way to use the methods for researchers with no programming background or no experience in using Linux operating systems.
Figure 4 .
Advanced methods and strategies for enzyme thermal stability design
Three aspects of methods are classified by blue, yellow and red blocks. Sequence-only and sequence-structure-based methods are distinguished by blue and red circles, respectively. Red citations indicate that the methods or strategies are freely accessible via web servers.
Challenges for Applying Machine Learning in Enzyme Thermostability Design
The employment of machine learning to cope with the accumulation of biological data brings not only new opportunities for enzyme design but also new challenges for corresponding data management and model training. Here, we present the challenges that may be encountered in guiding enzyme thermal stability design using machine learning.
Issues and challenges related to datasets
Insufficient, irrelevant, and imbalanced data points hinder the training of machine learning models. When the input label is not constant for the molecular characteristics of the enzyme but varies with the experimental conditions, this type of data can cause the trained model to deviate significantly from the task objective [126]. Therefore, corrective and standardized management should be established to improve the quality of the database.
Challenges in training models
The training process of the model requires monitoring the relevant metrics to prevent over- and underfitting of the models. When overfitting occurs, the model has been overfitted to the training dataset, which reduces the model’s ability to generalize to unseen data. To prevent overfitting, strategies such as data augmentation, dropout, cross-validation, and multi-model combination are usually helpful [127]. The underfitting situation is generally caused by insufficient training data or high model complexity, which can be countered by appropriately reducing the model complexity.
In recent years, protein language models, a kind of pre-training model, such as ESM-1b and ProtGPT2, have been gaining attention as sequence feature extractors, which can greatly reduce the training cost of models and improve the performance of models for downstream tasks [ 128– 130] . This approach is known as transfer learning and has been used in the past in natural language processing fields, such as bidirectional encoder representations from transformers (BERT) [131].
Interpretability of machine learning models for enzyme design
As previously mentioned, traditional machine learning methods are the primary choice when facing a new field, benefiting from the interpretability of models and prediction results. In contrast, the inherent black-box modeling framework of deep learning leads to weak interpretability and challenges the complex design of enzymes. Therefore, interpretable neural networks should be given more attention, which is crucial for the field of enzyme design [132].
Evaluation and utilization of existing models
A mature machine learning model must be subject to a complex evaluation process that includes statistical evaluation, biochemical experiments, and computational simulations. For statistical evaluation, the choice of metrics is critical. Both biochemical experiments and fine-grained computational simulations can be relatively costly. Therefore, high-throughput techniques and sophisticated experimental design can largely improve experimental efficiency.
A large number of models are available for the design of enzymes with thermal stability. However, the use of these models is hampered by the fact that most of these methods require skills to use the Linux operating system and some programming languages. Furthermore, the differences in the interfaces of the methods and the required compilation environment dependencies make it difficult for researchers to evaluate and select the best one among different methods. Nevertheless, similar platforms exist in other areas of life science that provide informative insights into this issue. For example, critical assessment of protein structure prediction (CASP) and critical assessment of protein function annotation algorithms (CAFA) provide benchmarks for evaluating computational methods of protein structure and function, respectively [ 133, 134] .
Conclusion and Future Direction
This review summarizes recent advances in machine learning–based methods for designing the thermal stability of enzymes. Among the introduced related datasets, a large amount of data was obtained by model prediction and manually collected from published literature. These available and accessible datasets provide a rich material basis for the training of machine learning models. After decades of development, models using different data, such as sequences, structures, and thermodynamics data, have contributed to the design of the thermal stability of enzymes. In addition, machine learning-based methods provide a sustained and efficient way of exploring the sequence- or structure-fitness landscape of proteins [ 135, 136] . We also presented some challenges in applying machine learning for designing enzyme thermostability.
The revolutionary breakthrough of AlphaFold2, which has achieved high accuracy and efficiency in protein structure prediction, promises to be the fourth most efficient way to obtain protein structures [ 60, 137] . AlphaFold2 recently open-sourced more than 200 million protein structures, bridging the data gap between protein sequence and structure and enabling the incorporation of protein structure information into machine learning models [ 138, 139] . This breakthrough laid the foundation for the genome-wide applications of protein structure-based artificial intelligence models [140]. Diverse successful applications that provide feasible solutions to the design of the thermal stability of enzymes are available. The unified and standardized platform for managing and comparing data resources and models is highly conducive to large-scale utilization and promotion.
Approximately 20 years ago, researchers leveraged whole-genome sequencing technology to obtain the genetic code for encoding life [ 141, 142] . We now stand at a critical point between the explosive generation of big biological data and the booming field of machine learning. A large number of techniques have enabled us to completely understand and decode how nature creates life. Harnessing machine learning models to observe the flow of information about life provides us with a powerful grip to find the patterns behind it. We look forward to the next groundbreaking advances in the field of enzyme design achieved by machine learning.
Supporting information
Acknowledgments
The scientific calculations in this paper were performed on the HPC Cloud Platform of Shandong University. We would also like to express our sincere appreciation to the reviewers and editors for their valuable feedback and assistance in enhancing the quality of this publication, which played a vital role in its success.
COMPETING INTERESTS
The authors declare that they have no conflict of interest.
Funding Statement
This work was supported by the grants from the National Natural Science Foundation of China (No. 32100022) and the Key Research and Development Program of Shandong Province (No. 2020CXGC010601).
References
- 1.Kashefi K, Lovley DR. Extending the upper temperature limit for life. Science. . 2003;301:934. doi: 10.1126/science.1086823. [DOI] [PubMed] [Google Scholar]
- 2.Mykytczuk NCS, Foote SJ, Omelon CR, Southam G, Greer CW, Whyte LG. Bacterial growth at –15°C; molecular insights from the permafrost bacterium Planococcus halocryophilus Or1. ISME J. . 2013;7:1211–1226. doi: 10.1038/ismej.2013.8. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 3.Price PB, Sowers T. Temperature dependence of metabolic rates for microbial growth, maintenance, and survival. Proc Natl Acad Sci USA. . 2004;101:4631–4636. doi: 10.1073/pnas.0400522101. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 4.Wolfenden R, Snider MJ. The depth of chemical time and the power of enzymes as catalysts. Acc Chem Res. . 2001;34:938–945. doi: 10.1021/ar000058i. [DOI] [PubMed] [Google Scholar]
- 5.Arcus VL, Mulholland AJ. Temperature, dynamics, and enzyme-catalyzed reaction rates. Annu Rev Biophys. . 2020;49:163–180. doi: 10.1146/annurev-biophys-121219-081520. [DOI] [PubMed] [Google Scholar]
- 6.Wu S, Snajdrova R, Moore JC, Baldenius K, Bornscheuer UT. Biocatalysis: enzymatic synthesis for industrial applications. Angew Chem Int Ed. . 2021;60:88–119. doi: 10.1002/anie.202006648. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 7.Saravanan A, Kumar PS, Vo DVN, Jeevanantham S, Karishma S, Yaashikaa PR. A review on catalytic-enzyme degradation of toxic environmental pollutants: Microbial enzymes. J Hazard Mater. . 2021;419:126451. doi: 10.1016/j.jhazmat.2021.126451. [DOI] [PubMed] [Google Scholar]
- 8.Fryszkowska A, Devine PN. Biocatalysis in drug discovery and development. Curr Opin Chem Biol. . 2020;55:151–160. doi: 10.1016/j.cbpa.2020.01.012. [DOI] [PubMed] [Google Scholar]
- 9.Champreda V, Mhuantong W, Lekakarn H, Bunterngsook B, Kanokratana P, Zhao XQ, Zhang F, et al. Designing cellulolytic enzyme systems for biorefinery: from nature to application. J Biosci Bioeng. . 2019;128:637–654. doi: 10.1016/j.jbiosc.2019.05.007. [DOI] [PubMed] [Google Scholar]
- 10.Planas-Iglesias J, Marques SM, Pinto GP, Musil M, Stourac J, Damborsky J, Bednar D. Computational design of enzymes for biotechnological applications. Biotechnol Adv. . 2021;47:107696. doi: 10.1016/j.biotechadv.2021.107696. [DOI] [PubMed] [Google Scholar]
- 11.Parvizpour S, Hussin N, Shamsir MS, Razmara J. Psychrophilic enzymes: structural adaptation, pharmaceutical and industrial applications. Appl Microbiol Biotechnol. . 2021;105:899–907. doi: 10.1007/s00253-020-11074-0. [DOI] [PubMed] [Google Scholar]
- 12.Arbab S, Ullah H, Khan MIU, Khattak MNK, Zhang J, Li K, Hassan IU. Diversity and distribution of thermophilic microorganisms and their applications in biotechnology. J Basic Microbiol. . 2022;62:95–108. doi: 10.1002/jobm.202100529. [DOI] [PubMed] [Google Scholar]
- 13.Ajeje SB, Hu Y, Song G, Peter SB, Afful RG, Sun F, Asadollahi MA, et al. Thermostable cellulases / xylanases from thermophilic and hyperthermophilic microorganisms: current perspective. Front Bioeng Biotechnol. . 2021;9 doi: 10.3389/fbioe.2021.794304. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 14.Vivek K, Sandhia GS, Subramaniyan S. Extremophilic lipases for industrial applications: a general review. Biotechnol Adv. . 2022;60:108002. doi: 10.1016/j.biotechadv.2022.108002. [DOI] [PubMed] [Google Scholar]
- 15.Zhu D, Adebisi WA, Ahmad F, Sethupathy S, Danso B, Sun J. Recent development of extremophilic bacteria and their application in biorefinery. Front Bioeng Biotechnol. . 2020;8:483. doi: 10.3389/fbioe.2020.00483. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 16.Vieille C, Zeikus GJ. Hyperthermophilic enzymes: sources, uses, and molecular mechanisms for thermostability. Microbiol Mol Biol Rev. . 2001;65:1–43. doi: 10.1128/MMBR.65.1.1-43.2001. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 17.Kuhlman B, Bradley P. Advances in protein structure prediction and design. Nat Rev Mol Cell Biol. . 2019;20:681–697. doi: 10.1038/s41580-019-0163-x. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 18.Goldenzweig A, Fleishman SJ. Principles of protein stability and their application in computational design. Annu Rev Biochem. . 2018;87:105–129. doi: 10.1146/annurev-biochem-062917-012102. [DOI] [PubMed] [Google Scholar]
- 19.Marabotti A, Scafuri B, Facchiano A. Predicting the stability of mutant proteins by computational approaches: an overview. Briefings BioInf. . 2021;22:bbaa074. doi: 10.1093/bib/bbaa074. [DOI] [PubMed] [Google Scholar]
- 20.Musil M, Konegger H, Hon J, Bednar D, Damborsky J. Computational design of stable and soluble biocatalysts. ACS Catal. . 2018;9:1033–1054. doi: 10.1021/acscatal.8b03613. [DOI] [Google Scholar]
- 21.Romero-Rivera A, Garcia-Borràs M, Osuna S. Computational tools for the evaluation of laboratory-engineered biocatalysts. Chem Commun. . 2017;53:284–297. doi: 10.1039/C6CC06055B. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 22.Arnold FH. The nature of chemical innovation: new enzymes by evolution. Quart Rev Biophys. . 2015;48:404–410. doi: 10.1017/S003358351500013X. [DOI] [PubMed] [Google Scholar]
- 23.Xiong W, Liu B, Shen Y, Jing K, Savage TR. Protein engineering design from directed evolution to de novo synthesis. Biochem Eng J. . 2021;174:108096. doi: 10.1016/j.bej.2021.108096. [DOI] [Google Scholar]
- 24.Nirantar SR. Directed evolution methods for enzyme engineering. Molecules. . 2021;26:5599. doi: 10.3390/molecules26185599. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 25.Steipe B, Schiller B, Plückthun A, Steinbacher S. Sequence statistics reliably predict stabilizing mutations in a protein domain. J Mol Biol. . 1994;240:188–192. doi: 10.1006/jmbi.1994.1434. [DOI] [PubMed] [Google Scholar]
- 26.Siddiqui KS, Cavicchioli R. Cold-adapted enzymes. Annu Rev Biochem. . 2006;75:403–433. doi: 10.1146/annurev.biochem.75.103004.142723. [DOI] [PubMed] [Google Scholar]
- 27.Maffucci I, Laage D, Sterpone F, Stirnemann G. Thermal adaptation of enzymes: impacts of conformational shifts on catalytic activation energy and optimum temperature. Chem Eur J. . 2020;26:10045–10056. doi: 10.1002/chem.202001973. [DOI] [PubMed] [Google Scholar]
- 28.Timr S, Madern D, Sterpone F. Protein thermal stability. Prog Mol Biol Transl Sci. 2020, 170: 239–272 . [DOI] [PubMed]
- 29.Liao M, Somero GN, Dong Y. Comparing mutagenesis and simulations as tools for identifying functionally important sequence changes for protein thermal adaptation. Proc Natl Acad Sci USA. . 2019;116:679–688. doi: 10.1073/pnas.1817455116. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 30.Beadle BM, Shoichet BK. Structural bases of stability–function tradeoffs in enzymes. J Mol Biol. . 2002;321:285–296. doi: 10.1016/S0022-2836(02)00599-5. [DOI] [PubMed] [Google Scholar]
- 31.Tawfik DS. Accuracy-rate tradeoffs: how do enzymes meet demands of selectivity and catalytic efficiency? Curr Opin Chem Biol. 2014, 21: 73–80 . [DOI] [PubMed]
- 32.Teufl M, Zajc CU, Traxlmayr MW. Engineering strategies to overcome the stability–function trade-off in proteins. ACS Synth Biol. . 2022;11:1030–1039. doi: 10.1021/acssynbio.1c00512. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 33.Anfinsen CB. Principles that govern the folding of protein chains. Science. . 1973;181:223–230. doi: 10.1126/science.181.4096.223. [DOI] [PubMed] [Google Scholar]
- 34.Baker D. What has de novo protein design taught us about protein folding and biophysics? Protein Sci. 2019, 28: 678–683 . [DOI] [PMC free article] [PubMed]
- 35.Zeymer C, Hilvert D. Directed evolution of protein catalysts. Annu Rev Biochem. . 2018;87:131–157. doi: 10.1146/annurev-biochem-062917-012034. [DOI] [PubMed] [Google Scholar]
- 36.Anishchenko I, Pellock SJ, Chidyausiku TM, Ramelot TA, Ovchinnikov S, Hao J, Bafna K, et al. De novo protein design by deep network hallucination. Nature. . 2021;600:547–552. doi: 10.1038/s41586-021-04184-w. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 37.Tischer D, Lisanza S, Wang J, Dong R, Anishchenko I, Milles LF, Ovchinnikov S, et al. Design of proteins presenting discontinuous functional sites using deep learning. Biorxiv. 2020, doi: https://doi.org/10.1101/2020.11.29.402743
- 38.Wang J, Lisanza S, Juergens D, Tischer D, Anishchenko I, Baek M, Watson JL, et al. Deep learning methods for designing proteins scaffolding functional sites. Biorxiv. 2021, doi: https://doi.org/10.1101/2021.11.10.468128
- 39.Fox R, Roy A, Govindarajan S, Minshull J, Gustafsson C, Jones JT, Emig R. Optimizing the search algorithm for protein engineering by directed evolution. Protein Eng Des Sel. . 2003;16:589–597. doi: 10.1093/protein/gzg077. [DOI] [PubMed] [Google Scholar]
- 40.Wu Z, Kan SBJ, Lewis RD, Wittmann BJ, Arnold FH. Machine learning-assisted directed protein evolution with combinatorial libraries. Proc Natl Acad Sci USA. . 2019;116:8852–8858. doi: 10.1073/pnas.1901979116.1902.07231 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 41.Rocklin GJ, Chidyausiku TM, Goreshnik I, Ford A, Houliston S, Lemak A, Carter L, et al. Global analysis of protein folding using massively parallel design, synthesis, and testing. Science. . 2017;357:168–175. doi: 10.1126/science.aan0693. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 42.Currin A, Swainston N, Day PJ, Kell DB. Synthetic biology for the directed evolution of protein biocatalysts: navigating sequence space intelligently. Chem Soc Rev. . 2015;44:1172–1239. doi: 10.1039/C4CS00351A. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 43.Wu Z, Johnston KE, Arnold FH, Yang KK. Protein sequence design with deep generative models. Curr Opin Chem Biol. . 2021;65:18–27. doi: 10.1016/j.cbpa.2021.04.004. [DOI] [PubMed] [Google Scholar]
- 44.Yang KK, Wu Z, Arnold FH. Machine-learning-guided directed evolution for protein engineering. Nat Methods. . 2019;16:687–694. doi: 10.1038/s41592-019-0496-6. [DOI] [PubMed] [Google Scholar]
- 45.Repecka D, Jauniskis V, Karpus L, Rembeza E, Rokaitis I, Zrimec J, Poviloniene S, et al. Expanding functional protein sequence spaces using generative adversarial networks. Nat Mach Intell. . 2021;3:324–333. doi: 10.1038/s42256-021-00310-5. [DOI] [Google Scholar]
- 46.Ingraham J, Garg V, Barzilay R, Jaakkola T. Generative models for Graph-based protein design. Proc Adv Neural Inf Process Syst. 2019, 32: 15820–15831
- 47.Lopez R, Gayoso A, Yosef N. Enhancing scientific discoveries in molecular biology with deep generative models. Mol Syst Biol. . 2020;16:e9198. doi: 10.15252/msb.20199198. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 48.Karplus M, McCammon JA. Molecular dynamics simulations of biomolecules. Nat Struct Biol. . 2002;9:646–652. doi: 10.1038/nsb0902-646. [DOI] [PubMed] [Google Scholar]
- 49.Mazurenko S, Prokop Z, Damborsky J. Machine learning in enzyme engineering. ACS Catal. . 2019;10:1210–1223. doi: 10.1021/acscatal.9b04321. [DOI] [Google Scholar]
- 50.Greener JG, Kandathil SM, Moffat L, Jones DT. A guide to machine learning for biologists. Nat Rev Mol Cell Biol. . 2022;23:40–55. doi: 10.1038/s41580-021-00407-0. [DOI] [PubMed] [Google Scholar]
- 51.Li G, Rabe KS, Nielsen J, Engqvist MKM. Machine learning applied to predicting microorganism growth temperatures and enzyme catalytic optima. ACS Synth Biol. . 2019;8:1411–1420. doi: 10.1021/acssynbio.9b00099. [DOI] [PubMed] [Google Scholar]
- 52.Foroozandeh Shahraki M, Farhadyar K, Kavousi K, Azarabad MH, Boroomand A, Ariaeenejad S, Hosseini Salekdeh G. A generalized machine‐learning aided method for targeted identification of industrial enzymes from metagenome: a xylanase temperature dependence case study. Biotechnol Bioeng. . 2021;118:759–769. doi: 10.1002/bit.27608. [DOI] [PubMed] [Google Scholar]
- 53.Wang X-F, Gao P, Liu Y-F, Li H-F, Lu F. Predicting thermophilic proteins by machine learning. Curr Bioinform. 2020, 15: 493–502
- 54.McCulloch WS, Pitts W. A logical calculus of the ideas immanent in nervous activity. Bull Math Biophys. . 1943;5:115–133. doi: 10.1007/BF02478259. [DOI] [PubMed] [Google Scholar]
- 55.Hornik K, Stinchcombe M, White H. Multilayer feedforward networks are universal approximators. Neural Networks. . 1989;2:359–366. doi: 10.1016/0893-6080(89)90020-8. [DOI] [Google Scholar]
- 56.Renaud N, Geng C, Georgievska S, Ambrosetti F, Ridder L, Marzella DF, Réau MF, et al. DeepRank: a deep learning framework for data mining 3D protein-protein interfaces. Nat Commun. . 2021;12:1–8. doi: 10.1038/s41467-021-27396-0. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 57.Bileschi ML, Belanger D, Bryant DH, Sanderson T, Carter B, Sculley D, Bateman A, et al. Using deep learning to annotate the protein universe. Nat Biotechnol. . 2022;40:932–937. doi: 10.1038/s41587-021-01179-w. [DOI] [PubMed] [Google Scholar]
- 58.Shen J, Liu F, Tu Y, Tang C. Finding gene network topologies for given biological function with recurrent neural network. Nat Commun. . 2021;12:3125. doi: 10.1038/s41467-021-23420-5. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 59.Luo Y, Jiang G, Yu T, Liu Y, Vo L, Ding H, Su Y, et al. ECNet is an evolutionary context-integrated deep learning framework for protein engineering. Nat Commun , 2021, 12: 5743 . [DOI] [PMC free article] [PubMed]
- 60.Jumper J, Evans R, Pritzel A, Green T, Figurnov M, Ronneberger O, Tunyasuvunakool K, et al. Highly accurate protein structure prediction with AlphaFold. Nature. . 2021;596:583–589. doi: 10.1038/s41586-021-03819-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 61.Ahmed Z, Zulfiqar H, Khan AA, Gul I, Dao FY, Zhang ZY, Yu XL, et al. iThermo: a sequence-based model for identifying thermophilic proteins using a multi-feature fusion strategy. Front Microbiol. . 2022;13:790063. doi: 10.3389/fmicb.2022.790063. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 62.Crick F. The recent excitement about neural networks. Nature. . 1989;337:129–132. doi: 10.1038/337129a0. [DOI] [PubMed] [Google Scholar]
- 63.Griffith D, Holehouse AS. PARROT is a flexible recurrent neural network framework for analysis of large protein datasets. eLife. . 2021;10:e70576. doi: 10.7554/eLife.70576. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 64.Lu H, Diaz DJ, Czarnecki NJ, Zhu C, Kim W, Shroff R, Acosta DJ, et al. Machine learning-aided engineering of hydrolases for PET depolymerization. Nature. . 2022;604:662–667. doi: 10.1038/s41586-022-04599-z. [DOI] [PubMed] [Google Scholar]
- 65.Xia Y, Xia CQ, Pan X, Shen HB. GraphBind: protein structural context embedded rules learned by hierarchical graph neural networks for recognizing nucleic-acid-binding residues. Nucleic Acids Res. . 2021;49:e51. doi: 10.1093/nar/gkab044. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 66.Vaswani A, Shazeer N, Parmar N, Uszkoreit J, Jones L, Gomez AN, Kaiser L, et al. Attention is all you Need. Proc Adv Neural Inf Process Syst. 2017, 30: 5998–6008
- 67.Aloysius N, Geetha M, Nedungadi P. Incorporating relative position information in transformer-based sign language recognition and translation. IEEE Access. . 2021;9:145929–145942. doi: 10.1109/ACCESS.2021.3122921. [DOI] [Google Scholar]
- 68.Meier J, Rao R, Verkuil R, Verkuil R, Liu J, Sercu T, Rives A. Language models enable zero-shot prediction of the effects of mutations on protein function. Proc Adv Neural Inf Process Syst. 2021, 34: 29287–29303
- 69.Ferruz N, Hoecker B. Controllable protein design with language models. Nat Mach Intell. 2022: 1–12
- 70.Rao RM, Liu J, Verkuil R, Meier J, Canny J, Abbeel P, Sercu T. Alexander Rives Proceedings of the 38th International Conference on Machine Learning, 2021, PMLR. 139: 8844–8856
- 71.Gulli A, Pal S. Deep learning with Keras. Packt Publishing Limited, Birmingham, 2017.
- 72.Abadi M, Barham P, Chen J, Chen Z, Davis A, Dean J, Devin M, et al. Tensorflow: a system for large-scale machine learning. Proc OSDI. 2016, 16: 265–283
- 73.Paszke A, Gross S, Massa F, Lerer A, Bradbury J, Chanan G, Killeen T, et al. PyTorch: an imperative style, high-performance deep learning library. Proc Adv Neural Inf Process Syst. 2019, 32: 8024–8035
- 74.Pedregosa F, Varoquaux G, Gramfort A, Michel V, Thirion B, Grisel O, Blondel M, et al. Scikit-learn: machine learning in Python. J Mach Learn Res. 2011, 12: 2825–2830.
- 75.Goodfellow I, Pouget-Abadie J, Mirza M, Xu B, Warde-Farley D, Ozair S, Courville A, et al. Generative adversarial nets. Proc Adv Neural Inf Process Syst. 2014, 2: 2672–2680.
- 76.Ronneberger O, Fischer P, Brox T. U-net: convolutional networks for biomedical image segmentation. Proc Int Conf Med Image Comput Comput-Assisted Intervention. 2015, 234–241
- 77.Jia Y, Shelhamer E, Donahue J, Karayev S, Long J, Girshick R, Guadarrama S, et al. Caffe: convolutional architecture for fast feature embedding. Proc 22nd ACM International Conference on Multimedia. 2014, 675–678
- 78.Collobert R, Bengio S, Mariéthoz J. Torch: a modular machine learning software library. Technical Report. 02-46, 2002.
- 79.Dimitriadou E, Hornik K, Leisch F, Meyer D, Weingessel A. Misc functions of the Department of Statistics (e1071), TU Wien. R package R package. , 2008, 1: 5–24
- 80.Kuhn M. Building predictive models in R using the caret Package . J Stat Soft. . 2008;28:1–26. doi: 10.18637/jss.v028.i05. [DOI] [Google Scholar]
- 81.Hall M, Frank E, Holmes G, Pfahringer B, Reutemann P, Witten IH. The WEKA data mining software. SIGKDD Explor Newsl. . 2009;11:10–18. doi: 10.1145/1656274.1656278. [DOI] [Google Scholar]
- 82.Abeel T, Van de Peer Y, Saeys Y. A machine learning library. J Mach Learn Res. , 2009, 10: 931–934
- 83.Engqvist MKM. Correlating enzyme annotations with a large set of microbial growth temperatures reveals metabolic adaptations to growth at diverse temperatures. BMC Microbiol. . 2018;18:1–4. doi: 10.1186/s12866-018-1320-7. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 84.Chang A, Jeske L, Ulbrich S, Hofmann J, Koblitz J, Schomburg I, Neumann-Schaal M, et al. BRENDA, the ELIXIR core data resource in 2021: new developments and updates. Nucleic Acids Res. . 2021;49:D498–D508. doi: 10.1093/nar/gkaa1025. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 85.Xavier JS, Nguyen TB, Karmarkar M, Portelli S, Rezende PM, Velloso JPL, Ascher DB, et al. ThermoMutDB: a thermodynamic database for missense mutations. Nucleic Acids Res. . 2021;49:D475–D479. doi: 10.1093/nar/gkaa925. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 86.Nikam R, Kulandaisamy A, Harini K, Sharma D, Gromiha MM. ProThermDB: thermodynamic database for proteins and mutants revisited after 15 years. Nucleic Acids Res. . 2021;49:D420–D424. doi: 10.1093/nar/gkaa1035. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 87.Stourac J, Dubrava J, Musil M, Horackova J, Damborsky J, Mazurenko S, Bednar D. FireProtDB: database of manually curated protein stability data. Nucleic Acids Res. . 2021;49:D319–D324. doi: 10.1093/nar/gkaa981. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 88.Pucci F, Bourgeas R, Rooman M. High-quality thermodynamic data on the stability changes of proteins upon single-site mutations. J Phys Chem Reference Data. . 2016;45:023104. doi: 10.1063/1.4947493. [DOI] [Google Scholar]
- 89.Madhavan A, Arun KB, Binod P, Sirohi R, Tarafdar A, Reshmy R, Kumar Awasthi M, et al. Design of novel enzyme biocatalysts for industrial bioprocess: harnessing the power of protein engineering, high throughput screening and synthetic biology. Bioresource Tech. . 2021;325:124617. doi: 10.1016/j.biortech.2020.124617. [DOI] [PubMed] [Google Scholar]
- 90.Frappier V, Keating AE. Data-driven computational protein design. Curr Opin Struct Biol. . 2021;69:63–69. doi: 10.1016/j.sbi.2021.03.009. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 91.Vanella R, Kovacevic G, Doffini V, Fernández de Santaella J, Nash MA. High-throughput screening, next generation sequencing and machine learning: advanced methods in enzyme engineering. Chem Commun. . 2022;58:2455–2467. doi: 10.1039/d1cc04635g. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 92.Ku T, Lu P, Chan C, Wang T, Lai S, Lyu P, Hsiao N. Predicting melting temperature directly from protein sequences. Comput Biol Chem. . 2009;33:445–450. doi: 10.1016/j.compbiolchem.2009.10.002. [DOI] [PubMed] [Google Scholar]
- 93.Gado JE, Beckham GT, Payne CM. Improving enzyme optimum temperature prediction with resampling strategies and ensemble learning. J Chem Inf Model. . 2020;60:4098–4107. doi: 10.1021/acs.jcim.0c00489. [DOI] [PubMed] [Google Scholar]
- 94.Li G, Buric F, Zrimec J, Viknander S, Nielsen J, Zelezniak A, Engqvist MKM. Learning deep representations of enzyme thermal adaptation. Protein Sci. . 2022;31: e4480 doi: 10.1002/pro.4480. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 95.Meng C, Ju Y, Shi H. TMPpred: a support vector machine-based thermophilic protein identifier. Anal Biochem. . 2022;645:114625. doi: 10.1016/j.ab.2022.114625. [DOI] [PubMed] [Google Scholar]
- 96.Foroozandeh Shahraki M, Ariaeenejad S, Fallah Atanaki F, Zolfaghari B, Koshiba T, Kavousi K, Salekdeh GH. MCIC: automated identification of cellulases from metagenomic data and characterization based on temperature and pH dependence. Front Microbiol. . 2020;11:567863. doi: 10.3389/fmicb.2020.567863. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 97.Zhang Y, Guan F, Xu G, Liu X, Zhang Y, Sun J, Yao B, et al. A novel thermophilic chitinase directly mined from the marine metagenome using the deep learning tool Preoptem. Bioresour Bioprocess. , 2022, 9: https://doi.org/10.1186/s40643-022-00543-1 . [DOI] [PMC free article] [PubMed]
- 98.Cui Y, Sun J, Wu B. Computational enzyme redesign: large jumps in function. Trends Chem. , 2022, 4: 409–419
- 99.Giollo M, Martin AJ, Walsh I, Ferrari C, Tosatto SC. NeEMO: a method using residue interaction networks to improve prediction of protein stability upon mutation. BMC Genomics. . 2014;15:1. doi: 10.1186/1471-2164-15-S4-S7. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 100.Laimer J, Hiebl-Flach J, Lengauer D, Lackner P. MAESTROweb: a web server for structure-based protein stability prediction. Bioinformatics. . 2016;32:1414–1416. doi: 10.1093/bioinformatics/btv769. [DOI] [PubMed] [Google Scholar]
- 101.Pucci F, Bourgeas R, Rooman M. Predicting protein thermal stability changes upon point mutations using statistical potentials: introducing HoTMuSiC. Sci Rep. . 2016;6:1–9. doi: 10.1038/srep23257. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 102.Yang Y, Urolagin S, Niroula A, Ding X, Shen B, Vihinen M. PON-tstab: protein variant stability predictor. Importance of training data quality. Int J Mol Sci. . 2018;19:1009. doi: 10.3390/ijms19041009. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 103.Kumar MDS. ProTherm and ProNIT: thermodynamic databases for proteins and protein-nucleic acid interactions. Nucleic Acids Res. . 2006;34:D204–D206. doi: 10.1093/nar/gkj103. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 104.Cao H, Wang J, He L, Qi Y, Zhang JZ. DeepDDG: predicting the stability change of protein point mutations using neural networks. J Chem Inf Model. . 2019;59:1508–1514. doi: 10.1021/acs.jcim.8b00697. [DOI] [PubMed] [Google Scholar]
- 105.Chen CW, Lin MH, Liao CC, Chang HP, Chu YW. iStable 2.0: predicting protein thermal stability changes by integrating various characteristic modules. Comput Struct Biotechnol J. . 2020;18:622–630. doi: 10.1016/j.csbj.2020.02.021. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 106.Kulandaisamy A, Zaucha J, Frishman D, Gromiha MM. MPTherm-pred: analysis and prediction of thermal stability changes upon mutations in transmembrane proteins. J Mol Biol. . 2021;433:166646. doi: 10.1016/j.jmb.2020.09.005. [DOI] [PubMed] [Google Scholar]
- 107.Li G, Panday SK, Alexov E. SAAFEC-SEQ: a sequence-based method for predicting the effect of single point mutations on protein thermodynamic stability. Int J Mol Sci. . 2021;22:606. doi: 10.3390/ijms22020606. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 108.Iqbal S, Li F, Akutsu T, Ascher DB, Webb GI, Song J. Assessing the performance of computational predictors for estimating protein stability changes upon missense mutations. Briefings BioInf. . 2021;22:bbab184. doi: 10.1093/bib/bbab184. [DOI] [PubMed] [Google Scholar]
- 109.Fang J. A critical review of five machine learning-based algorithms for predicting protein stability changes upon mutation. Briefings BioInf. . 2020;21:1285–1292. doi: 10.1093/bib/bbz071. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 110.Pussi F, Schwersensky M, Rooman M. AI challenges for predicting the impact of mutations on protein stability. arXiv. , DOI: arxiv-2111.04208 . 2111.04208, 2021 [DOI] [PubMed]
- 111.Usmanova DR, Bogatyreva NS, Bernad JA, Eremina AA, Gorshkova AA, Kanevskiy GM, Lonishin LR, et al. Self-consistency test reveals systematic bias in programs for prediction change of stability upon mutation. Bioinformatics. 2018, 34: 3653–3658 . [DOI] [PMC free article] [PubMed]
- 112.Romero PA, Krause A, Arnold FH. Navigating the protein fitness landscape with Gaussian processes. Proc Natl Acad Sci USA. . 2013;110:E193–E201. doi: 10.1073/pnas.1215251110. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 113.Bednar D, Beerens K, Sebestova E, Bendl J, Khare S, Chaloupkova R, Prokop Z, et al. FireProt: energy-and evolution-based computational design of thermostable multiple-point mutants. PLoS Comput Biol. . 2015;11:e1004556. doi: 10.1371/journal.pcbi.1004556. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 114.Wijma HJ, Floor RJ, Jekel PA, Baker D, Marrink SJ, Janssen DB. Computationally designed libraries for rapid enzyme stabilization. Protein Eng Des Sel. . 2014;27:49–58. doi: 10.1093/protein/gzt061. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 115.Cui Y, Chen Y, Liu X, Dong S, Tian Y, Qiao Y, Mitra R, et al. Computational redesign of a PETase for plastic biodegradation under ambient condition by the GRAPE strategy. ACS Catal. . 2021;11:1340–1350. doi: 10.1021/acscatal.0c05126. [DOI] [Google Scholar]
- 116.Barber-Zucker S, Mindel V, Garcia-Ruiz E, Weinstein JJ, Alcalde M, Fleishman SJ. Stable and functionally diverse versatile peroxidases designed directly from sequences. J Am Chem Soc. . 2022;144:3564–3571. doi: 10.1021/jacs.1c12433. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 117.Yang J, Anishchenko I, Park H, Peng Z, Ovchinnikov S, Baker D. Improved protein structure prediction using predicted interresidue orientations. Proc Natl Acad Sci USA. . 2020;117:1496–1503. doi: 10.1073/pnas.1914677117. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 118.Hiranuma N, Park H, Baek M, Anishchenko I, Dauparas J, Baker D. Improved protein structure refinement guided by deep learning based accuracy estimation. Nat Commun. . 2021;12:1. doi: 10.1038/s41467-021-21511-x. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 119.Goldenzweig A, Goldsmith M, Hill SE, Gertman O, Laurino P, Ashani Y, Dym O, et al. Automated structure-and sequence-based design of proteins for high bacterial expression and stability. Mol Cell. . 2016;63:337–346. doi: 10.1016/j.molcel.2016.06.012. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 120.Pinney MM, Mokhtari DA, Akiva E, Yabukarski F, Sanchez DM, Liang R, Doukov T, et al. Parallel molecular mechanisms for enzyme temperature adaptation. Science. . 2021;371:eaay2784. doi: 10.1126/science.aay2784. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 121.Singer JM, Novotney S, Strickland D, Haddox HK, Leiby N, Rocklin GJ, Chow CM, et al. Large-scale design and refinement of stable proteins using sequence-only models. PLoS One. . 2022;17:e0265020. doi: 10.1371/journal.pone.0265020. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 122.Wolpert DH, Macready WG. No free lunch theorems for optimization. IEEE Trans Evol Computat. . 1997;1:67–82. doi: 10.1109/4235.585893. [DOI] [Google Scholar]
- 123.Deng C, Ji X, Rainey C, Zhang J, Lu W. Integrating machine learning with human knowledge. iScience. . 2020;23:101656. doi: 10.1016/j.isci.2020.101656. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 124.Wu H, Chen Q, Zhang W, Mu W. Overview of strategies for developing high thermostability industrial enzymes: discovery, mechanism, modification and challenges. Crit Rev Food Sci Nutr. . 2021:1–18. doi: 10.1080/10408398.2021.1970508. [DOI] [PubMed] [Google Scholar]
- 125.Hait S, Mallik S, Basu S, Kundu S. Finding the generalized molecular principles of protein thermal stability. Proteins. . 2020;88:788–808. doi: 10.1002/prot.25866. [DOI] [PubMed] [Google Scholar]
- 126.Almeida VM, Marana SR. Optimum temperature may be a misleading parameter in enzyme characterization and application. PLoS ONE. . 2019;14:e0212977. doi: 10.1371/journal.pone.0212977. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 127.Bejani MM, Ghatee M. A systematic review on overfitting control in shallow and deep neural networks. Artif Intell Rev. . 2021;54:6391–6438. doi: 10.1007/s10462-021-09975-1. [DOI] [Google Scholar]
- 128.Rives A, Meier J, Sercu T, Goyal S, Lin Z, Liu J, Guo D, et al. Biological structure and function emerge from scaling unsupervised learning to 250 million protein sequences. Proc Natl Acad Sci USA. . 2021;118:e2016239118. doi: 10.1073/pnas.2016239118. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 129.Høie MH, Kiehl EN, Petersen B, Nielsen M, Winther O, Nielsen H, Hallgren J, et al. NetSurfP-3.0: accurate and fast prediction of protein structural features by protein language models and deep learning. Nucleic Acids Res. . 2022;50:W510–W515. doi: 10.1093/nar/gkac439. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 130.Ferruz N, Schmidt S, Höcker B. ProtGPT2 is a deep unsupervised language model for protein design. Nat Commun. . 2022;13:4348. doi: 10.1038/s41467-022-32007-7. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 131.Devlin J, Chang M, Lee K, Toutanova K. Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv. , DOI: arxiv-1810.04805 . 1810.04805, 2018
- 132.Jiménez-Luna J, Grisoni F, Schneider G. Drug discovery with explainable artificial intelligence. Nat Mach Intell. . 2020;2:573–584. doi: 10.1038/s42256-020-00236-4. [DOI] [Google Scholar]
- 133.Kryshtafovych A, Schwede T, Topf M, Fidelis K, Moult J. Critical assessment of methods of protein structure prediction (CASP)—Round XIII. Proteins. . 2019;87:1011–1020. doi: 10.1002/prot.25823. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 134.Zhou N, Jiang Y, Bergquist TR, Lee AJ, Kacsoh BZ, Crocker AW, Lewis KA, et al. The CAFA challenge reports improved protein function prediction and new functional annotations for hundreds of genes through experimental screens. Genome Biol. . 2019;20:1–23. doi: 10.1186/s13059-019-1835-8. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 135.Tian P, Best RB. Exploring the sequence fitness landscape of a bridge between protein folds. PLoS Comput Biol. . 2020;16:e1008285. doi: 10.1371/journal.pcbi.1008285. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 136.Ding X, Zou Z, Brooks Charles L. I. Deciphering protein evolution and fitness landscapes with latent space models. Nat Commun. . 2019;10:5644. doi: 10.1038/s41467-019-13633-0. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 137.Jones DT, Thornton JM. The impact of AlphaFold2 one year on. Nat Methods. . 2022;19:15–20. doi: 10.1038/s41592-021-01365-3. [DOI] [PubMed] [Google Scholar]
- 138.Thornton JM, Laskowski RA, Borkakoti N. AlphaFold heralds a data-driven revolution in biology and medicine. Nat Med. . 2021;27:1666–1669. doi: 10.1038/s41591-021-01533-0. [DOI] [PubMed] [Google Scholar]
- 139.Varadi M, Anyango S, Deshpande M, Nair S, Natassia C, Yordanova G, Yuan D, et al. AlphaFold Protein Structure Database: massively expanding the structural coverage of protein-sequence space with high-accuracy models. Nucleic Acids Res. . 2022;50:D439–D444. doi: 10.1093/nar/gkab1061. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 140.Pearce R, Zhang Y. Deep learning techniques have significantly impacted protein structure prediction and protein design. Curr Opin Struct Biol. . 2021;68:194–207. doi: 10.1016/j.sbi.2021.01.007. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 141.Sachidanandam R, Weissman D, Schmidt SC, Kakol JM, Stein LD, Marth G, Sherry S, et al. A map of human genome sequence variation containing 1.42 million single nucleotide polymorphisms. Nature. . 2001;409:928–933. doi: 10.1038/35057149. [DOI] [PubMed] [Google Scholar]
- 142.Venter JC, Adams MD, Myers EW, Li PW, Mural RJ, Sutton GG, Smith HO, et al. The sequence of the human genome. Science. . 2001;291:1304–1351. doi: 10.1126/science.1058040. [DOI] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.




