ABSTRACT
Protein sequence design is a highly challenging task, aimed at discovering new proteins that are more functional and producible under laboratory conditions than their natural counterparts. Deep learning‐based approaches developed to address this problem have achieved significant success. However, these approaches often do not adequately emphasize the functional properties of proteins. In this study, we developed a heuristic optimization method to enhance key functionalities such as solubility, flexibility, and stability, while preserving the structural integrity of proteins. This method aims to reduce laboratory demands by enabling a design that is both functional and structurally sound. This approach is particularly valuable for the synthetic production of proteins with anti‐inflammatory properties and those used in gene therapy. The designed proteins were initially evaluated for their ability to preserve natural structures using recovery and confidence metrics, followed by assessments with the AlphaFold tool. Additionally, natural protein sequences were mutated using a genetic algorithm and compared with those designed by our method. The results demonstrate that the protein sequences generated by our method exhibit much greater similarity to native protein sequences and structures. The code and sequences for the designed proteins are available at https://github.com/aysenursoyturk/HMHO.
Keywords: anti‐inflammatory proteins, deep learning, gene therapy, heuristic optimization, synthetic protein design
1. Introduction
Proteins, as the functional manifestations of genetic information, are the fundamental building blocks of cellular activities [1]. The structure and functionality of proteins are determined by the folding of amino acid sequences into specific three‐dimensional (3D) geometries [2]. This three‐dimensional conformation is the primary determinant of a protein's biological function. However, the laboratory processes required to determine these geometric shapes are labor‐intensive, require expertise, and are costly. To address these challenges, computational approaches have been developed [3, 4, 5]. The problem of predicting a protein's 3D structure from its amino acid sequence is a heavily studied topic in bioinformatics and structural biology. Initially, physically based approaches, such as the Rosetta method [6], were used for protein structure prediction. However, with advancements in artificial intelligence, deep learning algorithms have been developed to predict the 3D structures of proteins [5, 7].
Another significant challenge related to proteins is the efficient and effective design of protein sequences that can fold into desired 3D structures [8, 9, 10]. In the prediction of protein sequences from 3D structures, the amino acid sequence at different positions can form contact points between single or multiple chains, resulting in a wide range of potential sequence designs. Calculating these possibilities requires intensive computational processes. Despite advancements in computational capabilities for sequence design from 3D structures, the problem known as “inverse folding” or “protein design,” which involves designing a sequence from scratch for a desired 3D structure, remains a challenging task [11, 12].
Protein design is a vast and continually evolving field that aims to create or modify proteins to perform new or improved functions. These designed proteins, known as synthetic proteins, are molecules with specific structures and functions produced through protein engineering methods [13]. Synthetic proteins offer various advantages over natural proteins, particularly in drug development and industrial biotechnology, by aiding in the understanding of biological processes and optimizing targeted functions. While natural proteins are highly functional, they often present challenges for biotechnological applications due to their complexity, limitations for specific functions, poor expression in heterologous systems, limited solubility, and sensitivity to temperature [14].
1.1. Related Work
Protein design plays a crucial role in biotechnological applications [15, 16]. The functions of proteins are determined by their amino acid sequences and the 3D polymer structures formed by these sequences. One of the primary objectives of protein engineering is to design new proteins with specific functions. In this context, predicting amino acid sequences from the 3D structures of proteins has become a significant research focus in recent years [17, 18, 19].
Early studies in protein design relied on simple knowledge‐based and statistical methods [20]. Traditionally, the de novo design of proteins has been approached as an optimization problem, where the goal is to find the minima of an energy function [21]. This energy function is designed to evaluate a vast number of sequences for a given subunit, but this leads to increased computational complexity due to the multitude of possibilities. These studies aimed to demonstrate that specific structural motifs and secondary structure elements are associated with particular amino acid sequences [22]. Statistical approaches have employed sampling techniques such as Markov Chain Monte Carlo (MCMC) to create optimized sequences based on force fields or statistical potentials [23, 24, 25]. However, the limitations of sampling only a small portion of the vast search space have driven researchers to develop computational methods that yield more accurate results.
In recent years, machine learning and deep learning approaches have become widely applied in protein studies. Notably, algorithms developed for predicting protein structures from amino acid sequences have achieved remarkable success, significantly reducing the need for wet lab experiments. For protein design, methods based on deep neural networks and graph neural networks (GNNs) have also gained traction [26]. One such method is convolutional neural networks (CNNs), which have been highly successful in image recognition and processing and have begun to be effectively used for sequence generation in protein design. Tools like ProDCoNN, Anand, and DenseCPD utilize CNN architecture to predict amino acid sequences from protein 3D structures, particularly those based on x‐ray data [27, 28, 29]. However, while CNNs deliver excellent results in image processing, they struggle to capture the complex 3D structures and long‐range interactions inherent to proteins. Another approach used in sequence design is natural language processing (NLP). Tools like Ingraham, ABACUS‐R, and ProDESIGN‐LE employ encoder and transformer‐based decoders from NLP models to predict sequences from protein structures [30, 31]. Among these methods, the most successful results come from GNNs, specifically the message passing neural networks (MPNNs). MPNNs are a popular approach in chemistry for predicting the properties of molecules [26, 32]. In the MPNN architecture, a feature vector created by aggregating the states of a node and all its neighboring nodes along with potential edge functions is used as input to the model. Since all nodes in the data points representing a protein are interconnected, current machine learning methods can be designed by evaluating all connections within a vast search space. Although these algorithms can predict amino acid sequences capable of achieving the desired folds, they often overlook the functional properties of proteins by focusing solely on structure. This can result in the designed protein failing to be successfully synthesized in a laboratory setting. For instance, the laboratory applications of many protein‐based therapeutics and catalysts are limited by low stability. Additionally, protein solubility is crucial in laboratory applications as it ensures easier and more efficient purification processes [33]. Therefore, in protein design, it is essential to improve biophysical properties while maintaining the 3D structure.
1.2. Proteins Targeted for Design
Synthetic proteins are extensively used in medicine, agriculture, and industrial applications. Notably, in 2021, four of the top five best‐selling drugs were composed of synthetic proteins in their active forms [34]. Another example is the effort by protein engineers to enhance the thermostability (resistance to high temperatures) of a malaria invasion protein for use as a vaccine antigen [35]. Additionally, there is significant interest in producing synthetic protein forms aimed at increasing the activity of an enzyme capable of hydrolyzing (decomposing) polyethylene terephthalate (PET) [36]. Given the critical roles of synthetic proteins, those with anti‐inflammatory effects and those used in gene therapy have been selected for synthetic sequence production in the proposed method.
Gene therapy is rapidly advancing as a method for treating genetic disorders [37]. However, natural protein forms often lack sufficient quantity, stability, or functionality for therapeutic use. This situation necessitates the production of synthetic proteins for gene therapy applications [38]. To address this, more functional forms of five proteins used in gene therapy were selected for redesign.
The first protein is methylmalonyl‐CoA mutase (MUT), which is associated with elevated levels of methylmalonic acid in the body due to enzyme malfunction. Gene therapy involves introducing a healthy copy of the gene into diseased cells. The 3D structure of the protein, resulting from the translation of the transferred gene, is crucial for its function. Therefore, it is essential to ensure that the synthetic sequence produced has the desired structure [39, 40]. Another protein, Ornithine carbamoyltransferase (OTC), is involved in the urea cycle. This enzyme combines carbamoyl phosphate and ornithine to form citrulline and phosphate, and OTC deficiency is a genetic disorder [41]. Ongoing gene therapy studies on this enzyme highlight the need for designing synthetic forms for applicable laboratory research [42, 43]. Gaucher disease is associated with a deficiency in the Lysosomal acid glucosylceramidase (GCase) enzyme. Gene therapy aims to correct the defective GBA gene (encoding glucocerebrosidase) or provide a healthy copy to treat the disease or alleviate its symptoms [44, 45]. Fabry disease is related to the Alpha‐galactosidase A (GLA) enzyme. Producing synthetic sequences of GLA could offer a potential gene therapy approach for treating this disease. Synthetic designs can help create more stable and functional forms of GLA [46, 47]. Finally, the enzyme designed using the proposed method is Phenylalanine ammonia‐lyase (PAL), found in plants, fungi, and some bacteria. The PAL enzyme may serve as a potential therapeutic agent for treating PKU metabolic disease [48]. In addition to these complex enzyme structures, two proteins with relatively short amino acid sequences and anti‐inflammatory effects were also redesigned using the proposed method. The first is interleukin‐10 (IL‐10), an important cytokine that plays a role in regulating the immune system. IL‐10 is being investigated as a potential therapeutic agent for various inflammatory conditions, including inflammatory bowel disease (IBD), rheumatoid arthritis, and other autoimmune diseases [49, 50]. The α1‐Antitrypsin (AAT) protein, also known as Proteinase A Inhibitor 1, is released into the bloodstream and provides an important anti‐protease defense mechanism in the human body. Recombinant forms of this protein, produced from plasma, can be used for therapeutic purposes. Thus, synthetically produced, more stable, and soluble forms are more suitable for therapeutic applications [51].
We propose Heuristic Metropolis–Hastings Optimization (HMHO) for designing synthetic conformations that are crucial for therapeutic applications. The HMHO method overcomes the limitations of the restricted search space in MCMC sampling by exploring a subspace of protein space that is conducive to folding into functional structures. Its goal is to enhance the physical properties of proteins within this subspace while maintaining their structural integrity. Specifically, the HMHO method seeks to optimize the biophysical properties of proteins through MCMC searches within a subspace sampled by a protein model with high probability mass, thus achieving comprehensive optimization. The operational architecture of the HMHO method is depicted in Figure 1.
FIGURE 1.

Working architecture of the HMHO method. (A) Initially, a design is generated using a graph neural network model as part of the HMHO method. (B) From the resulting protein sequence, the three‐dimensional structure is redesigned by fixing all positions except for one randomly selected position. This redesigned structure is then fed back into the model. The biophysical properties of the redesigned sequence are compared with those of the initially designed sequence. The better solution is incorporated into the model for the next optimization step, and this process is repeated for a predetermined number of iterations.
2. Materials and Methods
2.1. Problem Definition
The protein sequence design problem involves searching for a sequence with specific properties in a multidimensional search space. However, in this study, the sequence search is not conducted over the entire space denoted by SA, but rather within a subspace suitable for folding into functional structures, denoted as SA′ = {R | For any subsequence of length L in R, a subspace sampled with high probability mass by a protein model.}. In this space, R represents the sequence of amino acids, while L represents the desired length of the sequence. If the objective function is defined as f, this function is a black‐box function that calculates the functional properties of the relevant protein. The goal is to maximize the function f: → V by modifying the initial sequence (s0).
2.2. Targeted Sequence Optimization in Protein Design: Navigating Functional Subspaces
Biological systems rely on the 3D structure of proteins for essential functions such as catalyzing chemical reactions, signal transduction and recognition, providing cellular structure and support, and executing desired biological processes. Therefore, the primary objective in protein design is to generate alternative sequences that preserve this structure. However, challenges arise in laboratory environments due to the low solubility of designed sequences, their instability under changing conditions (e.g., temperature), or a lack of sufficient flexibility. These issues can hinder the production and application of designed proteins [52]. To address these challenges, it is recommended to optimize not only the structure but also the functional properties of the proteins during the design process.
In this context, the target function to be optimized, denoted as f, is defined not as a single function but as a set of functions aimed at optimizing multiple parameters through multiobjective optimization methods. One of the key functions to be optimized is solubility. Protein solubility is crucial for maintaining the functional and biological activities of proteins, facilitating laboratory applications, supporting drug development processes, and enhancing their usability in industrial applications. The solubility value is calculated using the NetSolP tool, which predicts the solubility of amino acid sequences with higher accuracy and speed than other existing methods by utilizing artificial neural networks [33, 53].
The second function to be optimized is the instability index. Protein stability is critical for maintaining the structural integrity and functionality of proteins. For instance, the stability of proteins used as drug candidates is vital for preserving their therapeutic effects, while instability in laboratory studies can complicate manipulation, characterization, and analysis [33]. The instability index is a measure used to predict whether a protein will be stable under laboratory conditions [54]. An instability value is assigned to each of the 400 possible dipeptides, known as dipeptide instability weight values (DIWV), which have been experimentally validated. The instability index is computed using the following formula:
where L represents the length of the sequence; denotes the instability weight value for the dipeptide starting at position i.
The third function to be optimized is flexibility. Flexibility in proteins plays a significant role in various biological processes, including ligand binding, catalytic activity, structural adaptation, mobility, and drug design [55]. Flexibility is calculated using a formula proposed by Vihinen [56], which employs optimized parameters for a specific window size, determined to be 9 in previous studies. To calculate flexibility, subsequences within the specified window are analyzed, and a flexibility score is calculated for each subsequence based on the flexibility scores of each position within the subsequence. Different weights are applied to the front, back, and middle regions of the subsequence. The flexibility score for a subsequence is calculated using the following formula:
where N represents the length of the subsequence; Flexibility(i) denotes the flexibility score of the amino acid at position i within the subsequence; Weight(i), represents the weight used for position i; Flexibility(middle) denotes the flexibility score of the amino acid in the middle position; and Total weight is the sum of all weights.
The overall flexibility value for the protein is calculated by averaging the flexibility scores from the returned list values across the entire amino acid sequence.
2.3. HMHO
The Metropolis–Hastings algorithm is a type of MCMC method used to generate random samples from complex probability distributions [57]. The algorithm performs a random walk in the parameter space, and at each step, it decides whether to accept the next move by considering both randomness and the likelihood that the next step is more probable. The protein sequence design optimization problem can be framed as a black‐box optimization problem, where the Metropolis–Hastings algorithm facilitates effective exploration of the search space. This approach allows for the generation of functionally optimized sequences while preserving the 3D structure of the protein.
The HMHO algorithm is employed to obtain optimized versions of a specific protein sequence by taking into account the functional properties of the protein during the optimization process. The process begins with the redesign of a randomly chosen position within a protein sequence using the ProteinMPNN algorithm [26]. The solubility value, flexibility value, and instability index of the redesigned sequence are then calculated. The acceptability of the designed sequence is determined based on whether it improves upon the current solution or meets the Metropolis criterion. If the sequence is accepted, the optimization continues by redesigning another randomly chosen residue, and this process is repeated for a specified number of iterations. The pseudocode for this multiobjective heuristic optimization algorithm is shown in Supporting Information. At the end of this process, protein sequences are obtained that have the desired functions while preserving the structure and meeting the specified parameters.
2.4. Evaluating ProteinMPNN and HMHO Methods: A Comparison of Optimized and Native Protein Sequences
The solubility, stability, and flexibility properties of sequences optimized using the HMHO method were compared with those generated solely by the ProteinMPNN tool. For this comparison, the ProteinMPNN tool was run five times with a sampling temperature value of 0.2. Subsequently, the tool was executed 70 times, and the most frequently occurring amino acid at each position among the 70 generated sequences was selected to form a consensus sequence. Biophysical calculations were then performed for the resulting consensus sequence, and its 3D structural prediction was carried out using the AlphaFold tool.
Finally, the similarity between the sequences designed using the HMHO method and natural sequences was determined through pairwise alignment. The alignment matrix used to measure sequence similarity was BLOSUM62, with a gap open penalty of 12 and a gap extension penalty of 4. Sequence similarity and the corresponding alignment score were then calculated.
2.5. Structural Validation of the Redesigned Sequences
These proteins can be specifically designed to enact targeted genetic changes or to activate the cell's natural repair mechanisms. For example, the redesigned MUT protein serves as a therapeutic agent in genetic disorders like methylmalonic acidemia (MMA). A healthy methylmalonyl‐CoA mutase enzyme features a specific catalytic pocket that converts methylmalonyl‐CoA molecules into succinyl‐CoA [39, 40]. However, mutations in this enzyme can alter or completely disable this catalytic pocket. Therefore, understanding the 3D structure of proteins used in gene therapy is vital for evaluating their ability to perform the intended tasks in the body.
It is crucial to verify whether the structures of these designed proteins are preserved. To assess this, calculations have been performed to demonstrate the preservation of the structure and to simultaneously measure the reliability of the designed sequence using the HMHO method. These calculations include:
2.5.1. Recovery
This is the primary indicator of the ability to restore the original state of the designed protein. A high recovery value indicates that the designed protein sequence is similar to the reference sequence, and therefore, it will likely resemble the reference structure when folded. This value serves as an intermediate measure to assess structural similarity [55].
where n represents the number of data points; and denote the predicted and true values, respectively; 1, is the indicator function; it checks whether is equal to returning 1 if they are equal and 0 otherwise.
2.5.2. Perplexity
This metric provides the entropy‐dependent distribution width of the predicted posterior amino acid distribution. A low perplexity value indicates that the model better predicts the amino acid probabilities in the test set.
where N the number of amino acids in the test set; P(W i ), the probability of the predicted amino acid by the model.
2.5.3. Confidence
This metric measures the average estimated probability of the designed amino acids, which are used to assess the quality of a designed protein sequence without requiring a reference sequence. The estimated probabilities for each position are summed, and the confidence metric is obtained by taking their average [58].
where n represents the total number of positions; represents the designed amino acid at position i; and denotes the estimated probability for
These calculations serve as intermediate measures of structural similarity, and the similarity between the designed protein structure and the natural protein structure was verified using AlphaFold. To mitigate biases from different algorithms, various approaches were tried. TM (template modeling score) and RMSD values were used as validation parameters. The similarity between the two protein structures was measured using the TM score (template modeling score) [59]. Additionally, another control parameter involves calculating the root mean square deviation (RMSD) from the distances of equivalent residues in both structures. To account for differences arising from the algorithm used for structure validation, sc‐TM calculations were performed. This calculation is as follows:
where f represents the protein folding algorithm; sc‐TM represents the metric used to measure protein structure similarity; represents the protein folding algorithm used to predict the structures of the designed sequences; and represents the protein folding algorithm used to predict the structures of the reference sequences.
2.6. Benchmarking the HMHO Method Against Genetic Algorithms for Multiobjective Protein Design
The proposed HMHO method enables the design of proteins with improved functional properties while preserving their structural integrity. To demonstrate the superiority of HMHO over other optimization techniques, it was compared against a genetic algorithm, a well‐established method in the literature for designing proteins with enhanced functionality. In this comparative analysis, the genetic algorithm was used to optimize flexibility, solubility, and stability metrics by introducing mutations into the natural sequence [60]. The parameters for the genetic algorithm's multiobjective optimization were set as follows: a population size of 100, 100 generations, a mutation rate of 0.1, and 5 tournaments selected via the tournament selection method. A weighted sum approach was applied to address the multiobjective optimization. The selected parameters represent the most commonly used options; however, to ensure the reliability of the designed sequences and achieve an optimal solution to the problem, hyperparameter optimization was conducted. Using the Optuna library, parameters such as population size, number of generations, mutation rate, and tournament selection count were optimized to obtain the best fitness value. The pseudocode of the algorithm used is presented in Supporting Information.
2.7. Comparative Analysis of HMHO‐Designed Sequences and Experimentally Validated Homologous Proteins
The experimental validation of sequences generated using the HMHO method is of great importance; however, appropriate in vitro conditions must be established for this validation. To demonstrate that the designed sequences are indeed functional, it is also considered crucial to assess their similarity to experimentally verified sequences. For this purpose, homologous sequences of natural protein sequences were retrieved using a PSI‐BLAST search. The database selected for this search was the Swiss‐Prot database [61], which contains proteins curated through manual review by experts. These experts verify protein sequences and associated information, such as function, structural properties, and pathological relevance, by consulting experimental data, published literature, and reliable sources. This makes the Swiss‐Prot database relatively small, but highly reliable, with minimal annotation errors.
Subsequently, the distances between the homologous sequences and the sequences designed using the HMHO method, as well as the consensus sequence generated by running the ProteinMPNN tool 70 times (where the most frequently occurring amino acid at each position was selected), were calculated to construct a distance matrix. The distance matrix was computed using the—distmat‐out parameter of Clustal Omega.
3. Results
3.1. Optimization Results
First, IL‐10 and AAT, proteins with anti‐inflammatory effects that are targeted for synthetic development, were designed using the HMHO method. Figure 2 illustrates the optimization process of the HMHO method, showing the changes in the optimized parameters—solubility, stability, and flexibility—during the procedure.
FIGURE 2.

Optimization of biophysical properties of proteins with anti‐inflammatory effects.
Next, the HMHO method was used to design OTC, MUT, GCase, GLA, and PAL—proteins utilized in gene therapy. Figure 3 displays the optimization process of the HMHO method, illustrating the changes in the optimized parameters—solubility, stability, and flexibility—during the procedure.
FIGURE 3.

Optimization of biophysical properties of proteins used in gene therapy.
3.2. Comparative Analysis of Native Protein Sequences, ProteinMPNN‐Designed Sequences, and HMHO‐Designed Sequences
The predicted improvements in proteins designed using the HMHO method compared to the initial representatives are trivially evident from the optimization results; however, comparing them with sequences generated by ProteinMPNN is needed to evaluate the algorithm's performance. To this end, five distinct sequences were generated for the same protein using ProteinMPNN under identical priority conditions, and their functional outcomes were calculated. Additionally, the ProteinMPNN tool was executed 70 times to produce a consensus sequence based on the most frequent amino acids at each position. The functional properties of these sequences and pre‐ and post‐optimization measurements of the parameters optimized with the HMHO method are shown in Table 1. The 3D structure of this consensus sequence was generated using the AlphaFold tool, and the sc‐TM value representing its similarity to the native structure and comparative images of the 3D structures of the native sequences are presented in Figure 4. As shown in Table 1, the sequences generated using the HMHO method exhibit superior biophysical properties compared to both the native proteins and the sequences generated by ProteinMPNN. Furthermore, as illustrated in Figure 4, the consensus sequences fail to preserve the 3D structure of the proteins.
TABLE 1.
Results of functional properties of natural proteins, designed by HMHO and designed by ProteinMPNN sequences, consensus sequences.
| Natural proteins | Synthetic proteins designed by the HMHO method | Consensus | 1st sequence designed with ProteinMPNN | 2st sequence designed with ProteinMPNN | 3st sequence designed with ProteinMPNN | 4st sequence designed with ProteinMPNN | 5st sequence designed with ProteinMPNN | ||
|---|---|---|---|---|---|---|---|---|---|
| OTC | Instability index | 36.44 | 5.04 | 30.35 | 32.93 | 30.98 | 35.13 | 31.07 | 34.17 |
| Solubility | 0.6121 | 0.7338 | 0.7372 | 0.7814 | 0.7110 | 0.6364 | 0.6763 | 0.6838 | |
| Flexibility | 1 | 1 | 0.98 | 0.99 | 1.00 | 1.00 | 0.99 | 1.00 | |
| IL‐10 | Instability index | 53.16 | 15.00 | 21.94 | 36.04 | 20.02 | 21.54 | 31.18 | 34.37 |
| Solubility | 0.4186 | 0.7073 | 0.6542 | 0.57462 | 0.6695 | 0.7851 | 0.4003 | 0.6965 | |
| Flexibility | 0.96 | 0.99 | 1.00 | 0.97 | 0.99 | 1.00 | 0.98 | 1.00 | |
| AAT | Instability index | 33.78 | 18.64 | 65.86 | 48.11 | 76.08 | 42.30 | 64.48 | 58.47 |
| Solubility | 0.7511 | 0.9680 | 0.9270 | 0.9182 | 0.7976 | 0.9023 | 0.8673 | 0.9012 | |
| Flexibility | 1 | 1 | 1.00 | 0.99 | 0.99 | 1.00 | 0.99 | 1.00 | |
| MUT | Instability index | 40.65 | 7.10 | 29.45 | 30.53 | 30.01 | 30.49 | 31.71 | 31.50 |
| Solubility | 0.5291 | 0.6284 | 0.6707 | 0.6228 | 0.6752 | 0.6885 | 0.5673 | 0.6115 | |
| Flexibility | 1 | 1 | 1.00 | 0.99 | 0.97 | 0.96 | 1.00 | 0.99 | |
| GCase | Instability index | 39.08 | 5.47 | 42.50 | 40.61 | 38.22 | 38.61 | 43.41 | 47.41 |
| Solubility | 0.2062 | 0.2766 | 0.4087 | 0.3366 | 0.2876 | 0.3387 | 0.3207 | 0.3173 | |
| Flexibility | 0.97 | 0.99 | 0.99 | 1.00 | 1.00 | 1.00 | 0.97 | 1.00 | |
| GLA | Instability index | 41.09 | 5.05 | 34.43 | 31.90 | 31.75 | 32.89 | 43.62 | 37.33 |
| Solubility | 0.1767 | 0.2801 | 0.3582 | 0.3088 | 0.3372 | 0.3233 | 0.3190 | 0.3389 | |
| Flexibility | 0.98 | 0.99 | 1.00 | 0.99 | 0.96 | 1.00 | 1.00 | 1.00 | |
| PAL | Instability index | 33.67 | 5.07 | 29.81 | 27.13 | 29.92 | 29.56 | 34.99 | 30.67 |
| Solubility | 0.5272 | 0.5961 | 0.5063 | 0.4831 | 0.5064 | 0.5138 | 0.5037 | 0.4554 | |
| Flexibility | 1 | 1 | 0.97 | 0.97 | 0.98 | 1.00 | 1.00 | 0.99 |
FIGURE 4.

The consensus sequence obtained by running ProteinMPNN multiple times is evaluated for structural similarity to natural sequences using AlphaFold predictions. The structural comparison is quantified using the sc‐TM score and RMSD value.
3.3. Evaluation of the Similarity of Synthetic Sequences Generated by the HMHO Method to Native Protein Sequences and the Confidence, Recovery, and Perplexity Values of the Proposed Sequences
To demonstrate the similarity of the designed sequences to native sequences, pairwise alignments were performed. As shown in Table 2, the synthetic sequences align with the native sequences with remarkably high scores, and their percentage similarity is also significantly high. Additionally, the confidence values, perplexity values, and recovery values of the generated sequences are notably high, serving as additional metrics that enhance the reliability of the designed sequences. These results are also presented in Table 2. The evaluation metrics in Table 3 provide the initial assessment of structural integrity. For a more accurate structural control, the AlphaFold tool, which is widely used in both in silico and in vitro studies, was employed. The 3D structures of proteins designed using the HMHO method were re‐predicted alongside the natural protein structures, and the sc‐TM and RMSD values are presented in Table 3. In addition to structural analysis, the ConSurf [62] tool was used to examine whether the ligand binding regions and active residues present in the natural sequences were altered in the protein sequences designed using the HMHO method. ConSurf is both a database and a server for identifying functional regions in a protein. For each protein, the ConSurf tool was applied, and the number of highly conserved atoms was displayed in Table 3. Subsequently, the predicted structures of the HMHO‐designed sequences and the natural sequences were superimposed and compared. Conserved regions were identified, and it was determined whether these regions were present in the new designs. The number of altered residues is also provided in Table 3. In Figure 5, the first panel illustrates the superimposition of the native structure with the optimized structure. The second panel highlights the surface binding sites of the relevant protein, with dark purple regions indicating the highly conserved areas identified using the ConSurf tool. The final panel presents the superimposition of the structure designed using the HMHO method with the conserved regions.
TABLE 2.
Results of pairwise alignments, along with recovery values, confidence scores, and perplexity values.
| Recovery | Perplexity | Confidence | Alignment identity | Alignment score | |
|---|---|---|---|---|---|
| IL‐10 | 0.894039736 | 1.004944272 | 0.995080054 | 33.1% | 124 residues overlap, Score = 162.0 |
| AAT | 0.909090909 | 5.846752280 | 0.710351638 | 84.4% | 77 residues overlap, Score = 332.0 |
| OTC | 0.841121495 | 5.394784725 | 0.685351106 | 83.8% | 321 residues overlap, Score = 1360.0 |
| MUT | 0.842253521 | 1.000000596 | 0.999999403 | 83.7% | 708 residues overlap, Score = 3011.0 |
| GCase | 0.835820895 | 7.464276030 | 0.613397145 | 83.0% | 535 residues overlap, Score = 2332.0 |
| GLA | 0.822843822 | 1.0 | 0.999999553 | 82.3% | 429 residues overlap, Score = 1943.0 |
| PAL | 0.859375000 | 1.0 | 0.999888342 | 85.8% | 704 residues overlap, Score = 3069.0 |
TABLE 3.
sc‐TM and RMSD values in the 3D structure comparison of sequences designed with the HMHO method and natural sequences.
| AlphaFold predict | Consurf results | |||
|---|---|---|---|---|
| Sc‐TM | RMSD | The number of highly conserved atoms | The number of atoms in changed residues | |
| IL‐10 | 0.9421 | 1.24 | 149 atoms | 0 atoms |
| AAT | 0.8817 | 0.84 | 172 atoms | 0 atoms |
| OTC | 0.9950 | 0.48 | 305 atoms | 0 atoms |
| MUT | 0.9919 | 0.82 | 840 atoms | 0 atoms |
| GCase | 0.9431 | 0.38 | 663 atoms | 0 atoms |
| GLA | 0.9627 | 0.51 | 487 atoms | 0 atoms |
| PAL | 0.9947 | 0.61 | 1061 atoms | 0 atoms |
FIGURE 5.

The first image presents the native protein alongside the HMHO‐designed variant. The second image illustrates the conserved regions identified by Consurf. Finally, the conserved regions and the AlphaFold prediction of the protein designed using the HMHO model are displayed together.
3.4. Results of Sequences Optimized With Genetic Algorithm
The genetic algorithm was initially executed using the most commonly used parameters, and the comparison of the biophysical properties of the optimized protein sequences with those of natural protein sequences is presented in Table 4. Subsequently, the genetic algorithm was run with parameters optimized using the Optuna library to determine the most accurate parameters for the solution. The properties of the sequences optimized under these conditions are also shown in Table 4. Additionally, the best fitness function parameters for each protein during the optimization process are provided in Table 4.
TABLE 4.
Functional properties of natural sequences and functional properties designed by genetic algorithms.
| Measurements of natural proteins | Measurements of synthetic proteins designed using the genetic algorithm with non‐optimized parameters | Measurements of synthetic proteins designed using the genetic algorithm with optimized parameters | Parameters optimized with Optuna | |||||||
|---|---|---|---|---|---|---|---|---|---|---|
| Solubility | Flexibility | Instability index | Solubility | Flexibility | Instability index | Solubility | Flexibility | Instability index | ||
| IL‐10 | 0.4186 | 0.96 | 53.16 | 0.9579 | 0.9984 | 12.31 | 0.8389 | 0.99 | 11.21 |
Population size: 69 Generations: 27 Mutation rate: 0.12 Tournament size: 5 |
| AAT | 0.7511 | 1 | 33.78 | 0.6815 | 1 | 11.89 | 0.9675 | 1.01 | 10.02 |
Population size: 78 Generations: 47 Mutation rate: 0.06 Tournament size: 5 |
| OTC | 0.6121 | 1 | 36.44 | 0.7708 | 0.99 | 12.53 | 0.9043 | 1 | 11.69 |
Population size: 77 Generations: 36 Mutation rate: 0.06 Tournament size: 5 |
| MUT | 0.5291 | 1 | 40.65 | 0.5299 | 0.98 | 16.00 | 0.6399 | 1 | 10.97 |
Population size: 58 Generations: 22 Mutation rate: 0.05 Tournament size: 4 |
| GCase | 0.2062 | 0.97 | 39.08 | 0.5796 | 0.9870 | 11.99 | 0.5208 | 0.99 | 10.43 |
Population size: 79 Generations: 29 Mutation rate: 0.05 Tournament size: 5 |
| GLA | 0.1767 | 0.98 | 41.09 | 0.6462 | 0.9914 | 14.32 | 0.6203 | 0.99 | 9.76 |
Population size: 78 Generations: 33 Mutation rate: 0.05 Tournament size: 5 |
| PAL | 0.5272 | 1 | 33.67 | 0.5742 | 0.9885 | 16.26 | 0.6536 | 1 | 17.92 |
Population size: 75 Generations: 33 Mutation rate: 0.1 Tournament size: 5 |
Figure 6 illustrates the 3D structure of the sequences designed using the genetic algorithm with the most commonly used parameters and the 3D structure of the sequences designed with the optimized parameters. As observed from the figure, while the biophysical properties of the sequences improved regardless of whether the parameters were optimized, the 3D structure could not be preserved.
FIGURE 6.

Predicted natural structures and design structures in the AlphaFold tool of proteins whose biophysical properties are increased by genetic algorithms.
3.5. HMHO Sequence Design and Experimental Validation
To evaluate the sequential and phylogenetic conservation of the proteins designed by the HMHO method, comparisons were made with their naturally active counterparts. Homologous sequences for each protein were identified using PSI‐BLAST. Both the consensus sequence (generated by selecting the most frequently occurring amino acids across multiple runs of ProteinMPNN) and the HMHO‐designed sequences were aligned with the homologous sequences. The average distances of the aligned sequences were calculated and are presented in Table 5. As shown in Table 5, the sequences generated by the HMHO method exhibit significantly shorter distances compared to the consensus sequence, indicating a much higher similarity to the homologous sequences.
TABLE 5.
The average distances of sequences designed using the HMHO method and the consensus sequence generated by ProteinMPNN to homologous sequences.
| Protein name | Number of homologous sequences identified | Average distance to the consensus sequence | Average distance of HMHO‐designed sequence |
|---|---|---|---|
| IL‐10 | 48 | 0.800118 | 0.753863 |
| AAT | 15 | 0.718783 | 0.498075 |
| OTC | 650 | 0.723043 | 0.630094 |
| MUT | 99 | 0.812521 | 0.642848 |
| GCase | 90 | 0.722365 | 0.639418 |
| GLA | 332 | 0.816331 | 0.720388 |
| PAL | 558 | 0.731023 | 0.641370 |
4. Discussion
The protein design problem is considered challenging because it involves a high‐dimensional search space, which is difficult to navigate. Studies have shown that deep learning methods can solve this problem feasibly and with high accuracy [63, 64]. Although these methods can predict different sequences while preserving the 3D structure by observing the spatial arrangement of amino acids, they often overlook the biophysical properties of the proteins. The developed HMHO method enhances the properties of proteins throughout their 3D structure.
To demonstrate the applicability of this method, proteins with anti‐inflammatory effects and proteins used in gene therapy were selected. These are proteins currently undergoing clinical studies for potential treatment applications [39, 40, 41, 43, 44, 45, 46, 47, 48, 49, 50, 51]. Additionally, the fact that these proteins contain multi‐chain structures and are complex is crucial for demonstrating the method's applicability to different proteins. Figures 2 and 3 show the improvement of the selected proteins' parameters during the optimization process. As can be seen in Figures 2 and 3, solubility and instability index values are the parameters that show the most significant improvement during the optimization process. With the HMHO method, the solubility value of the IL‐10 protein was increased from 0.4186 to 0.7073, and a stable design was obtained by reducing the instability index value from 53.16 in its natural sequence to 15.00, indicating that it is now stable in vitro [65]. This allows for a stable and more soluble form of IL‐10 to be produced in a laboratory environment. The native sequences of the GCase and GLA proteins have very low solubility, but a higher solubility value could be achieved using the HMHO method.
It can be seen that the flexibility value does not differ much between the natural protein parameters given in Table 1 and the parameters of synthetic designs optimized with the HMHO method. This suggests that flexibility might not be a parameter that can be easily optimized. However, flexibility is significant because it grants proteins many features such as functional compatibility, ligand binding and release, a balance of stability and dynamism, and resistance to mutations [42, 66]. Therefore, it is suggested that the optimization process should use a deep learning method that estimates flexibility based on the amino acid sequence or 3D structure rather than a formula that calculates flexibility from the amino acid sequence. However, this approach could increase the optimization time. Table 1 presents the biophysical properties of the designs generated by running ProteinMPNN 5 times, as well as the biophysical properties of the consensus sequence (comprising the most frequently occurring amino acids) obtained after 70 iterations. As evident from the table, the HMHO method demonstrates superior performance in enhancing the biophysical properties of the optimized protein sequences. While ProteinMPNN produces designs with relatively higher solubility, it falls short compared to the HMHO method in terms of stability.
The developed HMHO method achieves greater recovery values than deep learning methods developed so far [67]. The perplexity and confidence values in Table 2 show the reliability of the designed sequences. A low perplexity value and a high confidence value are desirable. The Monte Carlo search includes the applicable subspace of the protein space thanks to the extraction of sequences by the deep learning method. Since the proposed HMHO method incorporates the Monte Carlo search based on deep learning, the confidence value of the designed sequences is expected to be high. The confidence value ranges from 0.61 to 0.99, indicating that the designed sequences are quite reliable. As seen in Table 2, the GCase protein has the lowest confidence value and the highest perplexity value. However, when examining the natural structure of this protein, the Flexible Loops make its task difficult because these extensions are not arranged in space to estimate them. Figure 5 shows the 3D structure of the GCase protein, and this extended Flexible Loops domain is outside the compact structure of the protein.
The alignment results presented in Table 2 demonstrate the pairwise alignment of sequences designed using the HMHO method with native protein sequences. As observed in the table, the sequences designed with the HMHO method achieved substantially greater alignment scores against native sequences.
It is crucial to show that these proteins, which have been redesigned by improving their biophysical properties, preserve their 3D structure because a protein's function is closely related to its 3D structure [68]. Control parameters were also established using structures predicted by AlphaFold, a tool frequently used for 3D structure prediction from amino acid sequences. However, protein folding algorithms have an inductive bias, and to avoid prediction errors that might affect the evaluation, both the designed sequence and the natural sequence were predicted by the protein folding algorithm, and the TM value and RMSD value were calculated. As shown in Table 3, the proteins designed using the HMHO method achieved very low TM and RMSD values, indicating a high degree of structural similarity to native conformations. Figure 5 provides visual alignments of the 3D structures, further illustrating this similarity. To enhance the reliability of the HMHO method, the 3D structures of consensus sequences generated from ProteinMPNN designs were also predicted using the AlphaFold tool, and their similarities to native structures were evaluated based on TM and RMSD values.
As presented in Figure 4, while the consensus sequences of some proteins showed significant structural similarity to their native forms, the 3D structures of other proteins exhibited notable deviations by diverging from the desired conformations. In contrast, the HMHO method, evaluated using AlphaFold predictions, resulted in 3D folds closely resembling the native structures, supported by low TM and RMSD values.
Sequences designed with the HMHO method are a biophysically enhanced method that can design proteins with structures similar to natural structures, but it is believed that comparison with other optimization methods would further increase their reliability. For this purpose, the solubility, flexibility, and instability index parameters were optimized using a genetic algorithm by mutating the natural sequence. Table 4 shows the improved values resulting from optimization, and it is seen that the mutated sequences exhibit better biophysical properties compared to the natural sequences. However, these mutated sequences show structures that are quite different from the 3D structure of natural proteins. Although the sequences mutated by the genetic algorithm achieved very successful biophysical results, they altered the 3D structure of the protein. AlphaFold outputs and structural differences are seen in Table 4 and Figure 6. When Figure 6 is examined carefully, it can be observed that even simple protein structures are dramatically changed by random mutation, affecting the 3D structure of the protein. The HMHO method provides very successful results in preserving the structure while improving biophysical properties. Considering these results, it is believed that the HMHO method is more appropriate for synthetic protein design.
Additionally, the HMHO method can be developed to include the optimization of other biophysical properties such as thermostability, viscosity, water retention capacity, etc. However, since the HMHO method is based on Monte Carlo simulation, it requires extensive computations, and hence, the full optimization of a protein requires significant time. In future studies, functional computation components could be integrated into the deep learning framework to develop end‐to‐end machine learning models that do not require simulation. Furthermore, the methodology might be boosted with recycling approaches such as recalculating the 3D structure of the designed sequences and refitting the sequence based on the updated structure iteratively. Such strategies might find pareto‐optimal functional states in the available protein space by expectation maximization. However, this also requires a computational overhead due to iterative fold predictions.
Finally, it is essential to test the resulting designs under laboratory conditions [69]. Creating appropriate in vitro conditions and testing the functional properties of the optimized sequences require laboratory efforts. However, to enhance confidence that the generated sequences can be replicated in a laboratory environment, homologous sequences of natural proteins were retrieved from a database of proteins, and the distances to these sequences were calculated and averaged. Upon examining the distance values provided in Table 5, it is evident that the sequences designed using the HMHO method demonstrate significantly closer relationships to the homologous sequences compared to the consensus sequence generated by ProteinMPNN.
Author Contributions
Ayşenur Soytürk Patat: software, writing – review and editing, visualization. Özkan Ufuk Nalbantoğlu: formal analysis, writing – review and editing.
Supporting information
Data S1. Supporting Information.
Data Availability Statement
The data that support the findings of this study are available in aysenursoyturk at https://github.com/aysenursoyturk/HMHO. These data were derived from the following resources available in the public domain: UniProt Protein Database, https://www.uniprot.org; NCBI GenBank, https://www.ncbi.nlm.nih.gov/genbank/.
References
- 1. Morris R., Black K. A., and Stollar E. J., “Uncovering Protein Function: From Classification to Complexes,” Essays in Biochemistry 66, no. 3 (2022): 255–285, 10.1042/EBC20200108. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 2. Dorn M., e Silva M. B., Buriol L. S., and Lamb L. C., “Three‐Dimensional Protein Structure Prediction: Methods and Computational Strategies,” Computational Biology and Chemistry 53 (2014): 251–276, 10.1016/j.compbiolchem.2014.10.001. [DOI] [PubMed] [Google Scholar]
- 3. Chothia C. and Lesk A. M., “The Relation Between the Divergence of Sequence and Structure in Proteins,” EMBO Journal 5, no. 4 (1986): 823–826, 10.1002/j.1460-2075.1986.tb04288.x. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 4. Berman H. M., Westbrook J., Feng Z., et al., “The Protein Data Bank,” Nucleic Acids Research 28, no. 1 (2000): 235–242, 10.1093/nar/28.1.235. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 5. Jumper J. M., Evans R., Pritzel A., et al., “Highly Accurate Protein Structure Prediction With AlphaFold,” Nature 596 (2021): 583–589, https://api.semanticscholar.org/CorpusID:235959867. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 6. Leman J. K., Weitzner B. D., Lewis S. M., et al., “Macromolecular Modeling and Design in Rosetta: Recent Methods and Frameworks,” Nature Methods 17, no. 7 (2020): 665–680, 10.1038/s41592-020-0848-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 7. Villegas‐Morcillo A., Gomez A. M., and Sanchez V., “An Analysis of Protein Language Model Embeddings for Fold Prediction,” Briefings in Bioinformatics 23, no. 3 (2022): bbac142, 10.1093/bib/bbac142. [DOI] [PubMed] [Google Scholar]
- 8. Huang P. S., Boyken S. E., and Baker D., “The Coming of Age of De Novo Protein Design,” Nature 537, no. 7620 (2016): 320–327, 10.1038/nature19946. [DOI] [PubMed] [Google Scholar]
- 9. Alford R. F., Leaver‐Fay A., Jeliazkov J. R., et al., “The Rosetta All‐Atom Energy Function for Macromolecular Modeling and Design,” Journal of Chemical Theory and Computation 13, no. 6 (2017): 3031–3048, 10.1021/acs.jctc.7b00125. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 10. Kuhlman B. and Bradley P., “Advances in Protein Structure Prediction and Design,” Nature Reviews. Molecular Cell Biology 20, no. 11 (2019): 681–697, 10.1038/s41580-019-0163-x. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 11. Rocklin G. J., Chidyausiku T. M., Goreshnik I., et al., “Global Analysis of Protein Folding Using Massively Parallel Design, Synthesis, and Testing,” Science 357, no. 6347 (1979): 168–175, 10.1126/science.aan0693. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 12. Gainza P., Nisonoff H. M., and Donald B. R., “Algorithms for Protein Design,” Current Opinion in Structural Biology 39 (2016): 16–26, 10.1016/j.sbi.2016.03.006. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 13. Singh R. K., Lee J. K., Selvaraj C., et al., “Protein Engineering Approaches in the Post‐Genomic Era,” Current Protein & Peptide Science 19, no. 1 (2018): 5–15, 10.2174/1389203718666161117114243. [DOI] [PubMed] [Google Scholar]
- 14. Sumida K. H., Núñez‐Franco R., Kalvet I., et al., “Improving Protein Expression, Stability, and Function With ProteinMPNN,” Journal of the American Chemical Society 146, no. 3 (2024): 2054–2061, 10.1021/jacs.3c10941. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 15. Pan X. and Kortemme T., “Recent Advances in de Novo Protein Design: Principles, Methods, and Applications,” Journal of Biological Chemistry 296 (2021): 100558, 10.1016/j.jbc.2021.100558. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 16. Röthlisberger D., Khersonsky O., Wollacott A. M., et al., “Kemp Elimination Catalysts by Computational Enzyme Design,” Nature 453, no. 7192 (2008): 190–195, 10.1038/nature06879. [DOI] [PubMed] [Google Scholar]
- 17. Pakhrin S. C., Shrestha B., Adhikari B., and KC D. B., “Deep Learning‐Based Advances in Protein Structure Prediction,” International Journal of Molecular Sciences 22, no. 11 (2021): 5553, 10.3390/ijms22115553. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 18. Li C., Zhang R., Wang J., Wilson L. M., and Yan Y., “Protein Engineering for Improving and Diversifying Natural Product Biosynthesis,” Trends in Biotechnology 38, no. 7 (2020): 729–744, 10.1016/j.tibtech.2019.12.008. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 19. Faber M. S. and Whitehead T. A., “Data‐Driven Engineering of Protein Therapeutics,” Current Opinion in Biotechnology 60 (2019): 104–110, 10.1016/j.copbio.2019.01.015. [DOI] [PubMed] [Google Scholar]
- 20. Rost B. and Sander C., “Prediction of Protein Secondary Structure at Better Than 70% Accuracy,” Journal of Molecular Biology 232, no. 2 (1993): 584–599, 10.1006/jmbi.1993.1413. [DOI] [PubMed] [Google Scholar]
- 21. Korendovych I. V. and DeGrado W. F., “De Novo Protein Design, a Retrospective,” Quarterly Reviews of Biophysics 53 (2020): e3, 10.1017/S0033583519000131. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 22. Dill K. A. and MacCallum J. L., “The Protein‐Folding Problem, 50 Years on,” Science 338, no. 6110 (1979): 1042–1046, 10.1126/science.1219021. [DOI] [PubMed] [Google Scholar]
- 23. Sun M. G. F. and Kim P. M., “Data Driven Flexible Backbone Protein Design,” PLoS Computational Biology 13, no. 8 (2017): e1005722, 10.1371/journal.pcbi.1005722. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 24. Shultis D., Mitra P., Huang X., et al., “Changing the Apoptosis Pathway Through Evolutionary Protein Design,” Journal of Molecular Biology 431, no. 4 (2019): 825–841, 10.1016/j.jmb.2018.12.016. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 25. Chevalier A., Silva D. A., Rocklin G. J., et al., “Massively Parallel De Novo Protein Design for Targeted Therapeutics,” Nature 550, no. 7674 (2017): 74–79, 10.1038/nature23912. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 26. Dauparas J., Anishchenko I., Bennett N., et al., “Robust Deep Learning–Based Protein Sequence Design Using ProteinMPNN,” Science (1979) 378, no. 6615 (2022): 49–56, 10.1126/science.add2187. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 27. Qi Y. and Zhang J. Z. H., “DenseCPD: Improving the Accuracy of Neural‐Network‐Based Computational Protein Sequence Design With DenseNet,” Journal of Chemical Information and Modeling 60, no. 3 (2020): 1245–1252, 10.1021/acs.jcim.0c00043. [DOI] [PubMed] [Google Scholar]
- 28. Anand N., Eguchi R., Mathews I. I., et al., “Protein Sequence Design With a Learned Potential,” Nature Communications 13, no. 1 (2022): 746, 10.1038/s41467-022-28313-9. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 29. Zhang Y., Chen Y., Wang C., et al., “ProDCoNN: Protein Design Using a Convolutional Neural Network,” Proteins: Structure 88 (2019): 819–829, https://api.semanticscholar.org/CorpusID:209447585. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 30. Huang B., Fan T., Wang K., et al., “Accurate and Efficient Protein Sequence Design Through Learning Concise Local Environment of Residues,” Bioinformatics 39, no. 3 (2023): btad122, 10.1093/bioinformatics/btad122. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 31. Liu Y., Zhang L., Wang W., et al., “Publisher Correction: Rotamer‐Free Protein Sequence Design Based on Deep Learning and Self‐Consistency,” Nature Computational Science 2, no. 8 (2022): 526, 10.1038/s43588-022-00305-1. [DOI] [PubMed] [Google Scholar]
- 32. Coley C. W., Jin W., Rogers L., et al., “A Graph‐Convolutional Neural Network Model for the Prediction of Chemical Reactivity,” Chemical Science 10, no. 2 (2019): 370–377, 10.1039/C8SC04228D. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 33. Rahban M., Ahmad F., Piatyszek M. A., Haertlé T., Saso L., and Saboury A. A., “Stabilization Challenges and Aggregation in Protein‐Based Therapeutics in the Pharmaceutical Industry,” RSC Advances 13, no. 51 (2023): 35947–35963, 10.1039/D3RA06476J. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 34. Feng Y., Su C., Mao G., et al., “When Synthetic Biology Meets Medicine,” Lifestyle Medicine 3, no. 1 (2024): lnae010, 10.1093/lifemedi/lnae010. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 35. Campeotto I., Goldenzweig A., Davey J., et al., “One‐Step Design of a Stable Variant of the Malaria Invasion Protein RH5 for Use as a Vaccine Immunogen,” Proceedings of the National Academy of Sciences of the United States of America 114, no. 5 (2017): 998–1002, 10.1073/pnas.1616903114. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 36. Lu H., Diaz D. J., Czarnecki N. J., et al., “Machine Learning‐Aided Engineering of Hydrolases for PET Depolymerization,” Nature 604, no. 7907 (2022): 662–667, 10.1038/s41586-022-04599-z. [DOI] [PubMed] [Google Scholar]
- 37. Shahryari A., Saghaeian Jazi M., Mohammadi S., et al., “Development and Clinical Translation of Approved Gene Therapy Products for Genetic Disorders,” Frontiers in Genetics 10 (2019), https://www.frontiersin.org/journals/genetics/articles/10.3389/fgene.2019.00868. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 38. Ki M. R. and Pack S. P., “Fusion Tags to Enhance Heterologous Protein Expression,” Applied Microbiology and Biotechnology 104, no. 6 (2020): 2411–2425, 10.1007/s00253-020-10402-8. [DOI] [PubMed] [Google Scholar]
- 39. Chandler R. J., Tsai M. S., Dorko K., et al., “Adenoviral‐Mediated Correction of Methylmalonyl‐CoA Mutase Deficiency in Murine Fibroblasts and Human Hepatocytes,” BMC Medical Genetics 8, no. 1 (2007): 24, 10.1186/1471-2350-8-24. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 40. Zhang S., Bastille A., Gordo S., et al., “Novel AAV‐Mediated Genome Editing Therapy Improves Health and Survival in a Mouse Model of Methylmalonic Acidemia,” PLoS One 17, no. 9 (2022): e0274774, 10.1371/journal.pone.0274774. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 41. Couchet M., Breuillard C., Corne C., et al., “Ornithine Transcarbamylase—From Structure to Metabolism: An Update,” Frontiers in Physiology 12 (2021), https://www.frontiersin.org/journals/physiology/articles/10.3389/fphys.2021.748249. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 42. Tokuriki N. and Tawfik D. S., “Stability Effects of Mutations and Protein Evolvability,” Current Opinion in Structural Biology 19, no. 5 (2009): 596–604, 10.1016/j.sbi.2009.08.003. [DOI] [PubMed] [Google Scholar]
- 43. Duff C., Alexander I. E., and Baruteau J., “Gene Therapy for Urea Cycle Defects: An Update From Historical Perspectives to Future Prospects,” Journal of Inherited Metabolic Disease 47, no. 1 (2024): 50–62, 10.1002/jimd.12609. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 44. Klein A. D. and Outeiro T. F., “Glucocerebrosidase Mutations Disrupt the Lysosome and Now the Mitochondria,” Nature Communications 14, no. 1 (2023): 6383, 10.1038/s41467-023-42107-7. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 45. Massaro G., Geard A. F., Liu W., et al., “Gene Therapy for Lysosomal Storage Disorders: Ongoing Studies and Clinical Development,” Biomolecules 11, no. 4 (2021): 611, 10.3390/biom11040611. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 46. Andreotti G., Citro V., and De Crescenzo A., “Therapy of Fabry Disease With Pharmacological Chaperones: From In Silico Predictions to In Vitro Tests,” Orphanet Journal of Rare Diseases 6 (2011): 66. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 47. Coura R. d. S. and Nardi N. B., “The State of the Art of Adeno‐Associated Virus‐Based Vectors in Gene Therapy,” Virology Journal 4, no. 1 (2007): 99, 10.1186/1743-422X-4-99. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 48. Wu Z., Gui S., Wang S., and Ding Y., “Molecular Evolution and Functional Characterisation of an Ancient Phenylalanine Ammonia‐Lyase Gene (NnPAL1) From Nelumbo nucifera: Novel Insight Into the Evolution of the PAL Family in Angiosperms,” BMC Evolutionary Biology 14, no. 1 (2014): 100, 10.1186/1471-2148-14-100. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 49. Carlini V., Noonan D. M., Abdalalem E., et al., “The Multifaceted Nature of IL‐10: Regulation, Role in Immunological Homeostasis and Its Relevance to Cancer, COVID‐19 and Post‐COVID Conditions,” Frontiers in Immunology 14 (2023), https://www.frontiersin.org/journals/immunology/articles/10.3389/fimmu.2023.1161067. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 50. Kuang P. P., Liu X. Q., Li C. G., et al., “Mesenchymal Stem Cells Overexpressing Interleukin‐10 Prevent Allergic Airway Inflammation,” Stem Cell Research & Therapy 14, no. 1 (2023): 369, 10.1186/s13287-023-03602-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 51. Grimstein C., Choi Y. K., Wasserfall C. H., et al., “Alpha‐1 Antitrypsin Protein and Gene Therapies Decrease Autoimmunity and Delay Arthritis Development in Mouse Model,” Journal of Translational Medicine 9, no. 1 (2011): 21, 10.1186/1479-5876-9-21. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 52. Qing R., Hao S., Smorodina E., Jin D., Zalevsky A., and Zhang S., “Protein Design: From the Aspect of Water Solubility and Stability,” Chemical Reviews 122, no. 18 (2022): 14085–14179, 10.1021/acs.chemrev.1c00757. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 53. Thumuluri V., Martiny H. M., Almagro Armenteros J. J., Salomon J., Nielsen H., and Johansen A. R., “NetSolP: Predicting Protein Solubility in Escherichia coli Using Language Models,” Bioinformatics 38, no. 4 (2022): 941–946, 10.1093/bioinformatics/btab801. [DOI] [PubMed] [Google Scholar]
- 54. Guruprasad K., Reddy B. V. B., and Pandit M. W., “Correlation Between Stability of a Protein and Its Dipeptide Composition: A Novel Approach for Predicting In Vivo Stability of a Protein From Its Primary Sequence,” Protein Engineering, Design & Selection 4, no. 2 (1990): 155–161, 10.1093/protein/4.2.155. [DOI] [PubMed] [Google Scholar]
- 55. Ferreira L. G., Dos Santos R. N., Oliva G., and Andricopulo A. D., “Molecular Docking and Structure‐Based Drug Design Strategies,” Molecules 20, no. 7 (2015): 13384–13421, 10.3390/molecules200713384. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 56. Vihinen M., Torkkila E., and Riikonen P., “Accuracy of Protein Flexibility Predictions,” Proteins: Structure, Function, and Bioinformatics 19, no. 2 (1994): 141–149, 10.1002/prot.340190207. [DOI] [PubMed] [Google Scholar]
- 57. Diaconis P. and Saloff‐Coste L., “What Do We Know About the Metropolis Algorithm,” Journal of Computer and System Sciences 57, no. 1 (1998): 20–36, 10.1006/jcss.1998.1576. [DOI] [Google Scholar]
- 58. Gao Z., Tan C., Zhang Y., Chen X., Wu L., and Li S. Z., “ProteinInvBench: Benchmarking Protein Inverse Folding on Diverse Tasks, Models, and Metrics,” in Advances in Neural Information Processing Systems, vol. 36, ed. Oh A., Naumann T., Globerson A., Saenko K., Hardt M., and Levine S. (Curran Associates, Inc., 2023), 68207–68220, https://proceedings.neurips.cc/paper_files/paper/2023/file/d73078d49799693792fb0f3f32c57fc8‐Paper‐Datasets_and_Benchmarks.pdf. [Google Scholar]
- 59. Zhang Y. and Skolnick J., “TM‐Align: A Protein Structure Alignment Algorithm Based on the TM‐Score,” Nucleic Acids Research 33, no. 7 (2005): 2302–2309, 10.1093/nar/gki524. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 60. Scott L. P. B., Chahine J., and Ruggiero J. R., “Using Genetic Algorithm to Design Protein Sequence,” Applied Mathematics and Computation 200, no. 1 (2008): 1–9, 10.1016/j.amc.2007.09.033. [DOI] [Google Scholar]
- 61. Bairoch A. and Apweiler R., “The SWISS‐PROT Protein Sequence Database and Its Supplement TrEMBL in 2000,” Nucleic Acids Research 28, no. 1 (2000): 45–48. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 62. Ben Chorin A., Masrati G., Kessel A., et al., “ConSurf‐DB: An Accessible Repository for the Evolutionary Conservation Patterns of the Majority of PDB Proteins,” Protein Science 29, no. 1 (2020): 258–267, 10.1002/pro.3779. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 63. Gao Z., Tan C., and Li S. Z., “AlphaDesign: A Graph Protein Design Method and Benchmark on AlphaFoldDB,” published online February 1, 2022, http://arxiv.org/abs/2202.01079.
- 64. Hsu C., Verkuil R., Liu J., et al., “Learning Inverse Folding From Millions of Predicted Structures,” bioRxiv 162 (2022): 8946–8970, 10.1101/2022.04.10.487779. [DOI] [Google Scholar]
- 65. Gamage D. G., Gunaratne A., Periyannan G. R., and Russell T. G., “Applicability of Instability Index for In Vitro Protein Stability Prediction,” Protein and Peptide Letters 26, no. 5 (2019): 339–347, https://api.semanticscholar.org/CorpusID:73473690. [DOI] [PubMed] [Google Scholar]
- 66. Joachimiak L. A., Walzthoeni T., Liu C. W., Aebersold R., and Frydman J., “The Structural Basis of Substrate Recognition by the Eukaryotic Chaperonin TRiC/CCT,” Cell 159, no. 5 (2014): 1042–1055, 10.1016/j.cell.2014.10.042. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 67. Ferruz N., Heinzinger M., Akdel M., Goncearenco A., Naef L., and Dallago C., “From Sequence to Function Through Structure: Deep Learning for Protein Design,” Computational and Structural Biotechnology Journal 21 (2023): 238–250, 10.1016/j.csbj.2022.11.014. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 68. Marsh J. A. and Teichmann S. A., “Structure, Dynamics, Assembly, and Evolution of Protein Complexes,” Annual Review of Biochemistry 84 (2015): 551–575, 10.1146/annurev-biochem-060614-034142. [DOI] [PubMed] [Google Scholar]
- 69. Rapp J. T., Bremer B. J., and Romero P. A., “Self‐Driving Laboratories to Autonomously Navigate the Protein Fitness Landscape,” Nature Chemical Engineering 1, no. 1 (2024): 97–107, 10.1038/s44286-023-00002-4. [DOI] [PMC free article] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Data S1. Supporting Information.
Data Availability Statement
The data that support the findings of this study are available in aysenursoyturk at https://github.com/aysenursoyturk/HMHO. These data were derived from the following resources available in the public domain: UniProt Protein Database, https://www.uniprot.org; NCBI GenBank, https://www.ncbi.nlm.nih.gov/genbank/.
