Skip to main content
GigaScience logoLink to GigaScience
. 2026 May 13;15:giag056. doi: 10.1093/gigascience/giag056

Accurate proteome-wide prediction of enzymes and catalytic sites using graph deep learning and protein language model

Yui-Lun Ng 1, Xiaomei Wang 2, Yingqi Li 3, Junzhe Huang 4,5,6,7, Hei Ming Lai 8,9,10,11,12, Jason Ying-Kuen Chan 13, Billy Wai-Lung Ng 14,15,16, Ho Ko 17,18,19,20,21,✉, Ka-Wai Kwok 22,23,✉
PMCID: PMC13217607  PMID: 42126889

Abstract

Identifying the enzyme functions of proteins and their catalytic residues is vital to our understanding of diverse cellular processes. However, existing frameworks that can concurrently determine the enzymatic functions and active sites of proteins are scarce, and still have much room for improvement in prediction performance. In this study, we present EC-LMGraph, a protein language model- and graph convolutional network-based framework to predict enzyme commission (EC) numbers from protein sequence features and structures, and saliency mapping to score representative residues attribute to the enzymatic functions. EC-LMGraph attained an average F1 score of 0.77 in 3rd-level EC number prediction, and 0.76 in 4th-level prediction, outperforming numerous other algorithms that were either sequence-based only, or additionally incorporated structural information. Benchmarking on the Mechanism and Catalytic Site Atlas dataset and a set of Parkinson’s disease-related proteins, we showed that EC-LMGraph showed a stronger emphasis on catalytic sites than the current state-of-the-art algorithm DeepFRI. Combining EC-LMGraph with AlphaFold2, our framework correctly determined the 3rd-level EC numbers of 229,160 proteins based purely on their predicted structures. We show that EC-LMGraph is capable of accurately predicting the 3rd/4th-level EC numbers, and pinpointing the key amino acid residues for many enzymes. EC-LMGraph is implemented and freely available at https://github.com/ngyuilun/EC-LMGraph.

Keywords: enzyme function prediction, catalytic site prediction, protein language model, graph convolution network, deep learning

Introduction

Enzymes constitute a large class of proteins. Identifying enzymes and revealing how they function is crucial for understanding the mechanisms of cellular processes and disease pathophysiology. Known enzymatic functions, along with other protein characteristics (e.g., amino acid sequence, variants, and other molecular functions), are comprehensively documented in an open access database, the UniProt Knowledgebase [1] (UniProtKB), serving as an indispensable resource in biomedical research. With tremendous efforts over the past two decades, the number of sequence entries that had been manually annotated in UniProtKB reached 573,230 (as of 2025–04), which is, however, only around 0.23% of all recorded entries. The remaining vast majority of proteins (>252 million) could only be annotated using UniProt’s rule-based automatic annotation systems [2]. Even with high-throughput assays, identifying or verifying the enzyme-catalyzed reactions of the unreviewed proteins, or just annotating the known catalytic domains of enzymes, would still demand massive amounts of time and effort.

To accelerate the process of protein function determination, a method for accurate prediction of enzyme functions with robust functional site annotation is strongly desired. Most existing data-driven enzyme prediction frameworks are sequence-based [3–11] (i.e., relying on just the primary structure), as protein sequences are abundantly available. Although many of the sequence-based tools exhibit robust performances in inferring enzymatic functions, most of them focus solely on identifying the enzyme classes without indicating the putative catalytic residues. While some sequence-based methods may effectively identify catalytic sites that involve consecutive residues, due to the lack of geometric coordinates, they may not accurately detect catalytic or important residues that are physically close in the folded protein but distant in sequence (e.g., located in a different fold).

Structure-based methods have been commonly adopted in drug discovery (e.g., in virtual screening with molecular docking and molecular dynamics simulation) [12–14], since the interactions of proteins and other molecules are constrained by their shapes, surface charges, and other structural properties. Despite variations in the primary sequence, similar enzyme functions are often mediated by a small number of residues within the active domains with highly conserved local structures. The secondary and tertiary structures of proteins can therefore be promising predictors for their enzymatic functions [15–18]. To capture the intrinsic three-dimensional (3D) structures of proteins, a graph-based representation is a powerful approach, as the chemical properties of amino acids and their pairwise interactions can be represented by nodes and edges, respectively [19]. To utilize such a representation in machine learning frameworks, graph convolution operators can propagate node information, such that node properties can be aggregated and integrated by a graph convolutional network (GCN) [20].

In a recent work, DeepFRI [16] employed a long short-term-memory language model and GCN to annotate protein functions and detect functional regions with an average Fmax score of ∼0.5–0.6. Nevertheless, the prediction performance still has room for improvement, particularly for the sequence feature extraction process that had not yet benefited from the recent advancements in large language models. Protein language models (pLMs), such as ProtT5 [21] and ESM-2 [22], leverage state-of-the-art transformer-based architectures [23] to generate embedded protein representations. These pLMs are comprised of billions of parameters (∼3 billion in ProtT5, and ∼15 billion in ESM-2), which enable them to extract relevant features from millions of sequences in UniRef (∼45 million on UniRef50, and ∼216 million on UniRef100). By utilizing these pre-trained models as a feature extraction module, it is possible to effectively capture the complex relationship between protein sequences and their functions. We hypothesized that the main limitation of several structure-based techniques for protein function prediction, specifically the relative scarcity of experimentally validated structures, could be alleviated by incorporating the protein sequence features captured by pLMs. These features have the potential to serve as supplementary input and enhance the performance in enzyme function prediction.

Here, we present EC-LMGraph, a deep learning framework that combines pLM with GCN to predict enzyme functions through learning graph representations of experimentally determined protein structures from the Protein Data Bank (PDB) [24]. In EC-LMGraph, we adopted a GCN architecture with a local extremum and graph convolution operators block, and performed training on three distinct, complementary types of data for each protein: the primary sequence, the pLM-embedded feature, and the structure graph. EC-LMGraph attained an average F1 score of 0.77 in 3rd-level enzyme commission (EC) number prediction, and 0.76 in 4th-level prediction, outperforming numerous other sequence-only or both sequence- and structure-based algorithms. To highlight protein regions essential for enzyme function prediction, we mapped the activation of EC-LMGraph on amino acid residues using model interpretability algorithms, and observed a high concordance between such regions and the catalytic residues curated in the Mechanism and Catalytic Site Atlas (M-CSA) database [25–27]. In head-to-head comparisons, EC-LMGraph outperformed DeepFRI on catalytic sites prediction for the M-CSA entries and a set of Parkinson’s disease (PD)-related enzymes. The EC-LMGraph source code and all prediction results are freely available and can be accessed online [28].

Results

Overview of the EC-LMGraph framework

The goal of our framework is to train GCNs, which combine the sequence embeddings and graph representations of protein structures to predict their EC numbers and catalytic residues accurately (Fig. 1). Protein structure data, whether determined through experimental methods or predicted using computational approaches, can be taken as input in our framework. The structure of each protein was first pre-processed into an adjacency matrix and a feature matrix to obtain a protein graph format. The adjacency matrix encodes the pairwise Euclidean distances between the alpha carbon of amino acid residues, whereas the feature matrix encodes the amino acid identity. To prevent excess information transmission between graph nodes while preserving sufficient details of protein structures, we applied distance-thresholding to the adjacency matrix using 9 Å as the optimal cutoff (Fig. 1, Fig. S1). We analyzed the trade-offs between precision and recall across thresholds ranging from 6 Å to 13 Å (Fig. S1B). We observed that precision remained consistently high (>0.84) for cutoffs between 8 Å and 10 Å. Regarding recall, the model’s ability to retrieve true positives was superior when the cutoff was between 7 and 9 Å, with the 9 Å cutoff achieving the maximum recall of 0.72. Consequently, the 9 Å cutoff yielded the highest overall F1 score (0.78), effectively balancing the capture of relevant structural features against the inclusion of noise. To determine whether a threshold preference exists for specific enzyme classes, we further stratified the performance of 3rd- and 4th-digit predictions across major EC categories (Table S1). A distance cutoff of 9 Å consistently demonstrated robustness across diverse functional categories. At the 3rd-digit level, the 9 Å threshold is optimal for six of the seven classes; Hydrolases (EC 3) is the sole exception, slightly favoring a 10 Å cutoff. Regarding 4th-level predictions, the 9 Å cutoff remains the top performer for five classes, while Lyases (EC 4) and Isomerases (EC 5) exhibit minor preferences for 7 Å and 8 Å, respectively. Despite these slight variations, the 9 Å threshold provides the most consistent and accurate performance across the majority of enzyme functions. Hence, we empirically chose the 9 Å cutoff as optimal for defining residue contacts to effectively learn graph structures. Through these preprocessing steps, each protein structure was modeled as a node property-preserved, thresholded, and undirected graph.

Figure 1.

A two-panel flowchart outlining the EC-LMGraph framework. The top panel illustrates the data pipeline from left to right: input protein structures from experimental or computational sources are translated into a graph representation consisting of a one-hot encoded feature matrix and a binarized adjacency matrix. These matrices feed into the EC-LMGraph model, which outputs enzyme class and catalytic residue predictions. The bottom panel breaks down the EC-LMGraph neural network architecture. An input graph passes through a protein language module using ProtTrans to create feature embeddings. These embeddings enter a graph learning module consisting of sequential layers: local extrema convolution with ReLU, ClusterGCN with ReLU, global mean pooling, and fully connected layers. The final outputs are the predicted enzyme classes and a visual representation that highlights key residues based on their positions or spatial locations.

Overview of EC-LMGraph framework. Protein structure data, whether determined through experimental methods or predicted using computational approaches, can be taken as input in EC-LMGraph framework. The amino acid sequence and structure are represented as a feature matrix (with one-hot encoding) and an adjacency matrix (binarized with a 9 Å distance cut-off), respectively. The EC-LMGraph models employ a protein language model (ProtTrans) to generate feature embedding for each amino acid sequence, and a graph convolution module to learn and predict EC numbers from the input sequences, embedded features, and structure graphs. To identify which residue contributes to the prediction of the enzyme class, explainability methods were employed to calculate the importance of each amino acid residue. The importance values can be mapped to the corresponding amino acid residues such that a visual representation that highlights key residues based on their positions or spatial locations can be generated.

EC-LMGraph employs a protein language module to transform the sequence into an efficient representation that captures the biological characteristics, and a graph convolution network module to disseminate the residual-level characteristics among residues located in close proximity in the three-dimensional space (Fig. 1). The protein language module incorporates a top-performing model from ProtTrans, namely ProtT5-XL-U50 [21]. The model was pre-trained using a vast dataset of over 2,122 M protein sequences from Big Fantastic Database (BFD) and further refined using additional 45 M protein sequences from the UniProt database. Note that this language model utilized self-supervised learning such that neither annotations nor labels were used to guide the training process, therefore it can make full use of the sequences in the UniProt database. The protein language module processes the input sequence and generates the corresponding feature embeddings. The adjacency matrix, feature matrix, and feature embeddings are then fed into the graph convolution module for learning the structure-function relationships (Fig. 1). The last layer of the graph convolution module is connected to global pooling operators and fully connected layers to output the final enzyme function predictions. In order to identify which residue contributes to the prediction of the enzyme class, we incorporated explainability methods to calculate the importance of each amino acid residue. By mapping the importance values onto the corresponding amino acid residues, our framework provides a visual representation that highlights key residues based on their positions or spatial locations.

We experimented on five types of graph convolutions, including the widely used graph convolutional layer (GCNConv) [20], graph attention (GATv2Conv) [29], hypergraph attention (HypergraphConv) [30], an efficient graph clustering algorithm (ClusterGCNConv) [31], and local extremum convolution (LEConv) [32], to investigate their efficacies in learning the structural representations of proteins. We compared the architectures incorporating these layers on (i) 3rd/4th-level EC class prediction performance, and (ii) number of catalytic sites identified for enzymes in the M-CSA database (Fig. S2A, also see later sections on catalytic site prediction evaluation). Overall, GCN attained the highest F1 score in EC number prediction (Fig. S2A), yet performed poorly on catalytic site prediction (Fig. S2B). ClusterGCNConv also demonstrated a relatively high score in EC prediction performance (Fig. S2A), and outperformed GCN for catalytic site identification (Fig. S2B). While LEConv allowed the most accurate catalytic site inference (Fig. S2B), an architecture solely based on LEConv exhibited more variable EC prediction performance (Fig. S2A). Balancing across the metrics, the combined LEConv-ClusterGCNConv (LE-ClusterGCN) architecture achieved relatively high scores in EC number classification (Fig. S2A), while permitting reasonably accurate catalytic site identification (Fig. S2B).

Many enzymes in nature are multi-functional, catalyzing multiple distinct reactions and thus carrying several EC numbers. To accommodate these cases, we have designed EC-LMGraph to compute independent probabilities for every EC class across the functional hierarchy (see the “Materials and methods” section). A challenge in multi-label classification is that the amount of training samples in each class is highly imbalanced, with many more negative than positive samples (e.g., for any given protein functional class, there are far more proteins that do not belong to the class) [33, 34]. Without measures to tackle this problem, graph neural network (GNN) classifiers often over-classify the majority of negative classes and fail to discriminate against the positive classes in the training processes [35]. Such an issue would be even more severe for 3rd/4th-level EC classes than 1st/2nd-level ones. To tackle the data imbalance issue and improve the classification performance, we employed a focal loss function [36] to guide the model to focus on learning from the minority class. Focal loss assigns a higher weight to the misclassified examples, thereby reducing the contribution of well-classified examples and increasing the contribution of misclassified examples during the training process. We evaluated the model’s performance using focusing parameters (γ) ranging from 1 to 4 and compared these results to the baseline binary cross-entropy (BCE) loss within a five-fold cross-validation setting. The model optimized with baseline BCE loss attained an F1 score of ∼0.68. By introducing the focal loss, we can accomplish a substantial improvement across all γ values tested, such that the F1 scores can consistently reach up to ∼0.77–0.78 (Fig. S3A). All tested γ values yield consistently high scores across folds, suggesting that the focal loss mechanism itself drives the performance gain. As the performance variations among the tested γ parameters were marginal, we adopted the standard setting of γ = 2. Apart from comparing different graph convolutions (Fig. S2), ablation studies were performed to assess the effectiveness of the protein language modules and graph network architecture. The use of the protein language module substantially improved the EC number prediction performance, from an average F1 score of ∼0.48 to ∼0.78 (Fig. S3B). To assess the importance of our proposed graph network architecture, additional models were trained by replacing the graph convolution module with fully connected network or 1D-convolutional neural network (Fig. S3B). Both architectures demonstrated a decrease in prediction accuracy, as the fully connected network combined with the protein language model only achieved an F1 score of 0.46, while the convolutional neural network attained a F1 score of 0.48. These results emphasize the importance of graph convolution module in learning and integrating structural information to achieve optimal performance. To assess the influence of mainstream protein language models on enzyme prediction, we trained LE-ClusterGCN architecture using embeddings from ProtT5-XL-U50 [21], ProtBert [37], and ESM-2 [22]. All three models demonstrated comparable performance, yielding robust F1 scores ranging from 0.77 to 0.78 (Fig. S3B). The overlapping confidence intervals (ProtT5: 0.76–0.79; ProtBert: 0.78–0.79; ESM-2: 0.78–0.78) indicate that the prediction accuracy is not heavily dependent on the specific pLM architecture. This consistency suggests that LE-ClusterGCN architecture effectively integrates structural graph features with high-dimensional sequence embeddings, regardless of whether the embeddings are derived from encoder-only (ProtBert, ESM-2) or encoder-decoder (ProtT5) architectures. Given this robustness, we chose ProtT5 embeddings as the representative sequence feature extractor for the subsequent analysis. Settling on the LE-ClusterGCN-based design, the parameters of EC-LMGraph were optimized through backpropagation of focal loss. The optimal model with the highest validation score was chosen for further evaluations.

EC-LMGraph performance evaluation with temporal holdout validation

We composed a temporal holdout dataset to evaluate the performance of EC-LMGraph in a realistic scenario, by identifying newly annotated protein structures in the UniProtKB database between two releases, namely 2022_01 (Feb 2022) and 2025_02 (Apr 2025). Protein structures with enzyme functions annotated in release 2022_01 were used as the training set, while the test set contained those newly annotated enzymes in release 2025_02. Protein functions that were annotated in both releases were categorized as previously known enzymes and therefore were not included in the test set to avoid data leakage. This test set represents a diverse collection of protein functions that have been identified and documented over a one-year period. To focus on a more specific level of enzyme function (i.e., the 3rd/4th-digits), enzymes lacking annotations for their third digit were excluded. As the experimental validated structures deposited in PDB can represent the same protein sequence, a sequence clustering algorithm CD-HIT [38] was applied on the sequence of the PDB structures to cluster the data using an identity cut-off of 95%. This procedure reduced the presence of homologous proteins between the training set and test set. The training set contained 16,263 (∼83%) protein structures, and the test set contained 3,378 (∼17%) structures. To provide more reliable estimates of the model’s performance, we applied a five-fold cross-validation to the training set and employed iterative stratification [39] to ensure that the numbers of samples in each EC class were balanced. By employing a training process that includes five-fold cross-validation and utilizing the test set as an independent dataset, the models were evaluated on unseen protein structures such that any model overfitting can be detected and mitigated.

We first trained and assessed the performance of our GCN architecture with six evaluation metrics, namely precision, recall, accuracy, specificity, F1 score, and Matthews Correlation Coefficient (MCC) (Fig. 2A, Fig. S3C). Given that this is a multi-label classification task with highly imbalanced classes, the evaluation metrics were computed using the micro-average, as it can reflect the overall performance and is less influenced by the performance of rare classes. When testing on the test set of 3,378 structures, the 3rd- and 4th-digit EC predictions of LE-ClusterGCN attained a minimum of 0.99 in accuracy and specificity (Fig. S3C), showing that the trained models could discriminate the negative class samples. On the performance of predicting positive cases, the models achieved mean recall scores of 0.71 for 3rd- and 4th-digit predictions (i.e., these models can predict the enzyme functions of at least 71% proteins in these classes) (Fig. 2A). Overall, we attained an average F1 score of 0.77 for 3rd- and 4th-level prediction (Fig. 2A, also see Fig. S3B for MCC). Regarding the accuracy of predicted positive cases, the GCN models demonstrated mean precision scores of 0.85 for 3rd-digit, 0.81 for 4th-digit, and 0.83 for 3rd/4th-digit predictions. To further validate the confidence of the predictions, higher thresholds for a positive prediction were applied to the predicted probabilities. Even with a threshold of 0.9, the F1 score attained remained at ∼0.77, while the precision can be improved to 0.86 at the expense of a slightly lower recall (Fig. 2B).

Figure 2.

A six-panel figure displaying performance evaluation charts for the EC-LMGraph model. The top row, comprising panels A through C, focuses on newly annotated enzyme structures. Panel A is a box plot showing precision, recall, and F1 scores categorized by the four EC class levels. Panel B is a similar box plot evaluating these same three metrics across five confidence thresholds ranging from 0.5 to 0.9. Panel C is a line graph comparing the F1 scores of four different models against maximum percentage sequence identity, with the EC-LMGraph line charting the highest performance. The bottom row, comprising panels D through F, displays grouped bar charts benchmarking twelve different methods on newly annotated enzyme sequences. Panel D groups the bars by precision, recall, and F1 score. Panel E groups the bars by the 1st, 2nd, 3rd, and 4th EC digits to compare F1 scores across the models. Panel F groups the bars by maximum percentage sequence identity to compare F1 scores. A shared legend at the bottom identifies the color coding for the twelve computational methods compared in the bottom row.

EC-LMGraph prediction performance. (A–C) Benchmarking results on newly annotated enzyme structures. (A) Performance of EC-LMGraph on the prediction of EC main class, subclass, sub-subclass, and sub-sub-subclass numbers, adopting a confidence threshold of 0.5. (B) Dependence of precision, recall, and F1 scores of EC-LMGraph on confidence thresholds ranging from 0.5 to 0.9. (C) Dependence of prediction performances on the training set maximum percentage of sequence identity cut-off for the compared methods. (D–F) Benchmarking results on newly annotated enzyme sequences. (D) and (E) Prediction performances on protein sequences with newly annotated EC numbers for the various compared methods. The input protein structures for the structure-based methods (i.e., EC-LMGraph, DeepFRI, and COFACTOR) were AlphaFold2-predicted structures. For (F), the dependence of prediction performances for 3rd/4th-digit EC numbers on the training set maximum percentage of sequence identity cut-off is shown for the compared methods.

We benchmarked our models against the current state-of-the-art sequence-based and structure-based methods, namely CLEAN [4], CLEAN-Contact [5], and DeepFRI [16], respectively (Fig. 2C). We compared the results using various degrees of sequence similarity (see “Materials and methods” section), ranging from as high as 95% (3,378 structures) to down to 30% (552 structures). EC-LMGraph consistently demonstrated a higher predictive power than the other two methods, as evidenced by the higher F1 scores across all similarity cut-off values (Fig. 2C). Specifically, EC-LMGraph achieved F1 scores of 0.76 at 95% similarity cut-off, and 0.49 at 30% similarity cut-off, CLEAN-Contact achieved F1 scores ranging from 0.43 to 0.60, while CLEAN and DeepFRI achieved F1 scores of 0.34–0.54 and 0.23–0.48, respectively (Fig. 2C).

To broaden the applicability of EC-LMGraph to sequences without experimentally determined structures, we also assessed its performance against several sequence-based methods using computationally predicted structures. We obtained a list of 297 sequences from UniProtKB with newly identified EC functions (i.e., annotated between releases 2022_01 and 2025_02) and utilized AlphaFold2 to derive their corresponding predicted structures. We then compared the performance of EC-LMGraph against eleven other methods, including the sequence alignment method (BLASTp) [3], sequence-based deep learning methods (CLEAN [4], CLEAN-Contact [5], DeepECTransformer [6], ProteInfer [7], ECPred [8], DeepEC [9], GraphEC [10], and ECPICK [11]), and structure-based methods (DeepFRI [16] and COFACTOR [15]). In the overall combined 3rd/4th-level EC evaluation, CLEAN-Contact and DeepECTransformer showed strong performance with F1 scores of ∼0.78 and ∼0.70, respectively, while EC-LMGraph followed closely behind (∼0.69) (Fig. 2D). When breaking down the evaluation by independent hierarchical levels, EC-LMGraph consistently outperformed the other methods on broader functional classes, attaining the highest F1 scores for predicting the 1st-, 2nd-, and 3rd-digit EC classes (0.90, 0.85, and 0.82, respectively) (Fig. 2E). Up to the 4th-digit prediction, CLEAN-Contact achieved the highest F1 score (0.74), followed by DeepECTransformer (0.65). EC-LMGraph achieved results similar to those of BLASTp, CLEAN, and ProteInfer, with scores in the ∼0.53–0.55 range (Fig. 2E). The relative performances remained consistent across sequence similarity levels (Fig. 2F). These results suggest that utilizing precise 3D structural graphs gives EC-LMGraph a distinct advantage in capturing the core catalytic mechanisms that define the 1st-, 2nd-, and 3rd-digit EC classes. In contrast, sequence-based methods utilize massive databases to access a substantial volume of sequence examples, effectively capturing the fine-grained substrate specificities that characterize 4th-digit EC classes. An analysis of the EC-LMGraph training dataset reveals class imbalance as the functional hierarchy deepens (Table S2). While the 1st-digit level contains abundant data (median of 906 samples per class), the dataset would fragment severely owing to the deepen functional hierarchy. At the 4th-digit level, the median drops to just 8 training samples per class, with 67.7% of all classes containing ≤ 10 examples (Table S2). It is worth noting that by incorporating a protein language module pre-trained on large sequence databases, EC-LMGraph is able to learn through a substantially smaller annotated training set (i.e., limited to proteins with experimentally validated structures), even when the numbers of samples in the 4th-digit classes are especially small.

In addition to prediction accuracy, the computational speed of an enzyme function prediction framework is crucial for effectively analyzing large-scale proteome datasets. We conducted an analysis to evaluate the computation time of these EC number prediction frameworks by randomly selecting 100, 200, 500, and 1,000 proteins and recording the time taken (Fig. S3D). Among the twelve evaluated methods, EC-LMGraph ranked fifth (∼240 seconds for 1,000 protein structures) and outperformed other structure-based methods such as DeepFRI and COFACTOR. Despite the substantial number of parameters (∼3B) in the protein language model, the architecture of EC-LMGraph allows for effective handling of the embedded protein representations and structural information while maintaining a comparable computation speed to other sequence-based methods.

These findings thus highlighted the advantages of combining protein language models with structure-based approach in predicting enzymatic functions, especially when the protein structure is available or can be predicted using computational methods.

Catalytic sites prediction based on EC-LMGraph saliency mapping

Proteins possess their enzymatic functions in specific regions where substrates bind to the catalytic domain(s) and undergo chemical reactions. Apart from determining the reaction(s) catalyzed, pinpointing the key amino acid residues constituting the catalytic sites is crucial to understanding the mechanisms of catalysis. Identification of catalytic sites is a complicated process, as these usually occupy only less than 1% of the volume of an enzyme [26], while mutagenesis experiments are time-consuming and resource-intensive. To explore the relationship between learned graph features and catalytic sites, we mapped the activation values of the EC-LMGraph models onto amino acid residues using various explainability methods, including saliency map (Saliency) [40], multiplies gradient with respect to input (InputXGradient) [41], guided backpropagation (GuidedBackprop) [42], and deconvolution (Deconvolution) [43]. These explainability methods were selected based on their model-agnostic nature, as they can be applied to any type of graph layer, regardless of its specific characteristics or structure. This characteristic enabled the selected explainability methods to seamlessly handle the five different types of graph layers that were tested (Fig. S2B), without encountering any limitations. The Saliency method was chosen as the optimal explainability method given that it demonstrated the highest correlation between the predicted and annotated catalytic sites (Fig. S4). The Saliency map derives the node importance based on the magnitude of gradient. A high gradient value suggests that this input node could lead to significant impact on the model’s output, thus highlighting the importance of that residue. By highlighting important amino acid residues for correct classification, we hypothesized that the so-obtained localization map can predict a subset of catalytic sites.

We evaluated the performance of EC-LMGraph catalytic site prediction using M-CSA [25–27], an expert-annotated database documenting enzyme catalytic residues derived from experiments. A catalytic residue was considered to be predicted by EC-LMGraph if (i) the 3rd-/4th-level EC number prediction for the protein structure is correct, and (ii) the activation value at the residue is among top 10% across the whole amino acid chain. Since the exact number of catalytic sites in an uncharacterized enzyme is rarely known, establishing the top 10% of predicted residues as the putative active site could help prioritize the catalytic hotspots for ease of downstream experimental validation. This approach ensures standardized performance comparisons by evaluating an algorithm’s capacity to rank actual catalytic residues within its top predictions. Furthermore, this relative threshold does not introduce length-dependent bias, as the random probability of capturing true sites remains statistically constant regardless of enzyme length. From M-CSA, we identified 4,007 catalytic residues for the protein structures with correct sub-subclass predictions in our dataset. As an illustrating example, for the bacterial leucyl aminopeptidase from V. proteolyticus (PDB: 1LOK) (EC 3.4.11.10), 5 out of 6 catalytic residues were predicted by EC-LMGraph–Saliency (Fig. 3A), and this was extremely unlikely to occur by chance (Fig. 3B). Likewise, further examples from the other major EC classes, including E. coli ribonucleoside reductase (PDB: 5CNV; EC 1.17.4.1), transaldolase B (PDB: 1ONR; EC 2.2.1.2), o-succinylbenzoate synthase (PDB: 1R6W; EC 4.2.1.113), A. pyrophilus glutamate racemase (PDB: 1B73; EC 5.1.1.3), human glutathione synthetase (PDB: 2HGS; EC 6.3.2.3), and rat cytochrome c oxidase (PDB: 1V54; EC 7.1.1.9), as depicted in Fig. 3C, highlighted EC-LMGraph–Saliency’s capability in predicting catalytic sites for enzymes of all seven major EC classes across different species.

Figure 3.

A five-panel figure illustrating EC-LMGraph saliency mapping and residue prediction. Panel A presents a 3D protein ribbon structure alongside horizontal linear sequence bars. Specific amino acid residues on both representations are highlighted using dark red and orange colors to indicate top activation values, and ground-truth catalytic sites are marked. Panel B is a histogram plotting frequency count against mean distance to the nearest catalytic sites. A single dotted vertical line on the left marks the EC-LMGraph predicted distance, while a bell curve on the right represents a random distribution. Panel C displays a grid of six smaller 3D protein ribbon structures from various enzymes, with key residues similarly highlighted in red and orange gradients and labeled with text. Panels D and E occupy the bottom right. Panel D is a line graph plotting the number of covered residues against window size positions, featuring two lines that represent coverage by residues with the top 5% and 10% saliency values. Panel E is a cumulative distribution plot comparing a solid curve for EC-LMGraph against a dashed curve for random predictions based on mean distance z-scores.

EC-LMGraph saliency mapping and catalytic residue prediction. (A) EC-LMGraph–Saliency-predicted catalytic residues of a bacterial leucyl aminopeptidase (PDB: 1LOK). Residues are colored according to the EC-LMGraph saliency values. The ground truth catalytic sites annotated in the Mechanism and Catalytic Site Atlas (M-CSA) are marked above the amino acid chain for reference. (B) Comparison of the mean distance (Å) between residues with top 10% activation values and the nearest M-CSA-annotated catalytic sites vs. that expected by random. The mean distance distribution expected by random was generated by randomly sampling the same number of residues across the whole amino acid chain. (C) Examples of enzymes from various species that belong to the different EC classes, showing residues with top EC-LMGraph saliency values and those that coincide with M-CSA-annotated catalytic sites. (D) Total number of M-CSA catalytic residues that are covered by residues with top 5% or 10% EC-LMGraph saliency values. (E) Cumulative distribution plot of the mean distance z-scores of the EC-LMGraph–Saliency predicted residues (to the nearest catalytic residues), compared to that expected with randomly drawn residues.

Overall, 1,480 M-CSA-annotated residues were located at sites with top 10% EC-LMGraph saliency values, accounting for 37% of the total catalytic residues on proteins with correct EC 3rd/4th-digit predictions (Fig. 3D). Since residues near the catalytic ones could also be important for determining the local conformation and hence enzyme function, we postulated that EC-LMGraph activation sites may cluster nearby (as also illustrated by the examples shown in Fig. 3A, C). Consistently, the distances between EC-LMGraph–Saliency-predicted residues were much closer to the M-CSA catalytic sites than by chance (Fig. 3E), with 75% of the annotated catalytic residues locating within ±5-residue windows of top 10% EC-LMGraph activation sites (Fig. 3D).

We additionally benchmarked the performance of EC-LMGraph against DeepFRI on the M-CSA dataset, specifically analyzing cases for which both models produced positive predictions. Higher percentages of M-CSA-annotated residues were located at sites with top 5% or 10% activation values for EC-LMGraph than DeepFRI (Fig. 4A), and similarly when we considered the proportions within ±5-residue windows (Fig. 4A). When considering the utilization of the 5% or 10% activation values as predicted sites and observing the ratio of correctly predicted sites over the overall predicted sites (Fig. S5A), the percentage of correctly predicted sites by EC-LMGraph is consistently higher than DeepFRI. The results indicate that the predicted sites by EC-LMGraph have a higher likelihood of being actual catalytic sites compared to DeepFRI. In addition to evaluating the prediction performance of catalytic residues based on amino acid position, we also tested using window sizes defined by a sphere radius ranging from 3 Å to 7 Å (Fig. S5B). A window size of 0–3 Å did not include any additional residues in the neighborhood; hence, the result was the same as using a residue position-based window size of 0 (i.e., exact position). We also observed that the position window sizes of ±1, ±3 and ±5 demonstrated similar prediction performance compared to Euclidean distance window sizes of 4 Å, 6 Å, and 7 Å, respectively. These suggested that the prediction accuracy achieved based on amino acid position aligns closely with the performance achieved using window size based on Euclidean distance. Classifying the M-CSA entries by the EC main class numbers, the superior catalytic site prediction performance of EC-LMGraph generalized across enzymes from the sub-subclasses of all seven main classes (Fig. 4B). In line with these, we found shorter distances between EC-LMGraph activation sites and the M-CSA annotated sites, than that obtained with DeepFRI (Fig. 4C, also see Fig. 4D for illustrating examples). EC-LMGraph also consistently outperformed DeepFRI across all evaluation metrics (Table S3). Notably, it achieved substantially higher recall (0.390 vs. 0.224) for superior coverage, more than doubled the MCC score (0.109 vs. 0.047), and improved the area under the precision-recall curve by ∼65% (0.220 vs. 0.133).

Figure 4.

A four-panel figure benchmarking EC-LMGraph against DeepFRI. The top half contains panels A, B, and C side-by-side. Panel A is a line graph comparing the number of covered residues against window size positions for both methods. It uses solid lines for EC-LMGraph and dashed lines for DeepFRI at 5% and 10% activation thresholds. Panel B is a horizontal stacked bar chart showing the number of catalytic sites identified by EC-LMGraph only, DeepFRI only, or both, broken down by the seven main EC classes and an overall summary. Panel C is a cumulative distribution plot comparing mean distance z-scores for EC-LMGraph, DeepFRI, and a random baseline, with intersection points marked by percentages. The bottom half, Panel D, directly compares predictions for tryptophan synthase, displaying EC-LMGraph in the upper section and DeepFRI in the lower section. Both sections share an identical layout: a 3D protein ribbon structure on the left highlighted predicted residues and annotated ground-truth sites, and a vertical dot plot on the right showing window size against the number of covered residues. 

Benchmarking EC-LMGraph against DeepFRI. (A) Total number of catalytic residues annotated in the Mechanism and Catalytic Site Atlas (M-CSA) that are covered by residues with top 5% or 10% activation values, with EC-LMGraph or DeepFRI. (B) Numbers of M-CSA catalytic sites identified by each method, overall and for the enzymes of each main class. (C) Cumulative distribution plots of the mean distance z-scores of the EC-LMGraph–Saliency and DeepFRI–Grad-CAM predicted residues (to the nearest catalytic residues), compared to that expected with randomly drawn residues. (D) Prediction of catalytic residues for the Salmonella typhimurium tryptophan synthase (PDB: 1A50, chain B) by EC-LMGraph (upper panels) vs. DeepFRI (lower panels).

Collectively, these results showed that the features EC-LMGraph learned for enzyme function prediction tend to reside at or near catalytic residues, with EC-LMGraph being more catalytic site-emphasized than DeepFRI.

Proteome-wide enzyme class and functional residue prediction based on predicted structures

Although advancements in X-ray crystallography and cryogenic electron microscopy have greatly accelerated structure identification (∼11,000 entries per year [44]), the growth of the protein structure dataset is still far from keeping pace with new sequence discovery. A common bottleneck for all structure-based function prediction frameworks is therefore the relative scarcity of experimentally determined protein structures. Building on the observation that applying EM-LMGraph on predicted structures performed favorably in comparison to sequence-based methods (Fig. 2D–F), we speculated that the use of high-quality predicted structures such as those by AlphaFold [45, 46] and RoseTTaFold [47, 48] would allow a much broader applicability of EC-LMGraph.

We retrieved the fourth release of 995,411 AlphaFold2-predicted structures, including model organism proteomes (326,175), global health proteomes (238,274), and Swiss-Prot (430,962). Among these, we identified 279,352 protein structures with enzyme functions annotated in UniProtKB (excluding unreviewed entries from UniProtKB TrEMBL), among which 265,331 came from the EC sub-subclasses for which we had sufficient dataset sizes for GCN training. EC-LMGraph correctly determined the 3rd-level EC numbers for 229,160 (∼86.4%) of the proteins based on predicted structures (Fig. 5A). Among entries with 4th-level EC numbers annotated, our method accurately predicted 127,636 out of 147,458 proteins, representing 86.6% of the annotated records. For the human proteome, EC-LMGraph correctly predicted 4,047 out of 5,065 (79.9%) 3rd-level, and 2,139 out of 3,115 (68.7%) 4th-level EC numbers (Fig. 5B). We also quantified the prediction performance with evaluation metrics (pooled across sub-subclasses, see Table S4). Across species, the F1 scores and MCCs reflected the highest model performances for human and several model organisms, including mouse, rat, zebrafish, and fruit fly, with F1 scores and MCCs of 0.72–0.78 (Table S4). EC-LMGraph therefore permits proteome-wide prediction of enzyme functions in various species even with only algorithm-predicted structures.

Figure 5.

A four-panel figure illustrating enzymatic predictions based on AlphaFold2 structures. Panel A displays five sets of horizontal stacked bar charts categorized by dataset, including Overall, Global health proteomes, Model organism proteomes, Swiss-Prot, and Homo sapiens. Each set contains two bars representing the 3rd and 4th digit Enzyme Commission predictions. The bars are color-coded in dark and light shades to show the proportion of correct versus incorrect predictions, alongside numerical totals and F1 scores. Panel B is a line graph plotting the number of covered residues against window size positions. Two lines track the 5% and 10% top activation value thresholds, with individual data points labeled with percentages. Panels C and D map the predictions for two proteins. Both panels follow an identical layout: the left side displays a 3D protein ribbon structure with predicted catalytic residues highlighted. The right side shows horizontal sequence bars that map these high-activation residues along the amino acid sequence, paired with ground-truth UniProt annotations.

Enzymatic prediction based on AlphaFold2 (AF2)-predicted full-length structures. (A) Number of enzymes with EC numbers correctly predicted by EC-LMGraph based on AF2-predicted structures. (B) Numbers and proportions of catalytic residues covered by residues with top 5% or 10% EC-LMGraph saliency values based on the annotation of Homo sapiens proteins from UniProt. (C) and (D) EC-LMGraph–Saliency-predicted residues on AF2-predicted structures.

As EC-LMGraph relies on the geometric arrangement of residues, the reliability of the AlphaFold2-predicted structure is crucial. We utilized the pLDDT (predicted Local Distance Difference Test), a per-residue confidence metric from AlphaFold2, to assess the confidence of the predicted atomic positions. To ensure the rigor of our predictions on computationally generated structures, we investigated the impact of these confidence scores on model performance by categorizing proteins into three groups: High (pLDDT > 90), Medium (70 < pLDDT ≤ 90), and Low (pLDDT ≤ 70). We observed that the quality of the predicted structure correlates with the accuracy of EC number prediction (Fig. S6). For the 3rd/4th-digit predictions, EC-LMGraph achieve an F1 scores of 0.921 (95% CI: 0.920–0.922) for high-confidence structures and 0.817 (95% CI: 0.814–0.820) for medium-confidence structures. Performance in the low-confidence group decreased to 0.615 (95% CI: 0.602–0.627). These results demonstrate that EC-LMGraph is highly effective for structures with high pLDDT scores, whereas predictions derived from low-confidence models (pLDDT ≤ 70) should be interpreted with caution.

We further tested whether the saliency mapping of EC-LMGraph can be similarly adopted for the identification of important functional residues based on predicted structures. We examined this on human proteins, by selecting a set of previously unsolved sequences and compared the top 10% saliency values with the UniProt-annotated catalytic residues. For proteins with correct EC 3rd/4th-digit predictions, 2,633 (∼29%) UniProt-annotated residues were located at sites with top 10% EC-LMGraph saliency values. As illustrating examples, for the human GPI-linked NAD(P)(+)–arginine ADP-ribosyltransferase 1 (UniProt: P52961; EC 2.4.2.31) and cytosolic phospholipase A2 zeta (UniProt: Q68DD2; EC 3.1.1.4), sites with top EC-LMGraph saliency values correspond very well with UniProt-annotated catalytic residues (Fig. 5C, D). Collectively, based on these results, we concluded that a combined AlphaFold2–EC-LMGraph–Saliency approach can be used to generate hypotheses regarding the positions of functionally important amino acid residues for proteins even with only predicted structures, highlighting putative catalytic or functionally important sites. In principle, EC-LMGraph can also be used in combination with other protein structure prediction algorithms.

Enzymatic prediction for Parkinson’s disease-related proteins

To evaluate the capability of EC-LMGraph in identifying disease-related enzyme functions, we performed a case study on a set of known human PD-related proteins [49–51]. The structures of these proteins had been excluded from the training set, and we tested whether EC-LMGraph could correctly predict their enzyme functions (Fig. 6A). In total, EC-LMGraph gave correct predictions for 8 (6 for 4th-digits) out of the 11 PD-related proteins examined with experimentally determined structures. These include the GTPase domain and its catalytic sites on leucine-rich repeat serine/threonine-protein kinase 2 (LRRK2, Fig. 6B), ubiquitin carboxyl-terminal hydrolase isozyme L1 (UCHL1, Fig. 6C), serine protease HTRA2 (Fig. 6D), protein deglycase DJ-1 (Fig. 6E), parkin (Fig. 6F), glucocerebrosidase (GBA, Fig. 6G), as well as SYNJ1 and POLG (Fig. 6A). For the other PD-related proteins with no known enzyme functions, EC-LMGraph incorrectly predicted VPS35 with EC 2.3.2 and 2.3.1.48, while FBXO7 and SCNA were mis-classified as EC 2.7.11.1.

Figure 6.

An eight-panel figure illustrating enzymatic predictions for Parkinson's disease-related proteins. Panel A is on the left. It is a table comparing annotated Enzyme Commission numbers against EC-LMGraph predictions. The table rows are split into two groups. One group uses experimental structures. The other uses AlphaFold2-predicted structures. Successful predictions are marked in bold text. The rest of the figure maps catalytic site predictions on 3D protein models. Panel B-H displays a 3D protein ribbon structure. Predictied catalytic residues are highlighted.

Enzymatic prediction for Parkinson’s disease (PD)-related proteins. (A) EC-LMGraph predictions for a set of PD-related proteins. (B–H) Catalytic sites prediction for PD-related enzymes using EC-LMGraph.

Similar results were obtained using AlphaFold2-predicted structures, where EC-LMGraph gave consistent results for UCHL1, HTRA2, parkin and GBA, and additionally predicted the known enzyme sub-subclass number for PTEN-induced kinase 1 (PINK1, Fig. 6H) and DNAJC6 (Fig. 6A). While DeepFRI also predicted the enzyme sub-subclass numbers for UCHL1, HTRA2, DJ-1, and parkin, EC-LMGraph consistently showed active site predictions closer to the known experimentally identified catalytic sites annotated in M-CSA and/or UniProt than DeepFRI (see Fig. S7). We thus concluded that EC-LMGraph will complement existing algorithms, and can be applied to predicting protein enzyme functions and corresponding active sites in disease-related settings.

Discussion

Recent advancements in machine learning methods have allowed unprecedented predictions of protein structure [46, 47], function [4, 7–9, 15–17, 52–54], protein–protein interaction [55–58], protein-substrate binding site [59, 60], and protein–nuclei acid interaction [61–63]. With EC-LMGraph, we made further advancements in protein enzyme function prediction. Combining (i) a protein language module that has learned an efficient representation of protein sequences, (ii) an architecture incorporating local extrema and graph clustering convolution operators blocks, (iii) the use of protein feature matrix and protein structure graphs with empirically optimized distance thresholding, and (iv) the increasing availability of solved enzyme structures, EC-LMGraph is capable of accurately predicting the 3rd/4th-level EC (i.e., enzyme sub-subclass or sub-sub-subclass) numbers, and pinpointing the key amino acid residues that constitute the catalytic sites for many enzymes.

Each component of EC-LMGraph is crucial to its performance. For instance, the incorporation of the protein language module led to substantial improvement in prediction accuracy compared to utilizing the graph convolution module alone. This substantiates the role of the protein language module as an effective sequence feature extractor, providing a more comprehensive representation of sequence data than conventional one-hot encoding or LSTM-based models [21]. For the graph convolution modules, the integration of the LEConv layer in the network can significantly boost the performance of catalytic site identification. This convolutional layer facilitates the consideration of both local and global node importance within the protein graph; thereby, the resultant activation mapping can be highly correlated to the annotated catalytic site [32]. In addition, the ClusterGCNConv utilizes a graph clustering algorithm to identify the most crucial nodes within a subgraph, thereby limiting the neighborhood search to this subset. Such an approach enables efficient learning of protein structural information, even for large proteins (e.g., those with over 1,000 amino acids) [31].

The learned graph features in EC-LMGraph can be interpreted as relying on the key residues in protein chains which are more conserved in each enzyme class or sub-subclass. In stringent head-to-head benchmarking, we showed that EC-LMGraph is a more catalytic site-emphasized than DeepFRI. As mutagenesis experiments are relatively costly and time-consuming, EC-LMGraph serves as a constructive tool to streamline experimental design by prioritizing the choice of residues from full sequences to those among or in the vicinity of top ∼10% activation scores. This may especially benefit studies demanding for identifying mutations in enzymes, accounting for changes of function and disease phenotypes [64]. In future work, it would be valuable to extend the framework to predict the mutation-induced loss of catalytic functions, due to the increasing need to understand the impact of genetic variations on enzyme activity. Experimental determination of mutated protein structures is often limited, making it challenging to directly assess the functional consequences of mutations. Thus, accurate prediction of protein structures for mutated variants would be a crucial aspect in training a framework capable of predicting the loss of catalytic function from these mutations. It is important to note that the negative labels in the training data should indicate a complete loss of function. Cases where mutations result in reduced metabolic rates instead of a complete loss may require specific handling or labeling strategies. By addressing these considerations, the framework can be extended to handle mutated protein structures and accurately estimate the impact of mutations on catalytic functions.

The prediction performance of EC-LMGraph certainly still has room for further improvement. With the number of solved enzyme structures rapidly increasing year by year, the capability of EC-LMGraph will also continue to grow, especially for sub-subclass models, which had few samples from prior to our temporal holdout dataset cut-off date to be trained on. Further architectural variations can be explored. For example, additional details in protein structure representations may be incorporated. Apart from using a contact map to represent whether the residues are within the predetermined thresholds, the inter-residue distance, orientation angles between adjacent residues or sidechain dihedral angles could be added to provide a more complete description of local protein conformations. EC-LMGraph currently utilizes static structural representations that do not explicitly account for environmental conditions like pH or temperature. While an enzyme’s intrinsic 3D fold dictates its fundamental EC classification, environmental factors largely influence its reaction kinetics and conformational dynamics. Incorporating environmental metadata is currently bottlenecked by data scarcity, future integration of high-throughput assays or molecular dynamics simulations could potentially overcome these hurdles. For many enzymes, difference in one amino acid residue with crucial physicochemical properties is sufficient to render a given catalytic site inactive. For algorithm-identified catalytic residues, additional filtering based on the known requirements for a given reaction could be used to refine prediction results. In addition to function predictions, recent deep learning approaches have also advanced in predicting enzymatic kinetic parameters. Methodologies utilizing protein language models and graph neural networks, such as UniKP [65] and DEKP [66], have been successfully applied to estimate turnover numbers and Michaelis constants directly from protein sequences and substrate structures. While these tools focus on the quantitative mechanics of enzyme-substrate efficiency, EC-LMGraph could complement this broader ecosystem by focusing on the preceding step: accurate, structure-aware prediction of the enzyme function and the localization of its catalytic sites. In future works, incorporating these with other variations in the design of GCNs may enable even more superior enzymatic function and active site predictions.

Materials and methods

EC number annotations of protein chain

Experimental validated structures from the PDB [24] and their corresponding EC number annotations from the UniProtKB [1] were retrieved to train the models. We first had to analyze the annotations in UniProtKB to compile a comprehensive catalog of protein functions for the PDB entries. This process is necessary because many of the structures deposited in the PDB can be multimers of several identical subunits (e.g., 1JU6: a homodimer structure consisting of two identical units) or structural subunits of large proteins (e.g., 7BV1: a complex consisting of the NSP7, NSP8, and NSP12 parts of R1AB_SARS2, where only NSP12 is found to have enzymatic functions). Therefore, the entire structure had to be separated into individual peptide chain structures to ensure that the EC number annotations were accurately associated with the corresponding chain structure.

Given the enzyme function annotations provided by UniProtKB accessions, protein chains were first extracted as sequence intervals labeled with EC numbers. Using the cross-references in each accession, a list of PDB identifiers associated with the sequence intervals was then collected, together with their chain and sequence positions. The start and end positions of the peptide chain were mapped onto the sequence intervals, given with the length-overlap and fraction-overlap computed. To avoid small ligands or protein connector subunits inaccurately categorized as enzymes, mapped records were excluded if (i) the length-overlap was <100 residues, and (ii) the fraction-overlap was <90%. Short sequences (i.e., <100 residues) therefore needed to have a high fraction-overlap of over 90% to be included. Records with annotated EC number only up to the 2nd-digit were excluded (i.e., only those with at least third-level EC numbers were included).

Dataset construction

Training set: A temporal holdout validation setting was employed to evaluate the performance of EC-LMGraph in a realistic scenario. The records were categorized into training or test data based on the difference in EC annotation between successive UniProtKB releases. Specifically, protein structures with EC functions annotated in release 2022_01 were categorized as the training set. With the above EC number annotations steps, 12,355 sequence intervals and their EC annotations were obtained from release 2022_01 (dated 2022 Feb 23), along with the corresponding 85,879 structures from the PDB database. Owing to the presence of numerous duplicate structures within the PDB, it is necessary to employ a sequence clustering algorithm to group the PDB structural data based on their sequence similarity. CD-HIT [38] algorithm was chosen, and an identity cutoffs of 95% was applied to minimize redundancy in the training set, resulting in a reduction of the training set to 16,263 protein structures.

Newly annotated enzyme structure set: To identify a set of newly annotated enzymes, we compared the annotation difference between UniProt release 2022_01 and a later release 2025_02 (dated 2025 Apr 9). Entries with a new EC number emerged in release 2025_02 were categorized as newly annotated enzymes. To ensure redundancy removal between training set and test set, the CD-HIT clustering algorithm with an identity cutoff of 95% was applied to these newly annotated records and the training set. After removing redundancies, records with a structural model deposited in PDB were categorized as newly annotated enzyme structures, which consist of 3,378 unique structures. Similarly, lower degrees of sequence identity cutoffs (30%, 50%, 70%) were applied to reduce sequence redundancy and demonstrate model robustness.

Newly annotated enzyme sequence set: Following the above steps, records without any structural model deposited in PDB were categorized as newly annotated enzyme sequences. Records annotated using the UniProt rule-based system were also excluded to ensure this set of newly annotated enzyme sequences was based on the latest experimental evidence. Several enzyme classes could have identified sequences but lack known protein structures, these classes were not included to ensure a fair comparison between sequence-based and structure-based methods. After exclusion, this newly annotated enzyme sequence set consisted of 297 sequence records, and their corresponding AlphaFold2-predicted structures were retrieved. Lower degrees of sequence identity cutoffs (30%, 50%, 70%) were applied to assess performance and evaluate robustness.

M-CSA set: M-CSA [25–27] serves as a comprehensive database specifically documenting enzyme catalytic residues derived from experiments. A total of 991 PDB entries and their corresponding 5,038 catalytic site residues were retrieved. To remove high similarity side chains or structures, CD-HIT with an identity cutoff of 95% was applied to cluster a non-redundancy set. This reduced the M-CSA set to 937 unique protein chains and 4,469 catalytic site residues.

AlphaFold2-predicted structure set: The latest release (fourth release) of AlphaFold2 Database [45, 46] contains over 200 million computational predicted protein structures. Not only does it provide the predicted structures of human proteome and the proteomes of 47 other key organisms, but it also covers the manually curated set from UniProtKB Swiss-Prot. Among these, we collected 995,411 predicted structures, including 326,175 model organism proteomes, 238,274 global health proteomes, and 430,962 Swiss-Prot. After excluding unreviewed entries from UniProtKB TrEMBL, this set consisted of 657,773 computationally predicted structures.

Protein graph formation

Protein graph data structure can be formed, with (i) the protein constituents (i.e., atoms and residues) represented by graph nodes, and (ii) metrics calculated from the intrinsic 3D atomic coordinates as the graph edges. For our graph models, the nodes are defined at the amino acid residue level. The residue coordinates were primarily based on the atomic coordinates of alpha carbon, while beta carbon and the centroid of each amino acid will be taken as reference in the absence of alpha carbon. This mitigates the effect of any missing backbone carbon atoms due to the low resolution of some experimental structures. To transform the twenty standard amino acid residues using the one-hot encoding [67] method, a node feature (sparse) matrix, Inline graphic is formed, where each matrix column denotes one amino acid type and L is the number of residues in the protein sequence.

To describe the spatial relationship between amino acid residues, we use the inter-residue separation of residues (in Å) to construct a contact map for each protein. A distance cutoff value was adopted to define which edges should be kept. This cutoff was empirically optimized to preserve the structural representation of the catalytic functional regions while discarding unnecessary edges, thereby avoiding noise propagation to neighbors and reducing the computational times required for model training. In prior works [16, 68], cutoff values of 6 Å or 10 Å were usually chosen (i.e., distances smaller than which define in-contact residue pairs) for protein structure graphs. As a result, a contact map of a given protein was devised as an unweighted (binary) adjacency matrix Inline graphic. The number of zero elements in the contact map/matrix is directly determined by the choice of the cutoff. With a higher cutoff, more non-zero edges would be preserved, and the computation time needed increases (Fig. S1A). To determine the optimal distance cutoff, we conducted an analysis using the ground truth labels of the 3,378 test set structures. We tested eight distance thresholds ranging from 6 Å to 13 Å and analyzed their impact on model performance metrics (Precision, Recall, and F1 score) (Fig. S1B). Given the ground truth labels of 3,378 structures in the test sets, precision was maximized within the 8–10 Å range, while recall peaked between 7 and 9 Å. Consequently, the 9 Å threshold yielded the highest overall F1 score (0.78). To investigate whether a threshold preference exists for specific enzyme classes, we evaluated the performance of 3rd- and 4th-digit predictions across the major EC categories. The 9 Å cutoff demonstrated consistent highest F1 score across the majority of classes (Table S1). Based on these findings, we selected 9 Å as the optimal distance cutoff for constructing the protein graphs. Furthermore, the average running time per epoch is illustrated in Fig. S1B. The average training time per epoch ranged from 44 seconds at 6 Å to 93 seconds at 13 Å. For our selected cutoff of 9 Å, the training time was approximately 60 seconds per epoch. Overall, we observed a linear increase in training time of approximately 11.3% for each 1 Å increment.

Protein language model and graph network architecture

The model architecture of EC-LMGraph consists of (i) a protein language module to transform the input sequence into an efficient representation, and (ii) a graph network module for learning the structural representation of a protein. The protein language module is a transformer-based model ProtT5-XL-U50 [21] pre-trained with over 2,122 M protein sequences from BFD and refined with over 45 M protein sequences from UniProt database. To generate the feature representation of a protein, the amino acids are first tokenized and encoded into a numerical format, which is subsequently parsed by the protein language model. The last hidden layers of the ProtT5-XL-U50 model are selected to generate feature embeddings as the latter hidden layer typically captures a higher-level representation of the input protein sequence. The feature embeddings, together with the adjacency matrix and feature matrix, are fed into the graph network module to predict the enzyme functions.

GCNs have shown to be effective in graph feature learning tasks including protein function predictions [16], protein–protein interaction [55], and drug–target interactions [69, 70]. We conducted experiments on five different types of graph convolutions, including the extensively used graph convolutional layer (GCNConv) [20], graph attention (GATv2Conv) [29], hypergraph attention (HypergraphConv) [30], an efficient graph clustering convolution (ClusterGCNConv) [31], and local extremum convolution (LEConv) [32]. We analyze the effect of these convolutional layers on the classification performance of EC function and the number of catalytic sites identified. Considering that the LEConv achieved an exceptionally high identification rate on the catalytic sites and the ClusterGCNConv could attain a relatively higher F1 score, we selected the LEConv and ClusterGCNConv as the building block of our graph convolution module, followed by a mean pooling layer and fully-connected layers for classification purpose.

The graph convolutional layer LEConv takes the feature matrix Inline graphic and the feature embeddings Inline graphic from protein language model as a concatenated matrix Inline graphic (i.e., Inline graphic), and the contact map Inline graphic as the inputs, computing the node embeddings for the next layer, Inline graphic:

graphic file with name TM0010.gif (1)

where, Inline graphic is the dimension of node embeddings in layer l. With the LEConv, the importance of node i with respect to its neighborhood nodes Inline graphic can be represented by the node embedding Inline graphic in the next layer, the node-wise formulation of Equation (1) as defined by Ranjan et al. [32] is

graphic file with name TM0016.gif (2)

where, Inline graphic denotes the edge between node i and its’ neighbor node j in Inline graphic, and Inline graphic are trainable weight matrices for layer l. The output after each LEConv undergoes a rectified linear activation function (ReLU) [71], i.e., Inline graphic, which sets the negative value to 0. Subsequently, the ClusterGCNConv outputs the node embeddings Inline graphic by taking the embeddings Inline graphic updated by LEConv operators and the contact map Inline graphic:

graphic file with name TM0027.gif (3)

The formulation of Equation (3) defined by Chiang et al. [31] is

graphic file with name TM0028.gif (4)

such that Inline graphic, Inline graphicis a diagonal enhancement value and two trainable weight matrices for layer l are denoted by Inline graphic. After the graph convolutional layers, a global mean pooling layer is applied on the resultant node embeddings to obtain a vector representation of the protein structure for graph classification. Such a pooling layer can ensure all protein graphs are represented by a fixed size vector independent of their number of nodes. Subsequent to the pooling layer, two fully connected layers with ReLU activation functions are used to compute the hidden representation from the pooled representation. At last, a fully connected layer is employed, where the number of output neurons corresponds to the total number of EC classes. To support multi-label prediction, each output neuron is activated by an independent sigmoid function, i.e., Inline graphic. This activation function maps the output of neurons to a value between 0 and 1, representing the probability belonging to the specific EC class. These independent sigmoid functions allow the model to simultaneously assign high probabilities to multiple classes.

The networks were supervised by minimizing the focal loss [36] function between the target class output y and the predicted probability p. Given a graph sample, the focal loss is defined as

graphic file with name TM0036.gif (5)

where the probability Inline graphic is

graphic file with name TM0038.gif (6)

The focusing parameter Inline graphic was set to 2, and the batch size was set to 20. Adam optimizer [72] were used to guide the neural network and update the weight parameters. The maximum number of training epochs was 500 and an early stopping criterion was set when training loss, validation loss and mean F1 score could not be further improved by 30 successive epochs. For ablation study, we evaluated the contribution of the protein language modules and the graph network architecture to the overall model performance. To achieve this, we trained additional models with specific modifications to isolate the impact of each component. First, we trained a model without the protein language module to assess its significance in extracting and enriching the protein sequence information. Second, we replaced the graph network module with common neural network architectures: (i) a fully connected network and (ii) a 1D-convolutional neural network. The framework was implemented on the PyTorch Geometric [73] (version 2.0.1) deep learning library, and all of the model training was executed with the use of NVIDIA GeForce RTX 3090 GPUs.

Explainable annotation of catalytic sites

Highly conserved residues within the active domains of protein, namely catalytic sites, can be identified to attribute the enzyme function. Predicting the residues that are directly involved in the catalytic functions could facilitate the design of mutagenesis experiment, thus confirming its enzymatic activity. The 3D conformations and groups of amino acids involved in the catalytic activities become highly relevant features for graph network models to learn and predict the corresponding EC classes. A method to explain the model prediction and quantify the importance of the input amino acid nodes could pinpoint a promising target set of residues in charge of the ultimate enzymatic functioning. Therefore, the activation maps of our network, which are expected to capture the structural features of functional regions, can be post-processed to quantify the learned feature map values onto the input amino acid nodes.

Quantitative evaluation was performed to assess the correspondence between the model-predicted catalytic residues and actual catalytic sites. M-CSA dataset [25] is chosen to derive a set of actual catalytic sites, considering that their catalytic site residues annotations have been experimentally validated. Given a positively predicted protein graph, we computed the activation values using various method, including saliency map (Saliency) [40], multiplies gradient with respect to input (InputXGradient) [41], guided backpropagation (GuidedBackprop) [42], and deconvolution (Deconvolution) [43]. Taking the top 10% residue positions as model-predicted catalytic residues, the saliency method was chosen given that a stronger correlation was observed between the model-predicted catalytic sites and the annotated catalytic sites. Specifically, the Saliency map was computed by assigning relevance scores to the input graph nodes (residues) based on the partial derivative of the model’s output. Next, these scores were aggregated as activation values by taking the absolute value of the gradient, which represents the strength of the activation associated with each input node. The activation values were then normalized into a scale of [0,1] in order to generate the final Saliency map. To evaluate the correspondence, function-specific activation values were computed to identify the top 10% predicted residues. A window size of 0 to ±5 was applied to assess the number of actual catalytic sites that can be covered. The prediction is also evaluated versus random coincidences, the same numbers of amino acid positions (i.e., 10% of the protein length) were randomly drawn across the amino acid chain. This proportional sampling method could reduce size-dependent effects that arise from proteins having different sizes and lengths, allowing for comparison across all proteins. The mean Euclidean distance between the predicted residue sites and the nearest M-CSA sites was calculated to illustrate whether the model-predicted residues would have a lower mean distance than the random drawn. This process is repeated 1,000 times, and the event counts are plotted. The predicted residues were mapped to the PDB structures upon the amino acid position, and then visualized using UCSF Chimera [74].

Availability of source code and requirements

Project name: EC-LMGraph

Project homepage: https://github.com/ngyuilun/EC-LMGraph

Operating system: Platform independent

Programming language: Python

Other requirements: Listed in the requirements.txt file provided in the repository

License: MIT license

RRID:SCR_027317

Additional files

Supplementary Fig. S1: Building protein graph representation and dependence of prediction performance on distance cut-off. (A) Protein graphs obtained with different distance cut-offs for the same example shown in Fig. 1A. (B) Dependence of EC-LMGraph performance (Precision, Recall, F1) and training time on distance cut-offs ranging from 6 to 13.

Supplementary Fig. S2: (A) Enzyme commission number prediction performance of different graph convolutional layers. (B) Number of M-CSA catalytic residues covered by residues with top 5% or 10% activation values using different graph convolutional layers.

Supplementary Fig. S3: (A) Prediction performance of EC-LMGraph optimized by binary cross-entropy loss compared to Focal Loss with varying focusing parameters (Inline graphic). (B) Comparison of the prediction performance of EC-LMGraph, fully connected network, and convolutional neural network on the learning of the protein language model. The performance of EC-LMGraph trained with ProtBert, with ESM-2, and without the protein language model is also included for comparison. (C) Accuracy, specificity, and Matthews correlation coefficient (MCC) of EC-LMGraph for the prediction of enzyme commission (EC) main class, subclass, sub-subclass, and sub-sub-subclass numbers. (D) Computation time of the various compared methods. The prediction framework was used to predict EC numbers for 100, 200, 500, and 1,000 randomly selected proteins.

Supplementary Fig. S4: Total number of M-CSA catalytic residues covered by residues with top 5% or 10% activation values using different explainability methods with EC-LMGraph.

Supplementary Fig. S5: (A) The ratio of the number of correctly predicted sites to the total number of predicted catalytic residues, evaluated across a window size range of ±1 to ±5. (B) Total number of catalytic residues annotated in the Mechanism and Catalytic Site Atlas (M-CSA) that are covered within the window sizes from 3 Å to 7 Å.

Supplementary Fig. S6: Performance of EC-LMGraph predictions relative to AlphaFold2 structural confidence (pLDDT). The performance was evaluated for High (>90), Medium (70–90), and Low (<70) pLDDT groups across all EC hierarchy levels. Error bars represent 95% confidence intervals derived from bootstrap resampling.

Supplementary Fig. S7: Catalytic sites prediction for Parkinson’s disease-related enzymes using DeepFRI–Grad-CAM. (A) Ubiquitin carboxyl-terminal hydrolase isozyme L1, (B) Serine protease HTRA2, (C) Protein deglycase DJ-1, and (D) Parkin.

Supplementary Table S1: Stratified evaluation of EC-LMGraph prediction performance (F1 score) across major enzyme commission (EC) classes using distance cutoffs ranging from 6 Å to 13 Å.

Supplementary Table S2: Distribution of training samples across the four hierarchical levels of the EC classification.

Supplementary Table S3: Performance comparison of catalytic site prediction between EC-LMGraph and DeepFRI.

Supplementary Table S4: EC-LMGraph prediction performance on AlphaFold2-predicted structures.

Competing interests

The authors declare no competing interests.

Supplementary Material

giag056_Supplemental_File
giag056_Authors_Response_To_Reviewer_Comments_original_submission
giag056_GIGA-D-25-00310_original_submission
giag056_GIGA-D-25-00310_revision_1
giag056_Reviewer_1_Report_original_submission

Reviewer 1 -- 9/20/2025

giag056_Reviewer_2_Report_original_submission

Reviewer 2 -- 11/25/2025

giag056_Reviewer_2_Report_revision_1

Reviewer 2 -- 4/5/2026

Acknowledgments

We acknowledge funding support from the Research Grants Council (14212525, 17204124, 17209021, 17210023, 14100122, C6027-19GF & C7074-21GF, C4026-21G, STG1/E-401/23-N, AoE/E-407/24-N, AoE/M-604/16) of the University Grants Committee of Hong Kong (H.K., K.W.K.); a Croucher Innovation Award (CIA20CU01) from the Croucher Foundation (H.K.); the Health and Medical Research Fund (21200872) from the Food and Health Bureau of Hong Kong (B.W.-L.N.); the Gerald Choa Neuroscience Institute (GCNI) Research Fund (B.W.-L.N.); the Excellent Young Scientists Fund from the National Natural Science Foundation of China (H.K.); the Lo’s Family Charity Fund Limited (H.K.); the Multi-scale Medical Robotics Center Limited (K.W.K.). Correspondence and requests for materials should be addressed to Ho Ko or Ka-Wai Kwok.

Contributor Information

Yui-Lun Ng, Department of Mechanical Engineering, Faculty of Engineering, The University of Hong Kong, Pok Fu Lam, Hong Kong.

Xiaomei Wang, Department of Mechanical Engineering, Faculty of Engineering, The University of Hong Kong, Pok Fu Lam, Hong Kong.

Yingqi Li, Department of Mechanical Engineering, Faculty of Engineering, The University of Hong Kong, Pok Fu Lam, Hong Kong.

Junzhe Huang, Division of Neurology, Department of Medicine and Therapeutics, Faculty of Medicine, The Chinese University of Hong Kong, Shatin, Hong Kong; Li Ka Shing Institute of Health Sciences, Faculty of Medicine, The Chinese University of Hong Kong, Shatin, Hong Kong; Margaret K. L. Cheung Research Centre for Management of Parkinsonism, Faculty of Medicine, The Chinese University of Hong Kong, Shatin, Hong Kong; Gerald Choa Neuroscience Institute, The Chinese University of Hong Kong, Shatin, Hong Kong.

Hei Ming Lai, Division of Neurology, Department of Medicine and Therapeutics, Faculty of Medicine, The Chinese University of Hong Kong, Shatin, Hong Kong; Li Ka Shing Institute of Health Sciences, Faculty of Medicine, The Chinese University of Hong Kong, Shatin, Hong Kong; Margaret K. L. Cheung Research Centre for Management of Parkinsonism, Faculty of Medicine, The Chinese University of Hong Kong, Shatin, Hong Kong; Gerald Choa Neuroscience Institute, The Chinese University of Hong Kong, Shatin, Hong Kong; Department of Psychiatry, Faculty of Medicine, The Chinese University of Hong Kong, Shatin, Hong Kong.

Jason Ying-Kuen Chan, Department of Otorhinolaryngology, Head and Neck Surgery, Faculty of Medicine, The Chinese University of Hong Kong, Shatin, Hong Kong.

Billy Wai-Lung Ng, Li Ka Shing Institute of Health Sciences, Faculty of Medicine, The Chinese University of Hong Kong, Shatin, Hong Kong; Gerald Choa Neuroscience Institute, The Chinese University of Hong Kong, Shatin, Hong Kong; Guangdong-Hong Kong-Macao Joint Laboratory for New Drug Screening, School of Pharmacy, Faculty of Medicine, The Chinese University of Hong Kong, Shatin, Hong Kong.

Ho Ko, Division of Neurology, Department of Medicine and Therapeutics, Faculty of Medicine, The Chinese University of Hong Kong, Shatin, Hong Kong; Li Ka Shing Institute of Health Sciences, Faculty of Medicine, The Chinese University of Hong Kong, Shatin, Hong Kong; Margaret K. L. Cheung Research Centre for Management of Parkinsonism, Faculty of Medicine, The Chinese University of Hong Kong, Shatin, Hong Kong; Gerald Choa Neuroscience Institute, The Chinese University of Hong Kong, Shatin, Hong Kong; Department of Psychiatry, Faculty of Medicine, The Chinese University of Hong Kong, Shatin, Hong Kong.

Ka-Wai Kwok, Department of Mechanical Engineering, Faculty of Engineering, The University of Hong Kong, Pok Fu Lam, Hong Kong; Department of Mechanical and Automation Engineering, Faculty of Engineering, The Chinese University of Hong Kong, Shatin, Hong Kong.

Author contributions

Yui-Lun Ng (Conceptualization [lead], Data curation [lead], Formal analysis [lead], Investigation [lead], Methodology [lead], Software [lead], Validation [lead], Visualization [lead], Writing – original draft [lead], Writing – review & editing [lead]), Xiaomei Wang (Investigation [supporting], Visualization [supporting], Writing – original draft [supporting]), Yingqi Li (Investigation [supporting], Writing – review & editing [supporting]), Junzhe Huang (Investigation [supporting], Writing – review & editing [supporting]), Hei Ming Lai (Writing – review & editing [supporting]), Jason Ying-Kuen Chan (Writing – review & editing [supporting]), Billy Wai-Lung Ng (Funding acquisition [supporting], Methodology [supporting], Writing – review & editing [supporting]), Ho Ko (Conceptualization [lead], Formal analysis [lead], Funding acquisition [lead], Investigation [lead], Methodology [lead], Project administration [lead], Resources [lead], Writing – original draft [lead], Writing – review & editing [lead]), and Ka-Wai Kwok (Conceptualization [lead], Formal analysis [lead], Funding acquisition [lead], Investigation [lead], Methodology [lead], Project administration [lead], Resources [lead], Writing – original draft [lead], Writing – review & editing [lead])

Data availability

The protein structure data and EC number annotations described in this manuscript were downloaded from wwPDB [75] and UniProt [76], respectively. The enzyme catalytic residues are based on annotations documented in M-CSA [77]. The AlphaFold2-predicted structures were downloaded from AlphaFold DB [78]. The post-processed training and test datasets, source code, models with their network weight, as well as all the entries of prediction results are freely available in the EC-LMGraph repository [28].

References

  • 1. UniProt Consortium . UniProt: a worldwide hub of protein knowledge. Nucleic Acids Res. 2019;47:D506–D515. 10.1093/nar/gky1049. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 2. MacDougall  A, Volynkin  V, Saidi  R, et al. UniRule: a unified rule resource for automatic annotation in the UniProt Knowledgebase. Bioinformatics. 2020;36:4643–4648. 10.1093/bioinformatics/btaa485. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 3. Altschul  S F, Gish  W, Miller  W, Myers  E W, et al. Basic local alignment search tool. J Mol Biol. 1990;215:403–410. 10.1016/S0022-2836(05)80360-2. [DOI] [PubMed] [Google Scholar]
  • 4. Yu  T, Cui  H, Li  J C, Luo  Y, et al. Enzyme function prediction using contrastive learning. Science. 2023;379:1358–1363. 10.1126/science.adf2465. [DOI] [PubMed] [Google Scholar]
  • 5. Yang  Y, Jerger  A, Feng  S, Wang  Z, et al. Improved enzyme functional annotation prediction using contrastive learning with structural inference. Commun Biol. 2024;7:1690. 10.1038/s42003-024-07359-z. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6. Kim  G B, Kim  J Y, Lee  J A, et al. Functional annotation of enzyme-encoding genes using deep learning with transformer layers. Nat Commun. 2023;14:7370. 10.1038/s41467-023-43216-z. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 7. Sanderson  T, Bileschi  M L, Belanger  D, Colwell  L J. ProteInfer, deep neural networks for protein functional inference. eLife. 2023;12:e80942. 10.7554/eLife.80942. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 8. Dalkiran  A, Rifaioglu  A S, Martin  M J, et al. ECPred: a tool for the prediction of the enzymatic functions of protein sequences based on the EC nomenclature. BMC Bioinf. 2018;19:334. 10.1186/s12859-018-2368-y. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 9. Ryu  J Y, Kim  H U, Lee  S Y. Deep learning enables high-quality and high-throughput prediction of enzyme commission numbers. Proc Natl Acad Sci USA. 2019;116:13996–14001. 10.1073/pnas.1821905116. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 10. Song  Y, Yuan  Q, Chen  S, Zhao  H, et al. Accurately predicting enzyme functions through geometric graph learning on ESMFold-predicted structures. Nat Commun. 2024;15:8180. 10.1038/s41467-024-52533-w. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 11. Han  S R, Park  M, Kosaraju  S, et al. Evidential deep learning for trustworthy prediction of enzyme commission number. Brief Bioinform. 2024;25:1–11. 10.1093/bib/bbad401. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12. Śledź  P, Caflisch  A. Protein structure-based drug design: from docking to molecular dynamics. Curr Opin Struct Biol. 2018;48:93–102. 10.1016/j.sbi.2017.10.010. [DOI] [PubMed] [Google Scholar]
  • 13. Ferreira  L, dosSantos  R, Oliva  G, et al. Molecular docking and structure-based drug design strategies. Molecules. 2015;20:13384–13421. 10.3390/molecules200713384. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 14. Meng  X-Y, Zhang  H-X, Mezei  M, et al. Molecular docking: a powerful approach for structure-based drug discovery. Curr Comput Aided-Drug Des. 2011;7:146–157. 10.2174/157340911795677602. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15. Zhang  C, Freddolino  P L, Zhang  Y. COFACTOR: improved protein function prediction by combining structure, sequence and protein–protein interaction information. Nucleic Acids Res. 2017;45:W291–W299. 10.1093/nar/gkx366. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16. Gligorijević  V, Renfrew  P D, Kosciolek  T, et al. Structure-based protein function prediction using graph convolutional networks. Nat Commun. 2021;12: 3168. 10.1038/s41467-021-23303-9. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 17. Smaili  F Z, Tian  S, Roy  A, et al. QAUST: protein function prediction using structure similarity, protein interaction, and functional motifs. Genomics, Proteomics Bioinforma. 2021;19:998–1011. 10.1016/j.gpb.2021.02.001. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 18. Liang  M, Nie  J. Prediction of enzyme function based on a structure relation network. IEEE Access. 2020;8:132360–132366. 10.1109/ACCESS.2020.3010028. [DOI] [Google Scholar]
  • 19. Zhou  J, Cui  G, Hu  S, et al. Graph neural networks: a review of methods and applications. AI Open. 2020;1:57–81. 10.1016/j.aiopen.2021.01.001. [DOI] [Google Scholar]
  • 20. Kipf  T N, Welling  M. Semi-supervised classification with graph convolutional networks. 5th Int Conf Learn Represent ICLR 2017–Conf Track Proc. 2017;1–14. Toulon, France: OpenReview.net. [Google Scholar]
  • 21. Elnaggar  A, Heinzinger  M, Dallago  C, et al. ProtTrans: toward understanding the language of life through self-supervised learning. IEEE Trans Pattern Anal Mach Intell. 2022;44:7112–7127. 10.1109/TPAMI.2021.3095381. [DOI] [PubMed] [Google Scholar]
  • 22. Lin  Z, Akin  H, Rao  R, et al. Evolutionary-scale prediction of atomic-level protein structure with a language model. Science. 2023;379:1123–1130. 10.1126/science.ade2574. [DOI] [PubMed] [Google Scholar]
  • 23. Vaswani A, Shazeer  N, Parmar  N, et al. Attention is all you need. Proc 31st Int Conf Neural Inf Process Syst. Curran Associates Inc: Red Hook, NY, 2017. [Google Scholar]
  • 24. Altunkaya  A, Bi  C, Bradley  A R, et al. The RCSB protein data bank: integrative view of protein, gene and 3D structural information. Nucleic Acids Res. 2016;45:D271–D281. 10.1093/nar/gkw1000. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 25. Ribeiro  AJM, Holliday  G L, Furnham  N, et al. Mechanism and Catalytic Site Atlas (M-CSA): a database of enzyme reaction mechanisms and active sites. Nucleic Acids Res. 2018;46:D618–D623. 10.1093/nar/gkx1012. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 26. Porter  C T. The Catalytic Site Atlas: a resource of catalytic sites and residues identified in enzymes using structural data. Nucleic Acids Res. 2004;32:129D–133. 10.1093/nar/gkh028. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 27. Furnham  N, Holliday  G L, deBeer  TAP, et al. The Catalytic Site Atlas 2.0: cataloging catalytic sites and residues identified in enzymes. Nucl Acids Res. 2014;42:D485–D489. 10.1093/nar/gkt1243. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 28. EC-LMGraph repository . https://github.com/ngyuilun/EC-LMGraph  Accessed 28 April 2026.
  • 29. Brody  S, Alon  U, Yahav  E. How attentive are graph attention networks?. ICLR 2022–10th Int Conf Learn Represent. Virtual Event, Amherst, MA: OpenReview.net, 2022;1–26. [Google Scholar]
  • 30. Bai  S, Zhang  F, Torr  PHS. Hypergraph convolution and hypergraph attention. Pattern Recognit. 2021;110:107637. 10.1016/j.patcog.2020.107637. [DOI] [Google Scholar]
  • 31. Chiang  W L, Li  Y, Liu  X, et al. Cluster-GCN: an efficient algorithm for training deep and large graph convolutional networks. Proc ACM SIGKDD Int Conf Knowl Discov Data Min. Anchorage, AK. New York: ACM, 2019; 10.1145/3292500.3330925. [DOI] [Google Scholar]
  • 32. Ranjan  E, Sanyal  S, Talukdar  P. ASAP: adaptive structure aware pooling for learning hierarchical graph representations. AAAI 2020–34th AAAI Conf Artif Intell. New York, NY. Palo Alto, CA:AAAI Press, 2020. 10.1609/aaai.v34i04.5997. [DOI] [Google Scholar]
  • 33. Li  Y, Huang  C, Ding  L, Li  Z, et al. Deep learning in bioinformatics: introduction, application, and perspective in the big data era. Methods. 2019;166:4–21. 10.1016/j.ymeth.2019.04.008. [DOI] [PubMed] [Google Scholar]
  • 34. Bonetta  R, Valentino  G. Machine learning techniques for protein function prediction. Proteins. 2020;88:397–413. 10.1002/prot.25832. [DOI] [PubMed] [Google Scholar]
  • 35. Johnson  J M, Khoshgoftaar  T M. Survey on deep learning with class imbalance. J Big Data. 2019;6:1–54. 10.1186/s40537-019-0192-5. [DOI] [Google Scholar]
  • 36. Lin  T-Y, Goyal  P, Girshick  R, et al. Focal loss for dense object detection. IEEE Trans Pattern Anal Mach Intell. 2020;42:318–327. 10.1109/TPAMI.2018.2858826. [DOI] [PubMed] [Google Scholar]
  • 37. Brandes  N, Ofer  D, Peleg  Y, et al. ProteinBERT: a universal deep-learning model of protein sequence and function. Bioinformatics. 2022;38:2102–2110. 10.1093/bioinformatics/btac020. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 38. Fu  L, Niu  B, Zhu  Z, et al. CD-HIT: accelerated for clustering the next-generation sequencing data. Bioinformatics. 2012;28:3150–3152. 10.1093/bioinformatics/bts565. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 39. Sechidis  K, Tsoumakas  G, Vlahavas  I. On the Stratification of Multi-label Data. Machine Learning and Knowledge Discovery in Databases. ECML PKDD 2011, Athens, Greece. Berlin, Heidelberg: Springer, 2011;145–158. [Google Scholar]
  • 40. Simonyan  K, Vedaldi  A, Zisserman  A. Deep inside convolutional networks: visualising image classification models and saliency maps. 2nd Int Conf Learn Represent ICLR 2014–Work Track Proc. Banff, Canada. Amherst, MA: OpenReview.net, 2014;1–8. [Google Scholar]
  • 41. Shrikumar  A, Greenside  P, Kundaje  A. Learning important features through propagating activation differences. 34th Int Conf Mach Learn ICML. Sydney, Australia. Brookline, MA, 2017;7:4844–662017. [Google Scholar]
  • 42. Springenberg  J T, Dosovitskiy  A, Brox  T, et al. Striving for simplicity: the all convolutional net. 3rd Int Conf Learn Represent ICLR 2015–Work Track Proc. San Diego, CA. Amherst, MA: OpenReview.net, 2015;1–14. [Google Scholar]
  • 43. Zeiler  M D, Fergus  R. Visualizing and understanding convolutional networks. ECCV 2014. Cham, Switzerland: Springer, 2014;818–833. [Google Scholar]
  • 44. Goodsell  D S, Zardecki  C, DiCostanzo  L, et al. RCSB Protein Data Bank: enabling biomedical research and drug discovery. Protein Sci. 2020;29:52–65. 10.1002/pro.3730. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 45. Varadi  M, Anyango  S, Deshpande  M, et al. AlphaFold Protein Structure Database: massively expanding the structural coverage of protein-sequence space with high-accuracy models. Nucleic Acids Res. 2022;50:D439–D444. 10.1093/nar/gkab1061. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 46. Jumper  J, Evans  R, Pritzel  A, et al. Highly accurate protein structure prediction with AlphaFold. Nature. 2021;596:583–589. 10.1038/s41586-021-03819-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 47. Baek  M, DiMaio  F, Anishchenko  I, et al. Accurate prediction of protein structures and interactions using a three-track neural network. Science. 2021;373:871–876. 10.1126/science.abj8754. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 48. Humphreys  I, Pei  J, Baek  M, et al. Computed structures of core eukaryotic protein complexes. Science. 2021;374:eabm4805. 10.1126/science.abm4805. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 49. Guadagnolo  D, Piane  M, Torrisi  M R, et al. Genotype-phenotype correlations in monogenic Parkinson disease: a review on clinical and molecular findings. Front Neurol. 2021;12:648588. 10.3389/fneur.2021.648588. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 50. Day  J O, Mullin  S. The genetics of parkinson’s disease and implications for clinical practice. Genes. 2021;12:1006. 10.3390/genes12071006. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 51. Lesage  S, Brice  A. Parkinson’s disease: from monogenic forms to genetic susceptibility factors. Hum Mol Genet. 2009;18:R48–R59. 10.1093/hmg/ddp012. [DOI] [PubMed] [Google Scholar]
  • 52. Altschul  S F, Madden  T L, Schäffer  A A, et al. Gapped BLAST and PSI-BLAST: a new generation of protein database search programs. Nucleic Acids Res. 1997;25:3389–3402. 10.1093/nar/25.17.3389. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 53. Kumar  N, Skolnick  J. EFICAz2.5: application of a high-precision enzyme function predictor to 396 proteomes. Bioinformatics. 2012;28:2687–2688. 10.1093/bioinformatics/bts510. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 54. Fu  Y, Gu  Z, Luo  X, et al. Learning a generalized graph transformer for protein function prediction in dissimilar sequences. Gigascience. 2024;13:giae093. 10.1093/gigascience/giae093. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 55. Jha  K, Saha  S, Singh  H. Prediction of protein–protein interaction using graph neural networks. Sci Rep. 2022;12:8360. 10.1038/s41598-022-12201-9. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 56. Du  X, Sun  S, Hu  C, et al. DeepPPI: boosting prediction of protein–Protein interactions with deep neural networks. J Chem Inf Model. 2017;57:1499–1510. 10.1021/acs.jcim.7b00028. [DOI] [PubMed] [Google Scholar]
  • 57. Sun  T, Zhou  B, Lai  L, et al. Sequence-based prediction of protein–protein interaction using a deep-learning algorithm. BMC Bioinf. 2017;18: 277. 10.1186/s12859-017-1700-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 58. Han  Y, Zhang  S W, Zhang  Q Q, et al. MGMA-PPIS: predicting the protein–protein interaction site with multiview graph embedding and multiscale attention fusion. Gigascience. 2025;14:giaf114. 10.1093/gigascience/giaf114. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 59. Tubiana  J, Schneidman-Duhovny  D, Wolfson  H J. ScanNet: an interpretable geometric deep learning model for structure-based protein binding site prediction. Nat Methods. 2022;19:730–739. 10.1038/s41592-022-01490-7. [DOI] [PubMed] [Google Scholar]
  • 60. Pan  X, Fang  Y, Li  X, et al. RBPsuite: rNA-protein binding sites prediction suite based on deep learning. Bmc Genomics [Electronic Resource]. 2020;21:884. 10.1186/s12864-020-07291-6. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 61. Zhang  J, Ghadermarzi  S, Katuwalala  A, et al. DNAgenie: accurate prediction of DNA-type-specific binding residues in protein sequences. Brief Bioinform. 2021;22:bbab336. 10.1093/bib/bbab336. [DOI] [PubMed] [Google Scholar]
  • 62. Zhang  F, Zhao  B, Shi W, et al. DeepDISOBind: accurate prediction of RNA-, DNA- and protein-binding intrinsically disordered residues with deep multi-task learning. Brief Bioinform. 2022;23:bbab521. 10.1093/bib/bbab521. [DOI] [PubMed] [Google Scholar]
  • 63. Zhang  S, Han  J, Liu  J. Protein–protein and protein–nucleic acid binding site prediction via interpretable hierarchical geometric deep learning. Gigascience. 2024;13:giae080. 10.1093/gigascience/giae080. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 64. Ma  W, Zheng  W, Qin  S, et al. DeepAnnotation: a novel interpretable deep learning-based genomic selection model that integrates comprehensive functional annotations. Gigascience. 2025;14:giaf083. 10.1093/gigascience/giaf083. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 65. Yu  H, Deng  H, He  J, et al. UniKP: a unified framework for the prediction of enzyme kinetic parameters. Nat Commun. 2023;14:8210. 10.1038/s41467-023-44113-1. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 66. Wang  Y, Cheng  L, Zhang  Y, et al. DEKP: a deep learning model for enzyme kinetic parameter prediction based on pretrained models and graph neural networks. Brief Bioinform. 2025;26:bbaf187. 10.1093/bib/bbaf187. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 67. Jing  X, Dong  Q, Hong  D, et al. Amino acid encoding methods for protein sequences: a comprehensive review and assessment. IEEE/ACM Trans Comput Biol and Bioinf. 2020;17:1918–1931. 10.1109/TCBB.2019.2911677. [DOI] [PubMed] [Google Scholar]
  • 68. Fout  A, Byrd  J, Shariat  B, et al. Protein interface prediction using graph convolutional networks. Adv Neural Inf Process Syst. 2017;6530–6539. 10.5555/3295222.3295399. [DOI] [Google Scholar]
  • 69. Shao  K, Zhang  Z, He  S, et al. DTIGCCN: prediction of drug-target interactions based on GCN and CNN. 2020 IEEE 32nd Int Conf Tools with Artif Intell. IEEE, 2020. [Google Scholar]
  • 70. Lin  S, Jia  P. scGraph2Vec: a deep generative model for gene embedding augmented by graph neural network and single-cell omics data. Gigascience. 2024;13:giae108. 10.1093/gigascience/giae108. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 71. Xu  B, Wang  N, Chen  T, Li  M. Empirical evaluation of rectified activations in convolutional network. arXiv 150500853, 2015.
  • 72. Kingma  D P, Ba  J L. Adam: a method for stochastic optimization. 3rd Int Conf Learn Represent ICLR 2015–Conf Track Proc., 2015, 1–15.
  • 73. Fey  M, Lenssen  J E. Fast graph representation learning with PyTorch geometric. ICLR 2019 Work Represent Learn Graphs Manifolds, San Diego, CA. Amherst, MA: OpenReview.net, 2015;1–9. [Google Scholar]
  • 74. Pettersen  E F, Goddard  T D, Huang  C C, et al. UCSF Chimera—a visualization system for exploratory research and analysis. J Comput Chem. 2004;25:1605–1612. 10.1002/jcc.20084. [DOI] [PubMed] [Google Scholar]
  • 75. wwPDB . https://www.wwpdb.org/ftp/pdb-ftp-sites. Accessed 28 April 2026.
  • 76. UniProt Database . https://www.uniprot.org/help/downloads. Accessed 28 April 2026.
  • 77. M-CSA Database . https://www.ebi.ac.uk/thornton-srv/m-csa/download. Accessed 28 April 2026.
  • 78. AlphaFold DB . https://alphafold.ebi.ac.uk/download. Accessed 28 April 2026.

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

giag056_Supplemental_File
giag056_Authors_Response_To_Reviewer_Comments_original_submission
giag056_GIGA-D-25-00310_original_submission
giag056_GIGA-D-25-00310_revision_1
giag056_Reviewer_1_Report_original_submission

Reviewer 1 -- 9/20/2025

giag056_Reviewer_2_Report_original_submission

Reviewer 2 -- 11/25/2025

giag056_Reviewer_2_Report_revision_1

Reviewer 2 -- 4/5/2026

Data Availability Statement

The protein structure data and EC number annotations described in this manuscript were downloaded from wwPDB [75] and UniProt [76], respectively. The enzyme catalytic residues are based on annotations documented in M-CSA [77]. The AlphaFold2-predicted structures were downloaded from AlphaFold DB [78]. The post-processed training and test datasets, source code, models with their network weight, as well as all the entries of prediction results are freely available in the EC-LMGraph repository [28].


Articles from GigaScience are provided here courtesy of Oxford University Press

RESOURCES