Abstract
Thirty to seventy percent of proteins in any given genome have no assigned function and have been labeled as the protein “unknome.” This large knowledge shortfall is one of the final frontiers of biology. Machine learning (ML) approaches are enticing, with early successes demonstrating the ability to propagate functional knowledge from experimentally characterized proteins. An open question is the ability of ML approaches to predict enzymatic functions unseen in the training sets. By integrating literature and a combination of bioinformatic approaches, we evaluated individually Enzyme Commission number predictions for over 450 Escherichia coli unknowns made using state-of-the-art ML approaches. We found that current ML methods not only mostly fail to make novel predictions but also make basic logic errors in their predictions that human annotators avoid by leveraging the available knowledge base. This underscores the need to include assessments of prediction uncertainty in model output and to test for “hallucinations” (logic failures) as a part of model evaluation. Explainable artificial intelligence analysis can be used to identify indicators of prediction errors, potentially identifying the most relevant data to include in the next generation of computational models.
Keywords: paralogs, orthologs, YciO, protein language models, Enzyme Commission, protein family, functional annotation
Many proteins in any genome, ranging from 30% to 70% of the genome, lack an assigned function. This knowledge gap limits the full use of the vast available genomic data. Machine learning has shown promise in transferring functional knowledge within isofunctional families, but it largely fails to predict novel functions not seen in its training data. Understanding these failures can guide the development of better machine-learning methods to help experts make accurate functional predictions for uncharacterized proteins.
Introduction
Determining protein function is not an easy task, and 30 yr after the first bacterial genome was sequenced, the functional annotation status of the proteome of most species is far from being accurate or complete, even for model organisms (Ghatak et al. 2019; Wood et al. 2019; de Crécy-lagard et al. 2022; Rocha et al. 2023). Experimental validation of protein function is a painstaking process, and with the explosion of whole genome sequences (Kyrpides 1999; Peterson et al. 2025; Torres et al. 2025), the gap between experimentally validated functions and those predicted through computational methods continues to widen. In UniProtKB (Bateman et al. 2023), the most widely used protein function database (Ramola et al. 2022), the estimates are that less than 0.5% to 15% of proteins have been linked to experimental data (Škunca et al. 2017).
The process of functional annotation of protein entries in databases starts with capturing information in the literature by biocurators (International Society for Biocuration 2018). This process links experimental characterizations of specific proteins in specific organisms to controlled vocabularies that describe validated functions, such as the Gene Ontology (GO) (Gene Ontology Consortium et al. 2023), the IUPAC Enzyme Commission (EC) numbers (https://iubmb.qmul.ac.uk/enzyme/), or biochemical reaction descriptors (e.g. Rhea (Bansal et al. 2022)). Text mining tools have accelerated the flow of information captured (Soldatos et al. 2015; Poux et al. 2017; Wei et al. 2019). However, this step remains a major bottleneck in the annotation workflow, which can result in mislabeling proteins as “unknown” when a function has been reported in the literature (see type 1 error in Table 1 and Fig. 1).
Table 1.
Main types of errors that lead to erroneous functional annotations of proteins.
| Error types | ||
|---|---|---|
| False unknowns | Annotated as unknown or general but function is known and published | |
| Details | Example | |
| 1 Failure to capture the literature | Not captured in any database or captured in some databases but not others Also annotated as vague when precise annotation is known | CT_611 is captured as folylpolyglutamate synthase in KEGG (ctr:CT_611), but this annotation is still not in UniProt as of April 2024 (O84617). See also examples in Price and Arkin (2024) |
| 2 Naming issues | Inconsistent naming of the same entities | See GroEL in Lockwood et al. (2019) |
| Propagation failures | Captured for one member of a family but not propagated to others | The MptE protein that encodes 6-hydroxymethyl-7,8-dihydropterin pyrophosphokinase in most Archaea (de Crécy-Lagard et al. 2012) is captured in BioCyc for Methanocaldococcus jannaschii DSM 2661 (https://biocyc.org/gene?orgid=MJ&id=MJ_RS08700-MONOMER) but not propagated to any archaeal homolog |
| Fusion and multidomain proteins | Only one out of multiple functions is captured or domain shuffling leads to miscalling | See examples in Henryt e al. (2016) and Hegyi and Gerstein (2001) |
| False knowns | Annotated as precise but wrong (should be general or another annotation) or incomplete | |
|---|---|---|
| Details | Example | |
| 3 Multiple functions | Protein has multiple functions because of fusions, moonlighting, or promiscuity, and only one of the functions is captured | For example, A5I019 has 2 functions QueD and PTPS-III, and only one is called in UniProt (de Crécy-Lagard 2014). See also examples in Price and Arkin (2024) |
| 4 Curation mistake | (1) The data were incorrectly captured by a biocurator; (2) functional annotations may become outdated | Ureidoglycolate lyase (Percudani et al. 2013) |
| 5 Experimental mistake | Finding has been refuted by other studies: (1) the published data are inconclusive; (2) inconsistency occurs when different databases or resources provide conflicting functional annotations for the same protein | DUF34 family was annotated as GTP cyclohydrolase 1B (Reed et al. 2021) |
| 6 Overannotation of paralogs | Annotation wrongly propagated to nonisofunctional paralogous groups | See examples in Schnoes et al. (2009), Zallot et al. (2016), and (Rembeza and Engqvist 2021) |
Error types 1–6 correspond to the errors numbered in Fig. 1. KEGG: Kyoto Encyclopedia of Genes and Genomes.
Fig. 1.
The major 2 types of challenges in accurate protein functional annotation. a) Challenge 1: to propagate the existing knowledge to the correct set of unannotated proteins. Challenge 2: to annotate proteins in a family with none of the members linked to any initial known information. Each circle represents a protein in a family with (bold border) or without (thin border) initially known function. The edges that connect 2 circles represent protein similarity above a preset threshold. The squares represent known functions with E.C.a to E.C.n as examples. b) The errors during propagation of annotation that could lead to erroneous results. The numbering of errors (red crosses) corresponds to the error types described in Table 1.
Using the principle of sequence similarity, also known as homology transfer, putative functions are assigned to proteins in newly sequenced genomes as a part of the genome annotation process (Seemann 2014; Thibaud-Nissen et al. 2016; Olson et al. 2023). Indeed, most functions of proteins in UniProt have been inferred computationally based on sequence similarity (Bateman et al. 2023). The methods used over the past 30 yr to automatically propagate functional annotations have been extensively reviewed (Storm and Sonnhammer 2002; Friedberg 2006; Lee et al. 2007; Satish Kumar et al. 2007; Devoid et al. 2013). This seemingly simple task can become quite difficult (challenge 1 in Fig. 1), since relying on sequence similarity alone can result in significant annotation errors (Schnoes et al. 2009; Rembeza and Engqvist 2021). These errors can arise from various sources, including human annotation mistakes or intrinsic features of the protein family itself, such as domain shuffling or domain fusions (Table 1). One prevalent error type affecting protein families (Rembeza and Engqvist 2021) is in the misannotation of paralogs (Schnoes et al. 2009; Zallot et al. 2016; Rembeza and Engqvist 2021) caused by the inherent functional complexity of enzyme families (Glasner et al. 2006; Gerlt et al. 2012) (see type 6 error in Table 1 and Fig. 1). Functional diversification through duplication and divergence (Todd et al. 2001; Das et al. 2015; Bordin et al. 2021) results in proteins with high degrees of similarity having different functional roles, creating nonisofunctional paralogous groups. Indeed, even within a single protein family, a difference in a few amino acids can affect substrate binding and/or catalysis, effectively changing the function (Roy et al. 2009; Ribeiro et al. 2020; Precord et al. 2023). Most current genome annotation pipelines annotate paralogs without considering the potential for functional divergence, leading to incorrect annotations for up to 80% of family members, with these types of errors increasing over time (Schnoes et al. 2009; Rembeza and Engqvist 2021).
Methods for inferring functions that incorporate additional information such as phylogeny, active site analyses, metabolic reconstruction, and sequence similarity networks (SSNs) combined with gene neighborhood or coexpression information can better separate nonisofunctional subfamilies (Zallot et al. 2016, 2021; Ribeiro et al. 2023) (Fig. 2). When hidden Markov models (HMMs) or signature motifs are generated to annotate the subfamilies, they can be integrated into annotation pipelines such as RefSeq (Li et al. 2021) or incorporated into rules such as the Unified Rule used by UniProtKB (MacDougall et al. 2020) additional curation, which has not kept pace with the exponential increase in sequenced genomes. In addition to misassignments among paralogs increasing with evolutionary distance, distinguishing between sub- and neofunctionalization can be problematic (Birchler 2025).
Fig. 2.
Computational workflow used to separate nonisofunctional paralogs in a protein superfamily. SSNs of protein families can separate paralogs by integrating different types of information, including genomic neighborhood context, structure, cooccurring genes, multiomics data, etc. Arrows indicate the flow of information. The final network shows the separation and reannotation of paralogous subclusters.
The task of functional annotation becomes even more challenging when one focuses on proteins that have not previously been functionally characterized. Estimates suggest that 30% to 70% of proteins in any given genome are in this “unknown” category (Ghatak et al. 2019; Wood et al. 2019; Lobb et al. 2020; Sajid et al. 2024). Some of these “unknowns” are likely to fill functional roles not yet linked to a specific gene (Karp 2004; Lespinet and Labedan 2006; Chen and Vitkup 2007; Niehaus et al. 2015). Others are likely to be nonorthologous replacements (or convergent evolution), where different (structurally diverse) protein families overlap in functional roles (Ribeiro et al. 2023). Finally, we expect some novel biological functions and predicting these requires extrapolation beyond the current knowledge base.
In summary, the annotation of a new genome faces 2 challenges: (i) to identify genes whose likely function can be inferred by propagating known annotations and (ii) to identify genes with potentially novel functions and make predictions. This is a thorny path with many different types of errors that can be made (Fig. 1b and Table 1). Computational approaches, however innovative, cannot prove a particular protein function; experimental validation is still necessary. Proof of function is not necessarily straightforward, as demonstrating a protein capable of participating in a chemical reaction is not proof that the protein has the responsibility of that function in the cellular context where it is identified (García-Contreras et al. 2012 ; Copley 2015). For proteins with available experimental evidence of function in 1 context, the computational propagation of the assigned functional role is the logical next step (Lee et al. 2007), although as described above, similarity does not guarantee functional conservation. For unknown proteins, the increasing amounts of genome-wide/transcriptome-wide/metabolome-wide experimental data can be used in combination with structural knowledge to develop predictions that can be tested experimentally (Hanson et al. 2010). Scientists are increasingly successful in these types of endeavors in a range of systems (Dembech et al. 2023; Rodríguez del Río et al. 2024), and there is tremendous hope that artificial intelligence (AI) methods will accelerate the pace at which predictions can be made.
For challenge 1, there has been a recent explosion of publications reporting the use of pretrained protein language models (PLMs) to predict protein functions (Ardern et al. 2023; Ayres et al. 2023; Derry et al. 2025; Durairaj et al. 2023; Kim et al. 2023; Prabakaran and Bromberg 2024; Rodríguez del Río et al. 2024; Hwang et al. 2024). Models have been developed to link protein sequences directly to vocabularies such as GO terms (Radivojac et al. 2013; Jiang et al. 2016; Zhou et al. 2019) or EC (EC classification) numbers (Hamamsy et al. 2023). EC numbers are a set of 4 numbers that are hierarchical with the first number the most general classification and the last number the most specific and whose purpose is to standardize the description of enzymatic activities. PLMs have been reported to accurately predict the first 2 digits of the EC number but not the last 2 (Sanderson et al. 2023). GloEC (Huang et al. 2024) and MAPred (Rong et al. 2024) are 2 recent PLM-based EC prediction tools that achieve precision and recall metrics ranging from 10% to 80%, depending on the test set (Fig. 3). When training data contain well-characterized proteins with well-described structures (e.g. cofactor-237 set) or well-separated functional groups (e.g. phosphorylase set), machine learning (ML) methods are quite proficient (Fig. 3a). Once the structure of the training data is more complex, and the algorithm must distinguish among similar enzymes with different functions, such as the carbohydrate dataset that comprised ∼350 glycosyl hydrolases or if the dataset contains enzymes not included in any training sets such as the Price dataset (Price and Arkin 2024), ML methods do poorly (Fig. 3b).
Fig. 3.
Comparison of combined macro-F1 or regular F1 scores using different ML methods to predict EC numbers. The F1 score is the harmonic mean of precision and recall. Recall is the proportion of true positives correctly identified out of all actual positives, while precision is the proportion of true positives among all predicted positives. A “positive” refers to an instance that actually belongs to the class of interest (here, the correct EC number). The macro-F1 score is calculated by computing the F1 score for each EC number category independently and then averaging these scores across EC numbers. a) Comparison of macro-F1 scores generated using 4 methods (GloEC, Protein Infer, DeepEC, and CLEAN) on 4 datasets (Carbohydrate esterase, New-432, Cofactor-237, and phosphorylase) (data extracted from Huang et al. 2024). b) Comparison of F1 scores generated using 8 methods including BlastP, DeepECTF, DeepEC, and CLEAN on 3 datasets (New-815, New-392, and Price) (data extracted from Rong et al. 2024). Dotted lines mark the BlastP scores. See initial publications for details and references for methods and datasets.
Several independent studies agree that the number of proteins in the model Gram-negative Escherichia coli K12 not linked to precise molecular and/or biological functions is between 1,200 and 1,400, or a quarter of the encoded proteins (Ghatak et al. 2019). Some of these proteins are likely responsible for functions that have been described in E. coli but never linked to a gene (called orphan enzymes (Karp 2004; Danchin et al. 2018; Kobras et al. 2021)) or functions that might have been described in another organism but could also be present in E. coli. What percentage of these currently unannotated proteins represent novel functions is difficult to determine. A recent study labeled over 450 of the unknowns in the model bacteria E. coli with substrate-level specificity ECs (all 4 numbers) using a supervised ML-based DeepECTransformer (DeepECTF) platform and validated 3 of these predictions in vitro (Kim et al. 2023). If corroborated, this report could be a real breakthrough in predictive modeling by linking true unknowns with their function, a process that usually takes years (Hanson et al. 2010; Niehaus et al. 2015). We designed a study to compare human curation to the reported DeepECTF predictions and examined in detail the potential of this new approach.
Methods
Bioinformatic analyses and data mining used to evaluate EC number prediction
The UniProt (www.uniprot.org) (Bateman et al. 2023), InterPro (www.ebi.ac.uk/interpro/) (Paysan-Lafosse et al. 2023), NCBI (www.ncbi.nlm.nih.gov) (NCBI Resource Coordinators 2016), BioCyc (www.biocyc.org) (Karp et al. 2019), and KEGG (Kyoto Encyclopedia of Genes and Genomes) (Kanehisa et al. 2021) knowledge bases were used to gather information and evaluate individually the 453 E. coli proteins of unknown function annotated with an EC number using the DeepECTF platform (Kim et al. 2023). In addition, PaperBLAST was used to query literature on each entry (https://papers.genomics.lbl.gov/cgi-bin/litSearch.cgi) (Price and Arkin 2017).
Comparing human curation with DeepECTF predictions
Based on the evidence gathered manually, the DeepECTF predictions for the 453 target unknowns were grouped into different categories with confidence scores (CSs) (Supplementary Table 1). Concordant predictions for all 4 EC numbers were given a CS of 2. These were of 2 types: proteins with the identical functional annotation already present in UniProt (with or without an EC number) and proteins with functions not captured by UniProt. Discordant predictions were cases where (i) another protein was known to perform the same function and experimental data showed that there were no redundancies; (ii) the predicted function was known to be absent in E. coli; and (iii) the published literature or comments in UniProt provided evidence for a different function were given a CS of 0. When proteins were members of families with many paralogous subgroups but no literature disproving the prediction was found and in other cases of uncertain calls, a CS of 1 was given.
Structural and phylogenetic analyses
Crystal structures were retrieved from the Protein Data Bank (www.rcsb.org) (Burley et al. 2023) and visualized using PyMOL (www.pymol.org) (The PyMOL Molecular Graphics System, Version 1.8, Schrödinger, LLC). Structure-based multisequence alignment derived from available crystal structures and AlphaFold structural models (https://alphafold.ebi.ac.uk/) were generated using PROMALS3D (http://prodata.swmed.edu/promals3d/promals3d.php) (Pei et al. 2008) and ESPript (espript.ibcp.fr/ESPript/ESPript). The percent conservation scores were calculated using the program AL2CO (http://prodata.swmed.edu/al2co/al2co.php) (Pei and Grishin 2001 ) using 4264 TsaC/Sua5 sequences and 1955 YciO sequences (Supplementary Data 1 and 2). To generate the TsaC/YciO phylogenetic tree, proteins were aligned using MUSCLE (Edgar 2004). The alignment was trimmed using BMGE (Criscuolo and Gribaldo 2010) and used to build the maximum likelihood tree using FastTree (Price et al. 2010) with LG + CAT model with bootstrap (1,000 replicates) and visualized using iTOL (Letunic and Bork 2021).
SSNs and gene neighborhood analyses
As an example of the computational analysis that is required to separate paralogs described in Fig. 1, we generated SSNs and the corresponding gene neighborhood networks for the PF01300 using EFI Enzyme Similarity Tool (EFI-EST, efi.igb.illinois.edu/efi-est) (Zallot et al. 2019). Briefly, 54,820 sequences of the PF01300 family between 150 and 350 aa in length were retrieved from UniProt and subjected to EFI-EST. Each node in the network represents 1 or multiple sequences that share no less than 70% identity. The initial SSN was generated with an alignment score (AS) cutoff set such that each connection (edge) represented a sequence identity above 40%. The nodes of paralogs were colored as given in the legend and visualized using Cytoscape (3.10.1) (Shannon et al. 2003). More SSNs were created by gradually increasing the AS cutoff in small increments (usually by 5 AS units). This process was repeated until most clusters were homogeneous in color (AS = 70). The genome neighborhood graphs were generated using Gene Graphics (https://genegraphics.net/) (Harrison et al. 2018).
Miscellaneous data extraction and analysis
Open AI ChatGPT 4.o was used to extract data from tables in publications and generate figures (https://chatgpt.com/) on August 20 using the prompt: “Extract data from Table X in attached pdf file.”
Explainable AI methods
To enhance the interpretability of DeepECTF's multilabel enzyme function predictions, we incorporated an explainable AI (XAI) module based on the local interpretable model-agnostic explanation (LIME) framework (Ribeiro et al. 2016). The approach was specifically adapted to address the multilabel nature of EC number predictions and to provide both local and global insights into the protein sequence segments or residues driving model decisions. Traditional LIME is designed for binary or multiclass outputs, but enzyme function prediction often involves multilabel assignments. To address this, we implemented a multilabel adaptation of LIME, which constructs independent explanation pipelines for each possible label. For each protein sequence input, the XAI module isolates the probability output for a single EC label using the model's sigmoid activation, enabling LIME to generate label-specific explanations. To optimize computational efficiency, explanations are generated primarily for the label with the highest predicted probability for each input sequence, focusing interpretability efforts on the most relevant predictions.
To balance explanation quality with computational feasibility, local LIME explanations are computed for a representative subsample of up to 500 protein sequences at a time. For each sequence, the module generates a local feature importance map, highlighting which residues in the protein sequence most influenced the model's prediction for the selected EC label. For proteins outside this subsample, residue-level importance is estimated based on the aggregate statistics from the explained subset. This approach enables the extraction of both local (residue-level) and global (dataset-level) feature importance profiles, which are subsequently exported for downstream analysis. To further dissect model behavior, feature importance scores derived from LIME explanations were stratified by prediction type (correct predictions [CORs], paralog errors, nonparalog errors, and repetitions [REPs]; see Supplementary Table 1). Residue-level importance values are normalized within each error type, and summary plots are generated to visualize the distribution and magnitude of important sequence segments across different error categories.
The XAI module is fully integrated into the DeepECTF workflow, allowing users to generate explainability outputs alongside standard predictions without additional user intervention (https://github.com/Dias-Lab/XAI_DeepProZyme). All explanation results, including local and global feature importance scores, are made available in standard tabular formats for further interpretation or visualization. This explainable AI approach provides actionable insights into the decision-making process of DeepECTF, supporting both the validation of CORs and the systematic investigation of model failures, facilitating the development of more robust and trustworthy protein function prediction systems.
Results
Most correctly AI-predicted EC numbers are generic or already in the training set
We analyzed 453 E. coli proteins functionally annotated in the Kim et al. (2023) study (Supplementary Table 1). One hundred and twenty-one had the same EC number in the August 2024 corresponding UniProt annotation (Bateman et al. 2023), and 87% of these were in the version used to generate the training dataset (Fig. 4; Supplementary Table 1a) making these an example of training data contamination. Fifteen proteins had the exact same function labeled in UniProt but without an EC number or with a partial EC number (Fig. 4; Supplementary Table 1b and c) and are examples of successful annotation propagation (Fig. 1). These 136 cases were all considered correct but not novel (CNN) predictions with CSs of 2 (Fig. 4).
Fig. 4.
Classification of DeepECTF predictions. The 453 EC number predictions for 453 E. coli unknowns were manually classified into categories by comparing the EC number in the UniProt database and labeled with CSs from 0 to 2. Data extracted from Supplementary Table 1. COR, correct prediction; UNC, uncertain; LSP, less precise; PLI, paralogs incorrect; NPI, nonparalog incorrect; REPs, repetitions; CNN, correct but not novel; CS, confidence score.
The remaining 317 predictions can be split into proteins with partial or different EC numbers present in the UniProt annotation (63 cases; Supplementary Table 1b) or with no EC number in the corresponding UniProt entry (254 cases; Supplementary Table 1c). For 27 predictions, more generic annotations than the existing UniProt annotations were given (less precise cases or LSP in Supplementary Table 1b and c). These were also given a CS of 2. The 290 remaining cases were subject to human curation. We combined PaperBLAST searches and UniProt and EcoCyc data analyses to group the 290 DeepECTF predictions into different categories with CSs from 0 to 2 (Supplementary Table 1; Fig. 4). The 42 predictions that could neither be validated nor refuted were labeled as uncertain with a CS of 1. Three cases, including YgfF discussed below, were validated by publications not captured in UniProt and can be considered successful predictions (CSs of 2). Combining the CORrect (COR, 3), the CNN (136), and the LSP or generic (27) predictions brings the number of CORs to 166 (36.6%). Two hundred and forty-five predictions (54%) were inconsistent with the existing evidence and given CSs of 0 (Supplementary Table 1b and c).
Manual analyses reveal logical inconsistencies in AI predictions of unknowns
The 245 predictions with CSs of 0 can be separated into 2 categories: those with published evidence refuting the annotation (77 of nonparalog incorrect [NPI] and 42 paralog incorrect [PLI] in Supplementary Table 1b and c) and those that were replications of identical EC numbers (126 REPs in Supplementary Table 1b and c). Examples of the first type (n = 119) are given in Table 2. For example, YjhQ/b4307 is predicted to be a mycothiol synthase (EC 2.3.1.189), but mycothiol is not a molecule synthesized by E. coli, and the remaining pathway genes are absent from the genome (BioCyc ID: PWY1G-0). YrhB/b3446 is predicted to be a 6-carboxytetrahydropterin synthase (EC 4.1.2.50), but E. coli already encodes this enzyme (QueD/b2765) and a queD mutant lacks this activity (Zallot et al. 2017).
Table 2.
Example where experimental evidence contradicts the DeepECTF predictions.
| Locus tag | UniProt EC | UniProt function | DeepEC EC | DeepEC function | Notes |
|---|---|---|---|---|---|
| Example of overpropagation mistakes | |||||
| b2736 | 1.1.1.411 | L-Threonate dehydrogenase | 1.1.1.60 | 2-Hydroxy-3-oxopropionate reductase | In vivo and in vitro data validate these proteins as involved D-Threonate degradation (Zhang et al. 2016 ) |
| b2738 | 4.1.1.104 | 3-Oxo-tetronate 4-phosphate decarboxylase | 4.1.2.17 | L-Fuculose-phosphate aldolase | |
| b2739 | 5.3.1.35 | 2-Oxo-tetronate isomerase | 5.3.1.22 | Hydroxypyruvate isomerase | |
| b4403 | 2.1.1.- | Uncharacterized tRNA/rRNA methyltransferase LasT | 2.1.1.200 | tRNA (cytidine32/uridine32-2′-O)-methyltransferase | This EC 2.1.1.200 activity is catalyzed by TrmJ/b2532, and the mutant is devoid of the Um/Cm32 modification (Purta et al. 2006) |
| b2160 | 2.7.1.- | Uncharacterized sugar kinase YeiI | 2.7.1.83 | Pseudouridine kinase | This activity is catalyzed by PsuK/b2166, and the mutant cannot use Psi as a C source (Preumont et al. 2008) |
| b3038 | 6.3.1.- | Putative acid–amine ligase YgiC | 6.3.1.8 | Glutathionylspermidine synthase | This activity is catalyzed by Gsp, and YgiC does not have the same activity (Sui et al. 2012) |
| b1267 | None | Uncharacterized protein YciO | 2.7.7.87 | L-Threonylcarbamoyladenylate synthase | YciO is a paralog of TsaC but does catalyze the same activity (El Yacoubi et al. 2009; Gerdes et al. 2011) |
| Example of refuted predictions | |||||
| b2438 | None | Bacterial microcompartment shell protein EutK (ethanolamine utilization protein EutK) | 2.1.1.223 | tRNA1Val (adenine37-N6)-methyltransferase | This activity is catalyzed by TrmN6/b2575, and the mutant does not make the modification (Golovina et al. 2009) |
| b0254 | None | HTH-type transcriptional regulator PerR (peroxide resistance protein PerR) | 2.4.2.29 | tRNA-guanosine34 preQ1 transglycosylase | This activity is catalyzed by Tgt/b0406, and the mutant does not insert preQ1 in tRNA (Noguchi et al. 1982) |
| b3446 | None | Uncharacterized protein YrhB | 4.1.2.50 | 6-Carboxytetrahydropterin synthase | This activity is catalyzed by QueD/b2765), and the mutant lacks this activity (Zallot et al. 2017) |
| b4307 | 2.3.1.- | Uncharacterized N-acetyltransferase YjhQ | 2.3.1.189 | Mycothiol synthase | Mycothiol is not synthesized by E. coli, and the remaining pathway enzymes are absent (BioCyc ID: PWY1G-0) |
Replication of identical EC numbers occurred for 126 proteins (Fig. 4; Supplementary Table 1b and c). REPs of EC numbers do occur in bacterial genomes, particularly with families that, despite having 4 EC numbers, have generic functions. For example, histidine kinases with different substrate specificities are frequent, with 29 annotated in E. coli (Supplementary Table 1d, top). However, analysis of the protein family domain membership showed that most of the REPs in these data were errors, except for the less specific annotation cases classified as CORs above (Supplementary Table 1c). For example, out of the 12 proteins annotated as histidine kinase (EC 2.7.13.3) in the 453 DeepECTF predictions, none of them have sequence similarity to histidine kinase families, and 8 have been annotated with different and experimentally validated functions (such as ferric enterobactin transport protein FepE for b0587) (Supplementary Table 1c). For the 15 proteins annotated as “protein-Npi-phosphohistidine-sugar phosphotransferase” (EC 2.7.1.69) (or PTS family proteins), 4 were indeed PTS transporters but were given less specific annotations than the ones in UniProt (Supplementary Table 1c). The 11 remaining were part of transporter families not related to PTS (Supplementary Table 1b). This type of error may be due to inherent limitations in how AI methods operate. Indeed, if the input features in the training data lack the biological structure and cannot leverage information to distinguish between different functions, the model is expected to make frequency-dependent predictions that reflect the training data, as demonstrated with histidine kinases.
Correct separation of paralogous group can confirm or refute AI-based functional predictions
Many of the incorrect or uncertain predictions (CSs of 0 or 1) were part of protein families with paralogs (Fig. 4; Supplementary Table 1b and c). For example, b2100 was annotated as a dehydro-2-deoxygluconokinase (EC 2.7.1.92) using DeepECTF and as an uncharacterized sugar kinase YegV (EC 2.7.1.-) in UniProt (Supplementary Table 1b). These 2 predictions differ by the fourth or last position of the EC number that specifies substrate specificity. This protein is a member of a superfamily of sugar kinases with multiple nonisofunctional paralogous subgroups that phosphorylate different substrates (Supplementary Table 1d, bottom). The dehydro-2-deoxygluconokinase (EC 2.7.1.92) activity is encoded by another member of this superfamily KdgK/b3526 (Supplementary Table 1d, bottom). Here, DeepECTF predicted correctly the first 3 digits of the EC number but not the last, making an overpropagation mistake (error 6 in Table 1 and Fig. 1). Correctly separating nonisofunctional paralogous subgroups in a superfamily is difficult, requiring extensive examination of all of the evidence and oftentimes experiments.
We performed an additional, in-depth analysis of the 3 proteins described using in vitro assays in the Kim et al. (2023) study (YgfF, YjdM, and YciO) and show that in vitro functionality does not always correspond with in vivo functionality.
YgfF analysis
YgfF is a member of the large short-chain dehydrogenase/reductase (SDR) superfamily (IPR002347). The Oppermann and Persson groups developed a nomenclature system and HMM-based classification (http://www.sdr-enzymes.org/) that distinguish different functional subgroups of the SDR superfamily (Persson et al. 2009; Kallberg et al. 2010). This resource predicts YgfF is part of the SDR63C/Glucose 1-dehydrogenase subgroup, the activity predicted and validated in the Kim et al. (2023) study. This prediction demonstrates the accurate propagation of functional annotation and is a successful prediction.
YjdM analysis
YjdM was predicted and shown to catalyze phosphonoacetate hydrolase (PhnA) (EC 3.11.1.2) activity in vitro. In E. coli, the yjdM gene is located upstream of the methylphosphonate catabolism operon (phnCDEFGHIJKLMNOP) (Fig. 5). Underscoring the challenges in annotations, the initial report that the protein was involved in phosphonate catabolism was later refuted with additional genetic analyses (Metcalf and Wanner 1993). The experimentally validated PhnA is part of a nonhomologous family and expression of members of this family in E. coli suggested that PhnA activity was not present in this organism (Kulakova et al. 1997). Genome neighborhoods of phnA show strong clustering with genes encoding phosphonoacetate transporters, phosphonoacetate sensing regulators, and, in some cases, enzymes involved in 2-aminoethylphosphonate catabolism (Kulakova et al. 2001). However, except for E. coli, yjdM genes are generally not close to phosphonate catabolism or transport genes (Fig. 5). In conclusion, the PhnA activity observed in vitro is not supported, and additional in vivo experiments are required to confirm the biological role of this enzyme. This prediction was given a CS of 1.
Fig. 5.
Gene neighborhoods and metabolic reconstructions do not link YjdM to phosphonate degradation. Most yjdM genes are not near phosphonate degradation operons. A SSN of 11,986 IPR004624 family members was generated, and the corresponding genomic neighborhood information is available Supplementary Table 4.
YciO analysis
YciO is a member of the same Pfam family (PF01300) as TsaC/Sua5, and DeepEC predicts YciO has the same function as TsaC/Sua5. TsaC and Sua5 (the latter known in bacteria as TsaC2) are 2 types of the well-characterized L-threonylcarbamoyladenylate synthase (EC 2.7.7.87) that catalyzes the first step in the synthesis of the universal tRNA modified nucleoside N-6-threonylcarbamoyladenosine or t6A (TsaC and Sua5 have a common catalytic domain and differ by the presence of an additional domain in Sua5) (Su et al. 2022; Pichard-Kostuch et al. 2023). The function of TsaC/Sua5 was first elucidated in 2009 (El Yacoubi et al. 2009). A structure-based multisequence alignment comparing YciO and TsaC/Sua5 shows that the active site residues of TsaC are largely conserved in YciO, suggesting similar catalytic activities (Fig. 6a). YciO catalyzed the synthesis of L-threonylcarbamoyladenylate from ATP, L-threonine, and bicarbonate, in vitro (Kim et al. 2023). However, the activity reported (0.14 nM/min TC-AMP production rate) for E. coli YciO is more than 4 orders of magnitude weaker than that of E. coli TsaC (2.8 μM/min) at the same enzyme concentration and similar reaction conditions (Swinehart et al. 2020), consistent with the possibility of a missing partner or a different biological substrate for YciO.
Fig. 6.
YciO harbors specific molecular surface features. a) Structure-based multisequence alignment of TsaC proteins, the TsaC domains of Sua5 proteins and YciO proteins, derived from available crystal structures and AlphaFold structural models. For crystal structures, the PDB IDs are indicated in the sequence name after the hyphen. Secondary structure elements from the crystal structures of E. coli TsaC and E. coli YciO are displayed above and below the sequences, respectively. YciO-specific conserved basic residues forming the positively charged surface patch of YciO are shaded in blue. Stars above and below the alignment indicate the crystallographically observed substrate binding residues in StSua5 and the corresponding putative substrate binding residues in EcYciO, respectively. Green stars indicate the Mg2+-ATP binding residues. Orange stars indicate the binding residues for the L-threonine substrate. StSua5, Sulfurisphaera tokodaii Sua5 (UniProt ID Q9UYB2); PaSua5, Pyrococcus abyssi Sua5 (UniProt Q9UYB2); EcTsaC, Escherichia coli TsaC (UniProt P45748); PvTsaN, TsaC domain of Pandoravirus TsaN (UniProt A0A291ATS8). TmTsaC2, Thermotoga maritima TsaC2 (by the authors, unpublished, UniProt Q9WZV6); AbYciO, Actinomycetales bacterium YciO (UniProt A0A1R4F4R9); BmYciO, Burkholderia multivorans YciO (UniProt A0A1B4MSV9); KpYciO, Klebsiella pneumoniae YciO (UniProt A6T7X1); PpYciO, Pseudomonas putida YciO (UniProt A5W0B1); SeYciO, Salmonella enterica TciO (UniProt A0A601PQ14); SfYciO, Shigella flexneri YciO (UniProt P0AFR6); EcYciO, Escherichia coli YciO (UniProt P0AFR4). b and c) Surface representations of the crystal structures of EcTsaC b) and EcYciO c), color-coded by surface electrostatic potential (top) and by positional sequence conservation score calculated from 4264 TsaC/Sua5 sequences and 1955 YciO sequences (bottom). The color keys for both panels are shown on the right. The conserved, YciO-specific, positively charged surface patch (6% of total molecular surface area) is encircled with a dashed line. The active center of TsaC and the putative active center of YciO are marked with asterisks. The figure illustrates that although the active site is conserved in both protein families, the positively charged surface patch is present and conserved only in the YciO family.
YciO does not perform the same function as TsaC/Susa5 in vivo experiments (El Yacoubi et al. 2009; Gerdes et al. 2011). Genome neighborhood and structural data suggest that the function of YciO may be related to rRNA rather than tRNA metabolism. The evidence related to rRNA is as follows: (i) the structure of YciO exhibits a large positively charged surface predicted to interact with RNA (Jia et al. 2002) (Fig. 6b and c). This large positively charged surface is conserved in YciO proteins and is absent in TsaC proteins. (ii) In many species, yciO genes are colocalized with rnm genes (Fig. 7), which encode the recently characterized RNase AM, a 5′ to 3′ exonuclease that matures the 5′ end of all 3 ribosomal RNAs in E. coli (Jain 2020).
Fig. 7.
The genes encoding members of the yciO and tsaC have different genome neighborhood contexts. a) SSN of 54,820 PF01300 family members between 150 and 350 aa in length generated using EFI-EST. Each node in the network represents 1 or multiple sequences that share no less than 70% identity. An edge, represented as a line, is drawn between 2 nodes with AS higher than 70 (similar in magnitude to the negative base-10 logarithm of a BLAST e-value of 1E-70). Paralogs were colored in pairs as indicated. Node borders were colored by gene neighborhood context, yciO and rnm cluster (red) and tsaC, dprA, smg, yrdD, aroE, and yrdB cluster (blue). For better visualization, clusters of less than 40 nodes and singular nodes were hidden. The proteins used in the SSN are available in Supplementary Table 3. b) Phylogenetic tree of TsaC and YciO proteins. Sixty-two PF01300 proteins were selected from each major cluster in the family SSN and 3 RibB proteins (UniProt Q60364, Q46TZ9, and P0A7J0) were used as the outgroup. Bootstrap values less than 0.75 were indicated by blue dots. Branches and leaves were colored by phyla. TsaC2 proteins that contain YciO/TsaC (PF01300) and Sua5 (PF03481) domains are enclosed in a bracket. The gene neighborhood information is available in Supplementary Table 4. Selected example organisms were indicated by arrows and a 2-letter code in the SSN. Each corresponding genome neighborhood schematic was indicated by a circle in the same color as the node. As, Arthrobacter saudimassiliensis; Ec, Escherichia coli; Pr, Pseudohaliea rubra; Sd, Shigella dysenteriae; Sx, Shewanella xiamenensis.
Models of enzyme evolution go through promiscuous stages (Ribeiro et al. 2023). TsaC is an enzyme predicted to have been present in the Last Universal Common Ancestor (Gagler et al. 2022; Pichard-Kostuch et al. 2023). YciO is a likely paralog of TsaC (Fig. 6), and as such, it is likely to have residual ancestral catalytic activity that can be detected in vitro. In summary, the functional puzzle is far from being solved for proteins of the YciO subgroup, and even if the existing data suggest a role in RNA metabolism, it cannot be the same as TsaC, and the EC number 2.7.7.87 prediction was given a CS of 0.
Feature importance patterns identify limitations in the DeepEC model
We explored if the use of XAI could help improve confidence in DeepECTF predictions using the 4 groups of proteins labeled with CORs, paralog errors, nonparalog errors, or REP errors (Fig. 4; Supplementary Table 1). Application of the XAI module to DeepECTF predictions enabled residue-level analysis of feature importance for EC prediction calls across protein sequences (Fig. 8). For CORs, the feature importance profiles identified a small number of residues—often corresponding to known catalytic or conserved domain positions, showing high contribution scores, indicating that features whose function relies on localized residues for function are more likely to be accurately predicted. In contrast, for all categories of erroneous predictions, feature importance profiles had small deviations from the null values, indicating that the model is not able to leverage information about the protein sequence in making its predictions.
Fig. 8.
Residue-level feature importance profiles for DeepECTF highlight the distinct interpretability signatures associated with model predictions. Each line plot displays the normalized contribution (average feature importance computed via LIME) of individual residues across protein sequences for 4 categories: CORs a), paralog errors b), nonparalog errors c), and hallucinations or repetitive predictions d). Shaded regions represent standard deviation across instances within each category.
Discussion
Our expert curation of ML-based EC number predictions for E. coli “unknowns” proteins reveals that these methods still have a lot of room to improve. Biases, data imbalance (under- or overrepresented domains, motifs, and functions), generic EC numbers, lack of data structures that enable the inclusion of feature information, architectural limitations (e.g. incapacity of capturing complex patterns, inability to integrate diverse data sources, and lack of regularization), and poor uncertainty calibration all contribute to challenges in training computational models and are likely contributors in frequency-dependent predictions (Pucci et al. 2018; Urban et al. 2020; Mardikoraem and Woldring 2023; Shahbazi et al. 2023). Efforts to improve the quality and consistency of training data and avoid data leakage between training and testing data can help minimize errors, but metrics evaluating the uncertainty in predicted protein functions, or an assessment of the likelihood of the label assignment, need to become standard model outputs. In the absence of these metrics, XAI can be used to provide insights into feature importance and model behavior. Many advanced ML models, particularly deep neural networks such as DeepECTF, function as “black boxes,” where the prediction algorithm is difficult to interpret. XAI techniques can help to dissect “black-box” models and understand what input features are driving COR and incorrect prediction (Chennam et al. 2023) and provide insights about which data to include in future models to improve predictions.
Our study also shows that the model's lack of inclusion of the broader accumulated knowledge in the fields of protein structure and function, biochemical pathways, and the absence of a mechanism to identify “illogical” predictions led to erroneous predictions for most proteins of unknown functions, although the model was able to successfully propagate functions among proteins that were included in the training data. We also want to emphasize that in vitro activity alone is not sufficient to validate the function of a protein in vivo (García-Contreras et al. 2012; Punekar 2018). Indeed, enzymes evolve by duplication, divergence, and subsequent sub-, neo-, or hypofunctionalization (Ribeiro et al. 2023; Birchler 2025). In vitro activities of many enzymes can show promiscuity, which is useful for biotechnological applications (Robinson et al. 2020) but does not guarantee that the protein plays that specific role in vivo (Copley 2015) as shown here with the YciO example. Best practices in functional annotations of enzymes couple biochemical and contextual evidence and reflect the GO Consortium definitions, capturing the cellular component as well as the molecular and biological functions (Ashburner et al. 2000).
PLM-driven approaches are becoming mainstream tools for propagating known functional annotations among isofunctional proteins (https://www.uniprot.org/help/ProtNLM). Computational models will be further enabled by the integration of complementary evidence such as structural data to identify active site signature residues (Derry et al. 2025; Sajid et al. 2024; Yu et al. 2023), gene neighborhood context (Urhan et al. 2024; Hwang et al. 2024; Jha et al 2025), and chemical reaction specificity (Qian et al. 2024), all of which help to distinguish nonisofunctional paralogous subgroups (Ribeiro et al. 2023). Recent models like MAPred (Rong et al. 2024) align with modern goals of interpretability and transparency in ML models (Chennam et al. 2023; Rong et al. 2024; Derry et al. 2025) and implement a feature-dropping approach that represents a step toward a more quantitative and nuanced “self” evaluation of model predictions. The future in this space is exciting, even though the limits of these models are just now beginning to be understood (Muir et al 2025).
Acknowledgments
We thank Alan J. Bridge (SIB) for the critical reading and input on early versions of the manuscript and the peer editors and reviewers who helped us greatly improve the manuscript.
Contributor Information
Valérie de Crécy-Lagard, Department of Microbiology and Cell Science, University of Florida, Gainesville, FL 32611, United States; Genetic Institute, University of Florida, Gainesville, FL 32611, United States.
Raquel Dias, Department of Microbiology and Cell Science, University of Florida, Gainesville, FL 32611, United States.
Nick Sexson, Department of Microbiology and Cell Science, University of Florida, Gainesville, FL 32611, United States.
Iddo Friedberg, Department of Veterinary Microbiology and Preventive Medicine, Iowa State University, Ames, IA 50011, United States.
Yifeng Yuan, Department of Microbiology and Cell Science, University of Florida, Gainesville, FL 32611, United States.
Manal A Swairjo, Department of Chemistry and Biochemistry, San Diego State University, San Diego, CA 92182, United States; The Viral Information Institute, San Diego State University, San Diego, CA 92182, United States.
Data availability
Supplemental files are available at Figshare (https://doi.org/10.6084/m9.figshare.29602418.v2). Supplementary Table 1 contains the manual review of EC number predictions for 450 E. coli unknowns made by DEEP EC. Supplementary Table 2 contains gene neighborhood information for selected yjdM (IPR004624 family) genes shown in Fig. 5. Supplementary Table 3 includes the identifiers for the PF01300 family members used in Fig. 7 and SSN, and Supplementary Table 4 contains the gene neighborhood information for selected PF01300 family genes. Supplementary Data 1 contains the TsaC sequences used to generate positional conservation scores for Fig. 6. Supplementary Data 2 contains YciO sequences used to generate the positional conservation scores for Fig. 7.
Funding
This work was funded by NationaI Institute of General Medical Sciences (GM110588 to MAS and VdC-L and GM145937 to IF) and Defense Advanced Research Projects Agency (HR0012530305 to VDC and RD).
Conflicts of interest. None declared.
Author contributions
VdC-L conceived the study, wrote the first draft of the manuscript, and did the analyses for Tables 1 and 2 and Supplementary Table 1. MAS created Fig. 6 and wrote sections of the manuscript. YY helped generate the data for Supplementary Tables 1 to 4 and created Figs. 1 to 5 and 7. RD ran the XAI analysis, wrote sections of the manuscript, and created Fig. 8. NS ran the mapping and preprocessing of the input data for the XAI analysis and contributed to the implementation of the XAI module in DeepECTF. IF wrote sections of the manuscript. All authors edited the final draft.
Literature cited
- Ardern Z, Chakraborty S, Lenk F, Kaster A-K. 2023. Elucidating the functional roles of prokaryotic proteins using big data and artificial intelligence. FEMS Microbiol Rev. 47:fuad003. 10.1093/femsre/fuad003. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Ashburner M et al. 2000. Gene ontology: tool for the unification of biology. Nat Genet. 25:25–29. 10.1038/75556. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Ayres G et al. 2023. HiFi-NN annotates the microbial dark matter with enzyme commission numbers. In: Machine learning for structural biology workshop, NeurIPS 2023. [Google Scholar]
- Bansal P et al. 2022. Rhea, the reaction knowledgebase in 2022. Nucleic Acids Res. 50:D693–D700. 10.1093/nar/gkab1016. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Bateman A et al. 2023. UniProt: the universal protein knowledgebase in 2023. Nucleic Acids Res. 51:D523–D531. 10.1093/nar/gkac1052. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Birchler JA. 2025. When is it subfunctionalization and when is it not? G3 (Bethesda). 15:jkae269. 10.1093/g3journal/jkae269. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Bordin N, Sillitoe I, Lees JG, Orengo C. 2021. Tracing evolution through protein structures: nature captured in a few thousand folds. Front Mol Biosci. 8:668184. 10.3389/fmolb.2021.668184. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Burley SK et al. 2023. RCSB Protein Data Bank (RCSB.org): delivery of experimentally-determined PDB structures alongside one million computed structure models of proteins from artificial intelligence/machine learning. Nucleic Acids Res. 51:D488–D508. 10.1093/nar/gkac1077. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Chen L, Vitkup D. 2007. Distribution of orphan metabolic activities. Trends Biotechnol. 25:343–348. 10.1016/j.tibtech.2007.06.001. [DOI] [PubMed] [Google Scholar]
- Chennam KK, Mudrakola S, Maheswari VU, Aluvalu R, Rao KG. 2023. Black box models for eXplainable artificial intelligence. In: Mehta M, Palade V, Chatterjee I, editors. Explainable AI: foundations, methodologies and applications. Springer. p. 1–24. [Google Scholar]
- Copley SD. 2015. An evolutionary biochemist's perspective on promiscuity. Trends Biochem Sci. 40:72–78. 10.1016/j.tibs.2014.12.004. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Criscuolo A, Gribaldo S. 2010. BMGE (Block Mapping and Gathering with Entropy): a new software for selection of phylogenetic informative regions from multiple sequence alignments. BMC Evol Biol. 10:210. 10.1186/1471-2148-10-210. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Danchin A, Ouzounis C, Tokuyasu T, Zucker JD. 2018. No wisdom in the crowd: genome annotation in the era of big data—current status and future prospects. Microb Biotechnol. 11:588–605. 10.1111/1751-7915.13284. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Das S et al. 2015. Functional classification of CATH superfamilies: a domain-based approach for protein function annotation. Bioinformatics. 31:3460–3467. 10.1093/bioinformatics/btv398. [DOI] [PMC free article] [PubMed] [Google Scholar]
- de Crécy-Lagard V et al. 2012. Comparative genomics guided discovery of two missing archaeal enzyme families involved in the biosynthesis of the pterin moiety of tetrahydromethanopterin and tetrahydrofolate. ACS Chem Biol. 7:1807–1816. 10.1021/cb300342u. [DOI] [PMC free article] [PubMed] [Google Scholar]
- de Crécy-Lagard V. 2014. Variations in metabolic pathways create challenges for automated metabolic reconstructions: examples from the tetrahydrofolate synthesis pathway. Comput Struct Biotechnol J. 10:41–50. 10.1016/j.csbj.2014.05.008. [DOI] [PMC free article] [PubMed] [Google Scholar]
- de Crécy-lagard V et al. 2022. A roadmap for the functional annotation of protein families: a community perspective. Database. 2022:baac062. 10.1093/database/baac062. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Dembech E et al. 2023. Identification of hidden associations among eukaryotic genes through statistical analysis of coevolutionary transitions. Proc Natl Acad Sci U S A. 120:e2218329120. 10.1073/pnas.2218329120. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Derry A, Tartici A, Altman RB. Protein functional site annotation using local structure embeddings. bioRxiv 562298. 10.1101/2023.10.13.562298, 25 June 2025, preprint: not peer reviewed. [DOI] [PMC free article] [PubMed]
- Devoid S et al. 2013. Automated genome annotation and metabolic model reconstruction in the SEED and model SEED. Methods Mol Biol. 985:17–45. 10.1007/978-1-62703-299-5_2. [DOI] [PubMed] [Google Scholar]
- Durairaj J et al. 2023. Uncovering new families and folds in the natural protein universe. Nature. 622:646–653. 10.1038/s41586-023-06622-3. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Edgar RC. 2004. MUSCLE: multiple sequence alignment with high accuracy and high throughput. Nucleic Acids Res. 32:1792–1797. 10.1093/nar/gkh340. [DOI] [PMC free article] [PubMed] [Google Scholar]
- El Yacoubi B et al. 2009. The universal YrdC/Sua5 family is required for the formation of threonylcarbamoyladenosine in tRNA. Nucleic Acids Res. 37:2894–2909. 10.1093/nar/gkp152. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Friedberg I. 2006. Automated protein function prediction—the genomic challenge. Brief Bioinform. 7:225–242. 10.1093/bib/bbl004. [DOI] [PubMed] [Google Scholar]
- Gagler DC et al. 2022. Scaling laws in enzyme function reveal a new kind of biochemical universality. Proc Natl Acad Sci U S A. 119:e2106655119. 10.1073/pnas.2106655119. [DOI] [PMC free article] [PubMed] [Google Scholar]
- García-Contreras R, Vos P, Westerhoff HV, Boogerd FC. 2012. Why in vivo may not equal in vitro—new effectors revealed by measurement of enzymatic activities under the same in vivo -like assay conditions. FEBS J. 279:4145–4159. 10.1111/febs.12007. [DOI] [PubMed] [Google Scholar]
- Gerdes S et al. 2011. Synergistic use of plant-prokaryote comparative genomics for functional annotations. BMC Genom. 12:S2. 10.1186/1471-2164-12-S1-S2. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Gerlt JA, Babbitt PC, Jacobson MP, Almo SC. 2012. Divergent evolution in enolase superfamily: strategies for assigning functions. J Biol Chem. 287:29–34. 10.1074/jbc.R111.240945. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Ghatak S, King ZA, Sastry A, Palsson BO. 2019. The y-ome defines the 35% of Escherichia coli genes that lack experimental evidence of function. Nucleic Acids Res. 47:2446–2454. 10.1093/nar/gkz030. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Glasner ME, Gerlt JA, Babbitt PC. 2006. Evolution of enzyme superfamilies. Curr Opin Chem Biol. 10:492–497. 10.1016/j.cbpa.2006.08.012. [DOI] [PubMed] [Google Scholar]
- Golovina AY et al. 2009. The yfiC gene of E. coli encodes an adenine-N6 methyltransferase that specifically modifies A37 of tRNA1Val (cmo5UAC). RNA. 15:1134–1141. 10.1261/rna.1494409. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Hamamsy T et al. Learning sequence, structure, and function representations of proteins with language models. bioRxiv 568742. 10.1101/2023.11.26.568742, 26 November 2023, preprint: not peer reviewed. [DOI]
- Hanson AD, Pribat A, Waller JC, de Crécy-Lagard V. 2010. “Unknown” proteins and “orphan” enzymes: the missing half of the engineering parts list–and how to find it. Biochem J. 425:1–11. 10.1042/BJ20091328. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Harrison KJ, de Crécy-Lagard V, Zallot R. 2018. Gene Graphics: a genomic neighborhood data visualization web application. Bioinformatics. 34:1406–1408. 10.1093/bioinformatics/btx793. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Hegyi H, Gerstein M. 2001. Annotation transfer for genomics: measuring functional divergence in multi-domain proteins. Genome Res. 11:1632–1640. 10.1101/gr.183801. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Henry CS et al. 2016. Systematic identification and analysis of frequent gene fusion events in metabolic pathways. BMC Genomics. 17:473. 10.1186/s12864-016-2782-3. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Huang Y, Lin Y, Lan W, Huang C, Zhong C. 2024. GloEC: a hierarchical-aware global model for predicting enzyme function. Brief Bioinform. 25:bbae365. 10.1093/bib/bbae365. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Hwang Y, Cornman AL, Kellogg EH, Ovchinnikov S, Girguis PR. 2024. Genomic language model predicts protein co-regulation and function. Nat Commun. 15:2880. 10.1038/s41467-024-46947-9. [DOI] [PMC free article] [PubMed] [Google Scholar]
- International Society for Biocuration . 2018. Biocuration: distilling data into knowledge. PLoS Biol. 16:e2002846. 10.1371/journal.pbio.2002846. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Jain C. 2020. RNase AM, a 5′ to 3′ exonuclease, matures the 5′ end of all three ribosomal RNAs in E. coli. Nucleic Acids Res. 48:5616–5623. 10.1093/NAR/GKAA260. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Jha N et al. 2025. Gaia: an AI-enabled genomic context-aware platform for protein sequence annotation. Sci Adv. 11:eadv5109. 10.1126/sciadv.adv5109. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Jia J et al. 2002. Crystal structure of the YciO protein from Escherichia coli. Proteins. 49:139–141. 10.1002/prot.10178. [DOI] [PubMed] [Google Scholar]
- Jiang Y et al. 2016. An expanded evaluation of protein function prediction methods shows an improvement in accuracy. Genome Biol. 17:184. 10.1186/s13059-016-1037-6. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Kallberg Y, Oppermann U, Persson B. 2010. Classification of the short-chain dehydrogenase/reductase superfamily using hidden Markov models. FEBS J. 277:2375–2386. 10.1111/j.1742-4658.2010.07656.x. [DOI] [PubMed] [Google Scholar]
- Kanehisa M, Furumichi M, Sato Y, Ishiguro-Watanabe M, Tanabe M. 2021. KEGG: integrating viruses and cellular organisms. Nucleic Acids Res. 49:D545–D551. 10.1093/nar/gkaa970. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Karp P. 2004. Call for an enzyme genomics initiative. Genome Biol. 5:401. 10.1186/gb-2004-5-8-401. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Karp PD et al. 2019. The BioCyc collection of microbial genomes and metabolic pathways. Brief Bioinform. 20:1085–1093. 10.1093/bib/bbx085. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Kim GB et al. 2023. Functional annotation of enzyme-encoding genes using deep learning with transformer layers. Nat Commun. 14:7370. 10.1038/s41467-023-43216-z. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Kobras CM, Fenton AK, Sheppard SK. 2021. Next-generation microbiology: from comparative genomics to gene function. Genome Biol. 22:123. 10.1186/s13059-021-02344-9. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Kulakova AN et al. 2001. Structural and functional analysis of the phosphonoacetate hydrolase (phnA) gene region in Pseudomonas fluorescens 23F. J Bacteriol. 183:3268–3275. 10.1128/JB.183.11.3268-3275.2001. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Kulakova AN, Kulakov LA, Quinn JP. 1997. Cloning of the phosphonoacetate hydrolase gene from Pseudomonas fluorescens 23F encoding a new type of carbon–phosphorus bond cleaving enzyme and its expression in Escherichia coli and Pseudomonas putida. Gene. 195:49–53. 10.1016/S0378-1119(97)00151-0. [DOI] [PubMed] [Google Scholar]
- Kyrpides NC. 1999. Genomes OnLine Database (GOLD 1.0): a monitor of complete and ongoing genome projects worldwide. Bioinformatics. 15:773–774. 10.1093/bioinformatics/15.9.773. [DOI] [PubMed] [Google Scholar]
- Lee D, Redfern O, Orengo C. 2007. Predicting protein function from sequence and structure. Nat Rev Mol Cell Biol. 8:995. 10.1038/nrm2281. [DOI] [PubMed] [Google Scholar]
- Lespinet O, Labedan B. 2006. Puzzling over orphan enzymes. Cell Mol Life Sci. 63:517–523. 10.1007/s00018-005-5520-6. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Letunic I, Bork P. 2021. Interactive tree of life (iTOL) v5: an online tool for phylogenetic tree display and annotation. Nucleic Acids Res. 49:W293–W296. 10.1093/nar/gkab301. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Li W et al. 2021. RefSeq: expanding the prokaryotic genome annotation pipeline reach with protein family model curation. Nucleic Acids Res. 49:D1020–D1028. 10.1093/nar/gkaa1105. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Lobb B, Tremblay BJ-M, Moreno-Hagelsieb G, Doxey AC. 2020. An assessment of genome annotation coverage across the bacterial tree of life. Microb Genom. 6:e000341. 10.1099/mgen.0.000341. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Lockwood S, Brayton KA, Daily JA, Broschat SL. 2019. Whole proteome clustering of 2,307 proteobacterial genomes reveals conserved proteins and significant annotation issues. Front Microbiol. 10:383. 10.3389/fmicb.2019.00383. [DOI] [PMC free article] [PubMed] [Google Scholar]
- MacDougall A et al. 2020. UniRule: a unified rule resource for automatic annotation in the UniProt Knowledgebase. Bioinformatics. 36:4643–4648. 10.1093/bioinformatics/btaa485. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Mardikoraem M, Woldring D. 2023. Protein fitness prediction is impacted by the interplay of language models, ensemble learning, and sampling methods. Pharmaceutics. 15:1337. 10.3390/pharmaceutics15051337. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Metcalf WW, Wanner BL. 1993. Evidence for a fourteen-gene, phnC to phnP locus for phosphonate metabolism in Escherichia coli. Gene. 129:27–32. 10.1016/0378-1119(93)90692-V. [DOI] [PubMed] [Google Scholar]
- Muir DF et al. 2025. Evolutionary-scale enzymology enables exploration of a rugged catalytic landscape. Science. 388:eadu1058. 10.1126/science.adu1058. [DOI] [PMC free article] [PubMed] [Google Scholar]
- NCBI Resource Coordinators . 2016. Database resources of the National Center for Biotechnology Information. Nucleic Acids Res. 44:D7–D19. 10.1093/nar/gkv1290. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Niehaus TD, Thamm AMK, de Crécy-Lagard V, Hanson AD. 2015. Proteins of unknown biochemical function: a persistent problem and a roadmap to help overcome it. Plant Physiol. 169:pp.00959.2015. 10.1104/pp.15.00959. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Noguchi S, Nishimura Y, Hirota Y, Nishimura S. 1982. Isolation and characterization of an Escherichia coli mutant lacking tRNA-guanine transglycosylase: function and biosynthesis of queuosine in tRNA. J Biol Chem. 257:6544–6550. 10.1016/S0021-9258(20)65176-6. [DOI] [PubMed] [Google Scholar]
- Olson RD et al. 2023. Introducing the Bacterial and Viral Bioinformatics Resource Center (BV-BRC): a resource combining PATRIC, IRD and ViPR. Nucleic Acids Res. 51:D678–D689. 10.1093/nar/gkac1003. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Paysan-Lafosse T et al. 2023. InterPro in 2022. Nucleic Acids Res. 51:D418–D427. 10.1093/nar/gkac993. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Pei J, Grishin NV. 2001. AL2CO: calculation of positional conservation in a proteinsequence alignment. Bioinformatics. 17:700–712. 10.1093/bioinformatics/17.8.700. [DOI] [PubMed] [Google Scholar]
- Pei J, Tang M, Grishin NV. 2008. PROMALS3D web server for accurate multiple protein sequence and structure alignments. Nucleic Acids Res. 36:W30–W34. 10.1093/nar/gkn322. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Percudani R, Carnevali D, Puggioni V. 2013. Ureidoglycolate hydrolase, amidohydrolase, lyase: how errors in biological databases are incorporated in scientific papers and vice versa. Database. 2013:bat071. 10.1093/database/bat071. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Persson B et al. 2009. The SDR (short-chain dehydrogenase/reductase and related enzymes) nomenclature initiative. Chem Biol Interact. 178:94–98. 10.1016/j.cbi.2008.10.040. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Peterson JK, MacDonald ML, Ellis VA. 2025. First whole-genome sequence of Triatoma sanguisuga (Le Conte, 1855), vector of Chagas disease. G3 (Bethesda). 15:jkae308. 10.1093/g3journal/jkae308. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Pichard-Kostuch A, Da Cunha V, Oberto J, Sauguet L, Basta T. 2023. The universal Sua5/TsaC family evolved different mechanisms for the synthesis of a key tRNA modification. Front Microbiol. 14:1204045. 10.3389/fmicb.2023.1204045. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Poux S et al. 2017. On expert curation and scalability: UniProtKB/Swiss-Prot as a case study. Bioinformatics. 33:3454–3460. 10.1093/bioinformatics/btx439. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Prabakaran R, Bromberg Y. Functional profiling of the sequence stockpile: a review and assessment of in silico; prediction tools. bioRxiv 548726. 10.1101/2023.07.12.548726, 12 April 2024, preprint: not peer reviewed. [DOI] [PMC free article] [PubMed]
- Precord TW et al. 2023. Catalytic site proximity profiling for functional unification of sequence-diverse radical S-adenosylmethionine enzymes. ACS Bio Med Chem Au. 3:240–251. 10.1021/acsbiomedchemau.2c00085. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Preumont A, Snoussi K, Stroobant V, Collet JF, van Schaftingen E. 2008. Molecular identification of pseudouridine-metabolizing enzymes. J Biol Chem. 283:25238–25246. 10.1074/jbc.M804122200. [DOI] [PubMed] [Google Scholar]
- Price MN, Arkin AP. 2017. PaperBLAST: text mining papers for information about homologs. mSystems. 2:e00039-17. 10.1128/mSystems.00039-17. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Price MN, Arkin AP. 2024. Interactive tools for functional annotation of bacterial genomes. Database. 2024:baae089. 10.1093/database/baae089. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Price MN, Dehal PS, Arkin AP. 2010. FastTree 2—approximately maximum-likelihood trees for large alignments. PLoS One. 5:e9490. 10.1371/journal.pone.0009490. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Pucci F, Bernaerts KV, Kwasigroch JM, Rooman M. 2018. Quantification of biases in predictions of protein stability changes upon mutations. Bioinformatics. 34:3659–3665. 10.1093/bioinformatics/bty348. [DOI] [PubMed] [Google Scholar]
- Punekar NS. 2018. In vitro versus in vivo: concepts and consequences. In: ENZYMES: catalysis, kinetics and mechanisms. Springer. p. 493–519. [Google Scholar]
- Purta E et al. 2006. The yfhQ gene of Escherichia coli encodes a tRNA:Cm32/Um32 methyltransferase. BMC Mol Biol. 7:23. 10.1186/1471-2199-7-23. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Qian W et al. 2024. A general model for predicting enzyme functions based on enzymatic reactions. J Cheminform. 16:38. 10.1186/s13321-024-00827-y. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Radivojac P et al. 2013. A large-scale evaluation of computational protein function prediction. Nat Methods. 10:e0147952. 10.1371/JOURNAL.PONE.0147952. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Ramola R, Friedberg I, Radivojac P. 2022. The field of protein function prediction as viewed by different domain scientists. Bioinform Adv. 2:vbac057. 10.1093/bioadv/vbac057. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Reed CJ, Hutinet G, de Crécy-Lagard V. 2021. Comparative genomic analysis of the DUF34 protein family suggests role as a metal ion chaperone or insertase. Biomolecules. 11:1282. 10.3390/biom11091282. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Rembeza E, Engqvist MKM. 2021. Experimental and computational investigation of enzyme functional annotations uncovers misannotation in the EC 1.1.3.15 enzyme class. PLoS Comput Biol. 17:e1009446. 10.1371/journal.pcbi.1009446. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Ribeiro AJM, Riziotis IG, Borkakoti N, Thornton JM. 2023. Enzyme function and evolution through the lens of bioinformatics. Biochem J. 480:1845–1863. 10.1042/BCJ20220405. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Ribeiro MT, Singh S, Guestrin C. 2016. Why should I trust you? In: KDD '16: proceedings of the 22nd ACM SIGKDD international conference on knowledge discovery and data mining. New York, NY: ACM. p. 1135–1144. 10.1145/2939672.2939778. [DOI]
- Ribeiro AJM, Tyzack JD, Borkakoti N, Thornton JM. 2020. Identifying pseudoenzymes using functional annotation: pitfalls of common practice. FEBS J. 287:4128–4140. 10.1111/febs.15142. [DOI] [PubMed] [Google Scholar]
- Robinson SL, Smith MD, Richman JE, Aukema KG, Wackett LP. 2020. Machine learning-based prediction of activity and substrate specificity for OleA enzymes in the thiolase superfamily. Synthetic Biology. 5:ysaa004. 10.1093/synbio/ysaa004. [DOI] [Google Scholar]
- Rocha JJ et al. 2023. Functional unknomics: systematic screening of conserved genes of unknown function. PLoS Biol. 21:e3002222. 10.1371/journal.pbio.3002222. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Rodríguez del Río ” et al. 2024. Functional and evolutionary significance of unknown genes from uncultivated taxa. Nature. 626:377–384. 10.1038/s41586-023-06955-z. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Rong D et al. 2024. Aug 11. Autoregressive enzyme function prediction with multi-scale multi-modality fusion [preprint]. arXiv 06391. 10.48550/arXiv.2408.06391. [DOI] [PMC free article] [PubMed]
- Roy A, Srinivasan N, Gowri VS. 2009. Molecular and structural basis of drift in the functions of closely-related homologous enzyme domains: implications for function annotation based on homology searches and structural genomics. In Silico Biol. 9:S41–S55. 10.3233/ISB-2009-0379. [DOI] [PubMed] [Google Scholar]
- Sajid S et al. 2024. The Y-ome conundrum: insights into uncharacterized genes and approaches for functional annotation. Mol Cell Biochem. 479:1957–1968. 10.1007/s11010-023-04827-8. [DOI] [PubMed] [Google Scholar]
- Sanderson T, Bileschi ML, Belanger D, Colwell LJ. 2023. ProteInfer, deep neural networks for protein functional inference. Elife. 12:e80942. 10.7554/eLife.80942. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Satish Kumar V, Dasika MS, Maranas CD. 2007. Optimization based automated curation of metabolic reconstructions. BMC Bioinformatics. 8:212. 10.1186/1471-2105-8-212. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Schnoes AM, Brown SD, Dodevski I, Babbitt PC. 2009. Annotation error in public databases: misannotation of molecular function in enzyme superfamilies. PLoS Comput Biol. 5:e1000605. 10.1371/journal.pcbi.1000605. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Seemann T. 2014. Prokka: rapid prokaryotic genome annotation. Bioinformatics. 30:2068–2069. 10.1093/bioinformatics/btu153. [DOI] [PubMed] [Google Scholar]
- Shahbazi N, Lin Y, Asudeh A, Jagadish HV. 2023. Representation bias in data: a survey on identification and resolution techniques. ACM Comput Surv. 55:1–39. 10.1145/3588433. [DOI] [Google Scholar]
- Shannon P et al. 2003. Cytoscape: a software environment for integrated models of biomolecular interaction networks. Genome Res. 13:2498–2504. 10.1101/gr.1239303. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Škunca N, Roberts RJ, Steffen M. 2017. Evaluating computational Gene Ontology annotations. Methods Mol Biol. 1446:97–109. 10.1007/978-1-4939-3743-1_8. [DOI] [PubMed] [Google Scholar]
- Soldatos TG, Perdigão N, Brown NP, Sabir KS, O’Donoghue SI. 2015. How to learn about gene function: text-mining or ontologies? Methods. 74:3–15. 10.1016/j.ymeth.2014.07.004. [DOI] [PubMed] [Google Scholar]
- Storm CE, Sonnhammer ELL. 2002. Automated ortholog inference from phylogenetic trees andcalculation of orthology reliability. Bioinformatics. 18:92–99. 10.1093/bioinformatics/18.1.92. [DOI] [PubMed] [Google Scholar]
- Su C, Jin M, Zhang W. 2022. Conservation and diversification of tRNA t6A-modifying enzymes across the three domains of life. Int J Mol Sci. 23:13600. 10.3390/ijms232113600. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Sui L, Warren JC, Russell JPN, Stourman NV. 2012. Comparison of the functions of glutathionylspermidine synthetase/amidase from E. coli and its predicted homologues YgiC and YjfC. Int J Biochem Mol Biol. 3:302–312. [PMC free article] [PubMed] [Google Scholar]
- Swinehart W et al. 2020. Specificity in the biosynthesis of the universal tRNA nucleoside N6-threonylcarbamoyl adenosine (t6A)-TsaD is the gatekeeper. RNA. 26:1094–1103. 10.1261/rna.075747.120. [DOI] [PMC free article] [PubMed] [Google Scholar]
- The Gene Ontology Consortium . 2023. The gene ontology knowledgebase in 2023. Genetics. 224:iyad031. 10.1093/genetics/iyad031. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Thibaud-Nissen F et al. 2016. P8008 the NCBI eukaryotic genome annotation pipeline. J Anim Sci. 94:184–184. 10.2527/jas2016.94supplement4184x. [DOI] [Google Scholar]
- Todd AE, Orengo CA, Thornton JM. 2001. Evolution of function in protein superfamilies, from a structural perspective 1 1 edited by A. R. Fersht. J Mol Biol. 307:1113–1143. 10.1006/jmbi.2001.4513. [DOI] [PubMed] [Google Scholar]
- Torres J, Rozo-Lopez P, Brewer W, Saleh Ziabari O, Parker BJ. 2025. Draft genome sequence of the glasshouse-potato aphid Aulacorthum solani. G3 (Bethesda). 15:jkaf013. 10.1093/g3journal/jkaf013. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Urban G, Torrisi M, Magnan CN, Pollastri G, Baldi P. 2020. Protein profiles: biases and protocols. Comput Struct Biotechnol J. 18:2281–2289. 10.1016/j.csbj.2020.08.015. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Urhan A, Cosma B-M, Manson AL, Abeel T. 2024. SAFPred: synteny-aware gene function prediction for bacteria using protein embeddings. Bioinformatics. 40:btae328. 10.1093/bioinformatics/btae328. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Wei C-H, Allot A, Leaman R, Lu Z. 2019. PubTator central: automated concept annotation for biomedical full text articles. Nucleic Acids Res. 47:W587–W593. 10.1093/nar/gkz389. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Wood V et al. 2019. Hidden in plain sight: what remains to be discovered in the eukaryotic proteome? Open Biol. 9:180241. 10.1098/rsob.180241. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Yu T et al. 2023. Enzyme function prediction using contrastive learning. Science. 379:1358–1363. 10.1126/science.adf2465. [DOI] [PubMed] [Google Scholar]
- Zallot R, Harrison KJ, Kolaczkowski B, de Crécy-Lagard V. 2016. Functional annotations of paralogs: a blessing and a curse. Life. 6:39. 10.3390/life6030039. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Zallot R, Oberg N, Gerlt JA. 2019. The EFI web resource for genomic enzymology tools: leveraging protein, genome, and metagenome databases to discover novel enzymes and metabolic pathways. Biochemistry. 58:4169–4182. 10.1021/acs.biochem.9b00735. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Zallot R, Oberg N, Gerlt JA. 2021. Discovery of new enzymatic functions and metabolic pathways using genomic enzymology web tools. Curr Opin Biotechnol. 69:77–90. 10.1016/j.copbio.2020.12.004. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Zallot R, Yuan Y, de Crécy-Lagard V. 2017. The Escherichia coli COG1738 member YhhQ is involved in 7-cyanodeazaguanine (preQ₀) transport. Biomolecules. 7:12. 10.3390/biom7010012. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Zhang X et al. 2016. Assignment of function to a domain of unknown function: DUF1537 is a new kinase family in catabolic pathways for acid sugars. Proc Natl Acad Sci U S A. 113:E4161-9. 10.1073/pnas.1605546113. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Zhou N et al. 2019. The CAFA challenge reports improved protein function prediction and new functional annotations for hundreds of genes through experimental screens. Genome Biol. 20:244. 10.1186/s13059-019-1835-8. [DOI] [PMC free article] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Data Availability Statement
Supplemental files are available at Figshare (https://doi.org/10.6084/m9.figshare.29602418.v2). Supplementary Table 1 contains the manual review of EC number predictions for 450 E. coli unknowns made by DEEP EC. Supplementary Table 2 contains gene neighborhood information for selected yjdM (IPR004624 family) genes shown in Fig. 5. Supplementary Table 3 includes the identifiers for the PF01300 family members used in Fig. 7 and SSN, and Supplementary Table 4 contains the gene neighborhood information for selected PF01300 family genes. Supplementary Data 1 contains the TsaC sequences used to generate positional conservation scores for Fig. 6. Supplementary Data 2 contains YciO sequences used to generate the positional conservation scores for Fig. 7.








