Significance
Metabolic networks and enzyme-molecule connections are still being elucidated. Here, by integrating large-scale gene coevolution analysis across the eukaryotic tree with metabolic network mapping, we build a coevolutionary framework of human metabolism. Using this information, we identify unmapped genes in metabolic networks, leading to the identification of C11orf54 as the long-missing decarboxylase of the pentose pathway. These findings explain how L-xylulose is formed in human metabolism and provide insights into an alternative route for sugar conversion linked to L-ascorbate (vitamin C) biosynthesis. More broadly, our approach offers a framework for mapping missing enzymes across species.
Keywords: metabolic networks, coevolution, orphan reactions, protein function identification, computational screening
Abstract
Excretion of L-xylulose is the hallmark of pentosuria, the fourth of Garrod’s inborn errors of metabolism, yet the molecular basis for L-xylulose formation remains unknown. Here, by projecting coevolutionary data for 511,114 orthogroups across 1,929 eukaryotic genomes onto metabolic maps, we screen for unmapped genes in human metabolism. Among these, we show that the DUF1907 domain of C11orf54 catalyzes formation of L-xylulose by establishing a zinc-coordinated Michaelis complex with β-keto-L-gulonate (BKG). The identification of BKG decarboxylase completes the pentose pathway, in which pentose sugars are produced by decarboxylation of nonphosphorylated hexose precursors. The pathway was present in the unicellular ancestor of animals and is conserved in all deuterostomes, in contrast to the alternative L-ascorbate (vitamin C) biosynthesis pathway. An increased flux toward pentoses may have represented an evolutionary tradeoff, favoring energy metabolism and redox cofactor balance at the expense of ascorbate biosynthesis in organisms, such as humans and other Haplorhini primates, where dietary vitamin C intake prevents scurvy.
The metabolism of all living organisms is sustained by enzymes, which catalyze chemical reactions essential for life (1–3). Over the years, advances in biochemistry and genetics have enabled the reconstruction of metabolic networks across species, identifying the enzymes responsible for thousands of biochemical transformations. These efforts have led to comprehensive databases cataloging metabolic reactions and their connections (4–6).
Despite this progress, gaps remain in our understanding of metabolism, even in well-studied organisms like humans. Many metabolic reactions lack a fully characterized enzyme, while others are known only from biochemical activity, with no associated gene sequence (7, 8). These unresolved cases, known as pathway holes (9) or orphan reactions/enzymes, have persisted for decades, limiting our ability to fully understand metabolic networks. Computational approaches, including sequence homology (9), genomic context analysis (10, 11), and structure-based screenings (12, 13), have been employed to address these gaps. In particular, evolutionary analyses reveal that enzymes working in the same pathway often coevolve—that is, their genes tend to be jointly present or absent across species—reflecting their functional interdependence (14–19). This principle can be leveraged to predict missing enzymes in metabolic pathways.
Here, we show that gene coevolution analysis can identify missing metabolic functions by detecting statistically significant patterns of gene cogain and coloss along the eukaryotic tree (20). Applying this approach to the human metabolic network, we uncover coevolutionary hotspots with unresolved reactions, leading to the identification of a key missing enzyme for the formation of L-xylulose in the pentose pathway. This metabolic route, which parallels the pentose phosphate pathway, generates five-carbon sugars through decarboxylation of nonphosphorylated six-carbon sugar acids (21–23). We experimentally validate the molecular identity of the missing decarboxylase, completing a long-standing gap in metabolism. This finding provides insights into the evolution of the pentose pathway and its role in human and metazoan physiology, including its connection to vitamin C biosynthesis.
Coevolution Hotspots in Metabolic Pathways
We performed a coevolution analysis using the cotr program (20) with the OrthoDB (24) v.11 set of orthogroups (n = 511,114) at the Eukaryota level (1,952 species) and obtained 520,770 coevolving orthogroup pairs with an adjusted P-value (p_adj) < 0.05. Of these, 41,258 pairs and 5,686 orthogroups contain at least one corresponding gene/protein in humans.
To integrate these data into metabolic context, we considered reactions in the KEGG “Metabolic pathways” map (rn01100), and assigned a reaction coevolution score (RCS) to each connection as described in the Methods (Fig. 1 and Dataset S1). Briefly, we mapped orthogroups encompassing human genes to KEGG reactions, covering a total of 2,246 reactions and 810 unique orthogroups. We then evaluated, at the gene level, the fraction of linked reactions (shortest path < 3) exhibiting significant cotr scores with the focus reaction (SI Appendix, Fig. S1). This analysis identified 457 human reactions (20%) with a positive RCS, significantly exceeding random expectation (81 ± 8.3 reactions, based on random permutations of reaction/gene associations; SI Appendix, Fig. S1D). Coloring the KEGG map according to RCS highlights distinct regions where coevolution is detected at the gene level, while other areas show little or no coevolution (Fig. 1A). In many cases, this procedure illuminates trajectories corresponding to well-defined metabolic pathways, or KEGG “modules”. Examples include purine biosynthesis, heme biosynthesis, and histidine degradation (Fig. 1B).
Fig. 1.

Coevolution in human metabolism. (A) KEGG map of metabolic pathways highlighted according to gene coevolution; lines represent reactions, circles represent compounds. Lines with a reaction convolution score (RCS)>0 are colored according to KEGG color codes, with a thickness proportional to RCS as indicated in the legend. (B) Map details showing De novo purine biosynthesis (M00048), Heme biosynthesis (M00868), and Histidine degradation (M00045) KEGG pathway modules, contoured with a halo proportional to their RCS values. Genes with an RCS>0 are in black. The start and end products are indicated. PRA: phosphoribosylamine; IMP: inosine monophosphate; 5-Ala: 5-aminolevulinic acid.
Of the 403 total modules in the map, 140 (~35%) were assigned to human metabolism for having a majority of reactions (>60%) mapped to human genes (SI Appendix, Fig. S2A). For these modules, we calculated a Module Coevolution Index (MCI), i.e. the fraction of reactions mapped to human genes with a positive RCS (Dataset S1). We observed a bimodal distribution of MCI values with most modules showing either low or high scores (SI Appendix, Fig. S2B). By assuming a cutoff of 0.5, we classified modules as Coevolution Hotspots (CH, 46 modules, 280 reactions) or Coevolution Coldspots (CC, 94 modules, 456 reactions). We asked whether the correct human gene for these reactions could be assigned based solely on coevolutionary information. In CCs, only 6% of the reactions had the correct gene within the top thirty positions of the cotr ranking. Conversely, in CHs, 67% of reactions had the correct gene within the top five positions, and 42% had it in the very first position (SI Appendix, Fig. S2C).
Bridging Gaps in Coevolution Hotspots
With the aim of applying coevolution data to fill pathway holes (PHs), we systematically classified those present in human metabolism, focusing on reactions connected to other human reactions but lacking KEGG Orthology (KO) records. This is a feature of reactions for which no genes are known in any organisms (global PH), as opposed to cases where the information is missing in some organisms (10). In the different KEGG maps, we identified 511 global PHs corresponding to 498 unique reactions potentially occurring in humans (SI Appendix, Fig. S3 and Dataset S1). We further distinguished between Central Pathway Holes (CPH, n = 201), which represent unassigned reactions surrounded by known human reactions, and Terminal Pathway Holes (TPH, n = 310), representing unassigned reactions connected to known human reactions only by one side (either substrate or product). Additionally, we recorded whether an alternative path exists in humans connecting substrates and products of the unmapped reaction (n = 61), thereby reducing its necessity.
Notably, ten reactions without assigned genes were found in human pathway modules, providing solid evidence that these reactions occur in humans (SI Appendix, Table S1). Four of these reactions are believed to be nonenzymatic, either because they occur spontaneously in solution (25–27) (R04293, R03314, and R07420) or are catalyzed by light (28) (R03311). Conversely, four other reactions have an assigned EC number (EC 4.1.1.34, 3.5.1.63, 3.2.1.103, and 3.1.6.18), providing evidence that a gene encoding an enzyme catalyzing the reaction should be present.
For each PH, we compiled a list of orthogroups ranked according to their significance with orthogroups mapped to nearby reactions (shortest path ≤ 2). Human genes corresponding to orthogroups not already mapped in the same KEGG map were considered potential candidates for the unassigned reaction (Dataset S1). The top-ranking candidates for the ten unassigned reactions found in human pathway modules are reported in SI Appendix, Table S1. Such indications should be regarded with caution due to the presence of false-positive (spurious) coevolutionary associations and the possibility that a coevolutionary link with pathway enzymes may arise from diverse biological functions, including regulation of enzyme activity, metabolite transport, and cofactor biosynthesis. Interestingly, the candidate orthogroup for the unassigned reaction in ubiquinone (coenzyme Q) biosynthesis (29, 30) corresponds to two human genes with unknown molecular function, COQ8A and COQ8B, whose role in ubiquinone biosynthesis is supported by genetic evidence (31, 32).
C11orf54 as the Candidate for a Gap in the Pentose Pathway
Among candidates for the unassigned reactions in human pathways, we focused on C11orf54 (chromosome 11 open reading frame 54), a gene with an unknown function, that was suggested by our analysis to be involved in a decarboxylation reaction (21, 22) (E.C. 4.1.1.34) within the glucuronate/pentose pathway (PP) (23) -a coevolution hotspot module (SI Appendix, Table S1). We chose C11orf54 as a case study because multiple lines of bioinformatics evidence converged to suggest its role in the unassigned reaction (Fig. 2).
Fig. 2.

Bioinformatics evidence linking C11orf54 to the unassigned pentose pathway reaction. (A) Network of significant cotr associations involving C11orf54 and pentose pathway genes (B) Part of the KEGG glucuronate/pentose pathway module highlighting coevolving reactions; the unassigned 3-Dehydro-L-gulonate decarboxylase (EC 4.1.1.34) is marked with an asterisk. (C) Structural overlap between C11orf54 (1XCR, slate blue) and ALDC (4BT2, purple); a zoom-in of the zinc-binding region is shown in the inset, showing Zn2+ coordination bonds (yellow dashes) and C11orf54 residues (sticks). (D) Comparison of the ALDC and 4.1.1.34 reactions highlighting their chemical similarity; the analogous parts of the two β-keto acid substrates are in red. (E) Evidence of functional association between the DUF1907 domain of C11orf54 and 3HCDH domains of CRYL1 provided by domain fusion in rotifers.
In the cotr analysis (Fig. 2A), C11orf54 is the best reciprocal hit for CRYL1, encoding the PP enzyme L-gulonate 3-dehydrogenase (P-value=8 × 10−6). In turn, CRYL1 is significantly associated by coevolution with another PP gene (DCXR). Two other genes of the pathway (SORD, XYLB) are significantly associated with each other but not with CRYL1 or DCXR. The unassigned 4.1.1.34 reaction of the pathway (Fig. 2B and SI Appendix, Fig. S4) involves decarboxylation of β-keto-L-gulonate (3-dehydro-L-gulonate) produced by the CRYL1 reaction to form L-xylulose, a pentose sugar. L-xylulose is then converted into D-xylulose-5P by the subsequent action of L-xylulose reductase (DCXR), sorbitol dehydrogenase (SORD), and xylulokinase (XYLB). An evolutionary link between C11orf54 and PP is supported by CladeOScope (18), where CRYL1 is the first hit when using C11orf54 as a query (SI Appendix, Fig. S5A), but not by STRING (33), where no associations among C11orf54 and PP enzymes are reported by coevolution (co-occurrence) or any other evidence (SI Appendix, Fig. S5 B and C).
The C11orf54 X-ray structure (PDB: 1XCR) has been previously determined at 1.7 Å resolution, revealing a zinc-binding site suggestive of an enzymatic activity (34). Based on structural comparisons with other zinc-containing enzymes, such as α-carbonic anhydrase and carboxypeptidase A2, various substrates were tested yielding only a weak esterase activity with 4-nitrophenyl acetate (PNPA) (34). However, a distant structural homology with α-acetolactate decarboxylase (ALDC, EC: 4.1.1.5) was suggested by A. Murzin through manual inspection of the protein fold (34). Consistently, our FoldSeek (35) and HHPred (36) searches revealed significant similarity between C11orf54 and ALDC. In particular, the C11orf54 structure can be superimposed to the ALDC structure from Brevibacillus brevis with an RMSD of 3.15 Å (Fig. 2C). The active site Zn2+ ion as well as three histidine residues (H266, H268, and H278) involved in metal coordination are well superimposed in the structure alignment. Conversely, two other active-site residues, H217 and R263, correspond to an arginine and a glutamate in ALDC enzymes (Fig. 2C). These residues have been shown to be highly conserved and functionally critical in ALDCs (37). The substitutions in C11orf54 may thus reflect differences in the catalyzed reaction. Notably, there are clear similarities between the ALDC and 4.1.1.34 reactions; both involve the decarboxylation of a β-keto-acid, and the acetolactate substrate of ALDC shows regio- and stereochemical similarities to the β-keto L-gulonate substrate of 4.1.1.34 (Fig. 2D).
Finally, according to the PFAM domain architecture analysis in the InterPro database (38), the C11orf54 domain of unknown function (DUF1907) is primarily present as a single domain, although there are other architectures with additional domains fused in sequence. The most frequent multidomain architecture, found in 28 distinct proteins of rotifers, comprises DUF1907 fused with the 3HCDH and 3HCDH_N domains of CRYL1 (Fig. 2E). Since enzyme fusions often involve genes encoding consecutive enzymes in a metabolic pathway (39) (Fig. 2B), this evidence further supports a role of C11orf54 in the unassigned reaction of the pathway.
C11orf54 Encodes the Pentose Pathway Decarboxylase
We decided to experimentally validate our hypothesis by recombinantly expressing the human C11orf54 coding sequence, encoding the putative β-keto-L-gulonate decarboxylase (BKGD). Additionally, we expressed other two enzymes of the pathway to produce β-keto-L-gulonate (BKG) and to confirm L-xylulose production. CRYL1, C11orf54, and DCXR were expressed and purified in the same manner (see Methods), yielding 70, 20, and 13 mg per liter of culture, respectively, in a pure form (SI Appendix, Fig. S6).
We assayed the activity of C11orf54 on the BKG substrate by 1H NMR (Fig. 3A). Due to the inability to obtain BKG in a pure and stable form, we monitored the reaction starting from L-gulonate obtained through basic hydrolysis of L-gulono-1,4-lactone (see Methods). In 1H NMR kinetic analysis, the CRYL1 oxidation of L-gulonate at C3 is observed as the loss of proton coupling of the C2 hydrogen and appearance of a singlet signal at 3.90 ppm attributable to the anomeric (cyclic) form of BKG. Upon magnification of a region near the water suppression, a singlet signal attributable to the linear keto isomer appears at 4.94 ppm, due to the existence of an equilibrium between open and closed forms of BKG (Fig. 3A). A more complete assignment of BKG protons could be achieved in reactions conducted with an excess of CRYL1 (SI Appendix, Fig. S7A). However, even under this condition, only partial formation of BKG is observed, consistent with the previously described equilibrium, which is shifted toward L-gulonate (21). In the absence of further additions, the BKG signal undergoes a slow decay (t½ = 5.3 h), which can be attributed to the nonenzymatic decarboxylation of the keto isomer (SI Appendix, Fig. S7 B and C). However, upon addition of recombinant C11orf54, the BKG and L-gulonate signals rapidly disappear, accompanied by the appearance of a set of signals which can be assigned to L-xylulose (Fig. 3A) based on correlation spectroscopy (SI Appendix, Fig. S8) and comparison with previous spectral assignments (40). The same signals were observed using DCXR in the reverse reaction with xylitol and NADP+ as substrates, supporting a connection of this enzyme with C11orf54 (SI Appendix, Fig. S9).
Fig. 3.

Recombinant human C11orf54 decarboxylates β-keto-L-gulonate. (A) Example of time-resolved partial 1H NMR spectra showing the conversion of L-gulonate into L-xylulose. The proton assignment of L-gulonate (5 mM) is shown at time 0′ in a reaction buffer (20 mM potassium phosphate, 10% D2O) containing NAD+ (0.5 mM), pyruvate (20 mM), and lactate dehydrogenase (LDH) for NAD+ recycling. After addition of CRYL1 (5 µM), singlet signals attributed to the open and furanose form of BKG are marked with asterisks; the region around the singlet at 4.94 ppm is shown with a 5x intensity enhancement. After addition of C11orf54 (5 µM), formation of L-xylulose spectra is observed, as confirmed by proton assignment; the C2 proton of lactate (L) produced by LDH is indicated (B) Extracted ion chromatograms (XICs) of m/z 193.0354 and 149.0455 from reaction mixtures containing CRYL1 (Bottom) or CRYL1+C11orf54 (Top). The peaks at 2.19 min correspond to BKG (193.0354 [M − H]−) and its in-source decarboxylation product. The peak at 3.36 min corresponds to L-xylulose (149.0455 [M − H]−) produced by enzymatic decarboxylation. XICs of the other reaction components are reported in SI Appendix, Fig. S10B. (C) Nonlinear fitting to the Michaelis-Menten equation of the initial reaction velocity of C11orf54 (0.25 μM), showing its dependency on BKG concentrations. Data are presented as independent experiments for each point.
To further validate the results obtained by 1H NMR, the reaction was also analyzed using UHPLC-HRMS (Fig. 3B and SI Appendix, Fig. S10). Two reaction mixtures, one containing CRYL1 alone and the other containing both CRYL1 and C11orf54, were incubated for 3 h and subsequently analyzed. The extracted ion chromatograms (XICs) for the BKG (m/z 193.0354) and L-xylulose (m/z 149.0455) signals show that, in the presence of CRYL1 alone (Fig. 3 B, Bottom), a peak appears at 2.19 min, consistent with the expected mass of BKG (193.0354 [M − H]−). At the same retention time, a signal corresponding to L-xylulose (149.0455 [M − H]−) is also detected, likely resulting from partial decarboxylation during high-temperature ionization, in agreement with the known instability of BKG. In contrast, when C11orf54 is also present (Fig. 3 B, Top), the peak at m/z = 193.0354 disappears, and a new one at m/z = 149.0455 emerges at 3.36 min, consistent with the enzymatic decarboxylation of BKG to L-xylulose. Mass spectra acquired at the two retention times (SI Appendix, Fig. S10A) confirmed the presence of the corresponding [M − H]− ions.
To characterize the kinetic parameters of C11orf54, we set up a coupled spectrophotometric assay exploiting the equilibrium of the NAD-dependent reaction of CRYL1 (Methods). The small amount of BKG produced by CRYL1 was used as the initial substrate concentration for the C11orf54 kinetics. After the addition of C11orf54, the equilibrium of CRYL1 is shifted, leading to the production of NADH. By analyzing the different slopes of the CRYL1 reaction rate, with CRYL1 used in excess as a reporter, we determined the kinetic constants of C11orf54 (Fig. 3B). The results showed a Km of 20.185 ± 2.505 μM and a kcatof 4.271 ± 0.304 s−1, with a kcat/Km of 211,590 ± 41,341 M−1 s−1. Notably, the turnover number of C11orf54 for BKG is two orders of magnitude higher than that reported for the generic PNPA substrate (34), supporting its physiological role as a BKG decarboxylase. Comparison of the reaction kinetics in the presence and absence of C11orf54, shows that the reaction proceeds slowly even without the activity of C11orf54 (SI Appendix, Fig. S11 A and B). This is due to the spontaneous decarboxylation of BKG, as previously observed by 1H NMR. In contrast, performing the same assay without CRYL1 showed no NADH production as well as using NADP+ as electron acceptor (SI Appendix, Fig. S11C). Based on our functional validation, we propose renaming C11orf54 as BKGD.
Crucial Residues for BKGD Catalysis
To identify BKGD residues involved in catalysis, we performed a docking analysis using the previously resolved crystallographic structure (PDB ID: 1XCR) (34). Docking BKG onto 1XCR with AutoDock Vina (41) (Fig. 4A) revealed multiple possible conformations. One conformation closely resembles the catalytic complex of α-acetolactate decarboxylase (37), with carboxyl, carbonyl, and hydroxyl groups of BKG coordinating the Zn2+ ion (Fig. 4A) and forming a six-coordinate octahedral complex. Additional hydrophilic interactions were observed, including H217 with the substrate C2 oxygen and R263 with both the carbonyl group and C4 oxygen; the carboxyl terminus (D315) interacts with H268, suggesting a role in modulating zinc coordination.
Fig. 4.

BKGD catalytically relevant residues. (A) Human BKGD structure (1XCR) with docked β-keto-L-gulonate (aquamarine, ball-and-stick). Noncarbon atoms follow CPK coloring. Hydrophilic interactions are highlighted with dashed lines for enzyme-zinc (yellow), substrate-zinc (magenta), enzyme–substrate (green), and enzyme–enzyme (orange). Zinc-coordinating residues are in slate blue and residues selected for site-directed mutagenesis in deep purple. (B) Portions of multiple sequence alignment of BKGD orthologs. Colors match the active site residues in (A); asterisks (*) indicate the end of sequences; species are indicated by their common name or genus name; complete sequences and scientific names are provided in SI Appendix, Fig. S12. (C) Bar plot of the catalytic efficiency of BKGD variants. Data points represent kcat/Km values, with error bars indicating SD of the fitting parameters.
A multiple sequence alignment (Fig. 4B and SI Appendix, Fig. S12) revealed that these residues are strictly conserved across metazoan orthologs. Based on this, we introduced three mutations in BKGD: two substitutions (H217N and R263Q) and one extension (D315extG). The substitutions targeted residues with a positive score in the BLOSUM62 matrix altering their amino acid properties while preserving their hydrophilic nature. For the extension, glycine was chosen due to its lack of a side chain, allowing us to investigate the potential role of the C-terminal aspartate.
All three variants were expressed under the same conditions as the wild-type (WT), yielding approximately 20 mg of purified protein per liter of culture (SI Appendix, Fig. S13A). Enzyme activity was assessed using CRYL1, as performed for WT. Circular dichroism analysis (SI Appendix, Fig. S13B) indicated no significant differences in secondary structure between WT and the variants. Functional assays revealed that H217N and R263Q lacked activity with β-keto-L-gulonate, whereas D315extG exhibited a Km of 33.97 ± 7.41 μM and a kcatof 2.71 ± 0.33 s−1, resulting in a catalytic efficiency (kcat/Km) of 79,719 ± 27,068 M−1s−1, lower than that of the WT (Fig. 4C and SI Appendix, Fig. S13C). We conclude that H217 and R263, two residues distinctive to BKGD compared to ALDC, are essential for catalysis, while D315 contributes to catalytic efficiency, providing a rationale for the unusual conservation of the C-terminal residue in the protein family.
The Complete Pentose Pathway and Its Connection to L-Ascorbate Biosynthesis
In various eukaryotes, PP and L-ascorbate biosynthesis pathway (ABP) are alternative branches of the catabolism D-glucuronate (Fig. 5A). This common oxidized form of glucose can be produced by the conversion of UDP-glucose, or the catabolism of myo-inositol (42, 43) and proteoglycans (44). D-glucuronate is reduced to L-gulonate by AKR1A1 (45). L-gulonate is a branching point (Fig. 5A): It can be converted to BKG by the enzyme CRYL1 in PP (46), or to L-gulonolactone by the enzyme RGN in ABP (47). The two possible fates of L-gulonate were identified almost 70 y ago in a study by Grollman and Lehninger (48), where under different conditions, L-gulonate was observed to be processed into L-ascorbate or into L-xylulose and CO2.
Fig. 5.

Evolution of the pentose pathway and its connection with vitamin C biosynthesis. (A) Scheme of the glucuronate degradation with the alternative branches of the pentose pathway (PP, slate blue) and L-ascorbate biosynthesis (ABP, orange) highlighted. Reduced NAD and NADP cofactors are shown in different colors to emphasize the consumption of NADPH and the formation of NADH. (B) The presence of PP (CRYL1, BKGD, DCXR, SORD, XYLB) and ABP (RGN, GULO) genes is shown across representative eukaryotic species with an emphasis on Metazoa (arrow). Highlighted in colors are metazoan species inferred to have lost PP, ABP, or both. The corresponding taxonomic groups are named and illustrated by Phylopic silhouettes (www.phylopic.org). The eukaryote chronogram is according to Timetree (www.timetree.org). Gene presence/absence is according to OrthoDB v.11.
PP enables the production of D-xylulose-5P, an intermediate of the nonoxidative phase of the pentose phosphate pathway (PPP), that can be isomerized to D-ribulose-5P and D-ribose-5P or feed into glycolysis and gluconeogenesis. The initial reactions of PP bear resemblance with the oxidative branch of PPP. In both pathways, the production of a pentose sugar from a hexose derivative involves oxidation of a sugar acid β-carbon followed by decarboxylation of the resulting β-keto acid. While in PPP these reactions are carried out by a single bidomain protein (49) (6-phosphogluconate dehydrogenase), in PP, with the possible exception of rotifers (Fig. 2E), they are catalyzed by two distinct enzymes and proteins with dehydrogenase (CRYL1) and decarboxylase (BKGD) activities (Fig. 5A). The product of the BKGD reaction is substrate of L-xylulose reductase (DCXR). DCXR mutations cause excretion of L-xylulose (50), pentosuria, one of the four Garrod’s inborn errors of metabolism (51). The subsequent reaction catalyzed by xylitol dehydrogenase (SORD) converts xylitol back to xylulose in the opposite (D) configuration, and phosphorylation by xylulokinase (XYLB) produces the common D-xylulose-5-phosphate intermediate. Besides acting on unphosphorylated precursors, a distinctive feature of PP with respect to PPP is that the pathway produces NADH and consumes NADPH (Fig. 5A).
According to the gene distribution across the eukaryotic tree (Fig. 5B), PP is predominantly found in Metazoa, with the exception of certain Alveolata (Perkinsus marinus) and nonmetazoan Opisthokonts. Plants and fungi possess homologs of the SORD and XYLB genes for the conversion of xylitol [which can also derive from the hemicellulose and pectin component L-arabinose (52)], but they typically lack the complete gene set for the conversion of L-gulonate to xylitol (CRYL1, BKGD, and DCXR). By contrast, a full set of these genes is found in extant lineages of the unicellular ancestors of animals (e.g. Capsaspora owczarzaki, Monosiga brevicollis) and in basal animals (Amphimedon queenslandica). Notably, PP genes were concurrently lost in some protostome lineages (Fig. 5B and SI Appendix, Fig. S14), such as Placozoa (Trichoplax adhaerens), nematodes, and trematodes, providing the coevolutionary signal detected by our analysis. In contrast, we found no instances of PP loss in any deuterostome species (Fig. 5B). The strict conservation of PP genes in deuterostomes also pertains to BKGD, despite the nonenzymatic decarboxylation of BKG (Fig. 2C).
Although ABP has a more widespread presence—being also found in fungi and plants—it has been lost multiple times in animals (23), particularly among several deuterostome lineages such as fishes (Teleostei), birds (Passeriformes), bats (Yangochiroptera), and primates (Haplorrhini). The absence of ascorbate biosynthesis from glucuronate in humans and other animals is supported by numerous experimental studies (53–56) and generally involves the loss of GULO rather than RGN (Fig. 5B). In contrast to RGN, which also participates in other cellular functions (57), GULO has a specialized role in different pathways for vitamin C biosynthesis (58). A metabolic consequence of ABP loss in animal species is that all glucuronate is directed toward PP and the production of D-xylulose-5P.
Discussion
We performed a computational screening to identify unmapped genes in metabolic pathways, focusing on gaps within coevolution hotspots. One such gap, the decarboxylase of the pentose pathway, had remained unresolved for 65 y. Our screening identified C11orf54 as a candidate gene, supported by multiple bioinformatics approaches. Biochemical validation confirmed C11orf54 as BKGD and revealed key catalytic residues. These findings complete the pentose pathway and refine our understanding of its metabolic and evolutionary significance.
In humans, C11orf54 is predominantly expressed in the kidneys and liver, corresponding with early activity assays performed on these tissue extracts (21, 22). The other PP genes, particularly CRYL1, have similar expression patterns (SI Appendix, Fig. S15). Previous studies have reported phenotypes for the knockdown of C11orf54 in human cells and of the orthologous meep gene in Drosophila melanogaster. Notably, C11orf54 silencing inhibited the PLC/PRF/5 and 293 T cell growth (59), while the meep knockdown caused hyperglycemia and reduced growth in larvae reared on high-sugar diet (60). In addition, a link to DNA repair has been reported for C11orf54 (59), hinting at a possible moonlighting activity. According to data from the International Mouse Phenotyping Consortium (www.mousephenotype.org), the knockout mouse (MGI:1918234) is significantly impaired in cardiovascular function and eye morphology. A possible connection of PP with ocular function in certain mammals is supported by the role of the pathway intermediate xylitol, along with other polyols, in eye defects such as cataracts (61); as well as the presence of CRYL1 (crystallin lambda 1) in rabbit lenses (46).
The C11orf54 family tree in the PANTHER database (62) shows that the gene is essentially single-copy in metazoans, with the exception of a few species (e.g., Drosophila, Xenopus) that display closely related in-paralogs likely arising from recent duplications. Gene distribution along the eukaryotic tree suggests that PP was present in the ancestors of animals but not in fungi, consistent with the experimental evidence of a different pathway in fungi for D-glucuronate catabolism (63, 64). In metazoans, a source of D-glucuronate are extracellular matrix proteoglycans, which are abundant in connective tissues, cartilage, and basement membranes. These proteoglycans undergo turnover and remodeling (65, 66), releasing sugar acids that can feed into PP.
Since individuals with pentosuria do not show significant metabolic impairments (50), PP was historically considered a minor or redundant pathway (23). This view, however, contrasts with the strict conservation of the pathway in nearly all animals besides Placozoa and parasitic worms. In comparison, ABP, despite producing a crucial metabolic cofactor, has been lost multiple times in evolution, especially in vertebrates (Fig. 5B). A possible explanation that reconciles evolutionary conservation and the phenotypes observed in humans and Drosophila is that PP plays a broader role in sugar metabolism and energy homeostasis. As a route using an alternative sugar source with respect to PPP, this pathway provides advantages under conditions of limited glucose availability. Additionally, PP provides a metabolic counterbalance of redox cofactors, consuming NADPH and regenerating NADP+ (Fig. 5A), thus sustaining pentose production under high NADPH/NADP+ ratio, a condition that feedback-inhibits PPP (67). Since PP and ABP serve as alternative branches of D-glucuronate metabolism, these advantages can explain why ascorbate biosynthesis has been lost in multiple lineages. The loss of ABP would increase PP flux, reinforcing its role in pentose production and redox balance, thereby compensating for the lack of ascorbate synthesis in organisms where dietary vitamin C intake can prevent scurvy.
A relevant result of our analysis is that in well-defined pathways of human metabolism, a limited number of reactions, some of which may not require an enzymatic activity, remain without an assigned gene (SI Appendix, Table S1). The number of potential holes in human metabolism is much larger (498 reactions, SI Appendix and Dataset S1) if one considers also unassigned reactions not included in KEGG modules but connected to human reactions in the metabolic map; there is less confidence, however, that these reactions have a physiological meaning in humans. A comparable figure is reported in the Metacyc database, which lists 402 pathway holes in humans (https://biocyc.org/HUMAN/missing-rxns.html).
We note that, in addition to identifying previously unmapped genes in metabolic networks, our analysis can be useful to provide independent support for gene-reaction assignments that lack experimental validation. A relevant example is the AMDHD1 (amidohydrolase domain containing 1) gene, whose assignment to histidine degradation based on the presence of the imidazolonepropionase domain (68) is strongly supported by coevolution with other genes of the pathway (Fig. 1B).
Although human metabolism was the focus of our work, the same procedure can be easily adapted to other organisms, with the potential to shed light on metabolic pathways present in less explored branches of the tree of life. We also note that our analysis was limited to global PHs, that is reactions for which no genes are known in any species, as opposed to reactions for which a gene is known in some organisms but not in humans (local PHs). A straightforward extension of our procedure would be to apply a similar analysis to local PHs in humans or other species.
Methods
Cotr Analysis and KEGG Mapping.
Cotr analysis was performed with the previously published pipeline (https://github.com/lab83bio/Cotransitions) applied to the OrthoDB database v. 11 at Eukaryota level (at2759) using a right-ladderized tree of 1,952 species. To map the orthogroups onto the KEGG PATHWAY database, we first identified the corresponding human genes using the odb11_OG2genes file. Then, we used the human proteome from UniProt to retrieve the KEGG ID of each human gene (hsa:). Finally, the KEGG IDs were mapped using the KEGG API to obtain the reactions to be analyzed in the KEGG PATHWAY maps. Code is available in Source Data (https://doi.org/10.6084/m9.figshare.29371550).
RCS and MCI Calculation.
For each reaction in different KEGG maps, we calculate a RCS corresponding to the fraction of significant cotr scores between the orthogroup mapped to the focus reaction and orthogroups mapped to linked nearby reactions (LNR), that is linked reactions within two steps in the metabolic map (SI Appendix, Fig. S1). Reactions with an RCS > 0 were considered coevolving reactions. For Kegg modules, we calculated a MCI as the fraction of mapped orthogroups with an RCS > 0 considering only comparisons among within-module reactions. Modules with an MCI ≥ 0.5 were considered coevolution hotspots, whereas modules with an MCI < 0.5 were considered coevolution coldspots. Code is available in Source Data (https://doi.org/10.6084/m9.figshare.29371550).
Pathway Holes and Candidate Genes.
We considered potential human pathway hole (PH) reactions lacking the KO identifier flanked by reactions with hsa identifier within one step. PHs surrounded by reactions with hsa identifiers both upstream (toward the reagents) and downstream (toward the products) within two steps were considered central pathway holes (CPH), whereas the others were considered terminal pathway holes (TPH). To provide an indication on the coevolution score of the pathway hole neighbors, pathway holes inherit the RCS of the LNR with the highest RCS (SI Appendix, Fig. S3). To select candidate orthogroups for pathway holes, we selected up to 50 unmapped orthogroups ranked by their significance with respect to LNR orthogroups. For each of the unmapped orthogroups, we then assessed the ranking of the most significant LNR orthogroup and selected as the best hit the candidate with the lowest combined rank. Code is available in Source Data (https://doi.org/10.6084/m9.figshare.29371550).
Docking Analysis.
Docking of BKG against BKGD was performed using AutoDock Vina 1.2.5 (41) with exhaustiveness = 100. The BKG ligand was constructed using 3D coordinates from PubChem (CID: 439273) and protonated at pH 7.4 with Avogadro 1.2.0. The corresponding pdbqt file was generated with the prepare_ligand command from AutoDock4.2. For the protein, the structure from PDB entry 1XCR, corresponding to human BKGD, was used. The necessary pdbqt file for the protein was obtained using the AutoDock4Zn protocol (69), which is suitable for zinc-dependent proteins. Of the nine conformations obtained, the most reliable one in terms of Zn2+ coordination (37) was selected. Docking output was visually inspected with PyMOL 2.5.0 Open-Source.
Bioinformatics.
Sequences for multiple sequence alignment were retrieved from orthogroup 2654314at2759 of OrthoDB v.11, selected to ensure broad species representation across the metazoan tree. Alignment was performed using Clustal Omega 1.2.3 and visualized with ESPript 3.0. Cytoscape 3.9.1 was used to generate protein interaction networks. Figures of protein structure and docking results were generated with PyMOL 2.5.0 Open-Source.
Expression and Purification.
Human CRYL1, C11orf54, and DCXR sequences were cloned into the pET28-TEV vector by GenScript. BL21 CodonPlus cells were transformed via electroporation and selected on agar plates. Expression was conducted for all constructs in autoinducing LB broth, prepared by adding 0.5 g/L glucose and 2 g/L lactose to standard LB medium. A single colony of the clone was inoculated into 1 L of autoinducing LB broth, with a preinduction phase at 37 °C for 8 h, followed by 20 °C for 16 h. Site-directed mutants of C11orf54 were purchased by GenScript, and the protein production was obtained as WT protocol.
Cell lysis was performed using sonication, and for sample preparation after overnight induction, the culture was centrifuged, and the pellets were resuspended in 30 mL of PBS. Afterward, the cell lysate was centrifuged again at 20,000×g for 40 min at 4 °C. All three proteins were successfully found in the soluble fraction, so the supernatant was loaded onto a 50 mL Superloop of the ÄKTA pure FPLC system and purified by Affinity Chromatography (AC) using a HisTrap 5 mL FF column.
During purification, Washing Buffer (25 mM TEA, 150 mM NaCl, and 1 mM DTT) and Elution Buffer (25 mM TEA, 150 mM NaCl, 500 mM imidazole, and 1 mM DTT) were used to elute the proteins. Protein fractions were collected and concentrated using Vivaspin™ centrifugation, then dialyzed in the Washing Buffer.
Biochemical Validation.
CRYL1 substrate L-Gulonate has been produced from L-gulono-1,4-lactone powder purchased by SIGMA by a basic hydrolysis reaction as described previously (66). Briefly, for a final solution of L-gulonate at 200 mM concentration, the sugar was resuspended in deionized H2O by adding 200 mM NaOH. After 1 h RT, 200 mM HCl is added to neutralize the basic pH.
Enzymatic assays were performed using a JASCO V-750 spectrophotometer. A coupled assay was designed to assess enzyme activity by leveraging the reaction catalyzed by the first purified enzyme, CRYL1, which generates the substrate for C11orf54. The reaction rate was determined by measuring absorbance at 340 nm in a quartz cuvette at 37 °C, corresponding to NADH formation during the first reaction. The reaction mixture contained 20 mM potassium phosphate (KP), 0.2 mM NAD+, and varying concentrations of L-gulonate (0.05 to 3.2 mM) as the baseline. The reaction was initiated by adding 10 μM CRYL1. After approximately 150 s, once CRYL1 reached equilibrium, 0.25 μM C11orf54 was added, and absorbance was recorded to determine initial velocities using an extinction coefficient of ε340 = 6,220 M−1 cm−1. Velocities were plotted using β-keto-L-gulonate, produced by CRYL1 in the first step, as the initial substrate concentration for C11orf54. Since β-keto-L-gulonate is equimolar to the NADH produced, its concentration was inferred using ε340 = 6,220 M−1 cm−1. Kinetic parameters were obtained by fitting the data to the Michaelis–Menten equation using nonlinear regression. The same procedure was applied to characterize the D315extG variant. Code is available in Source Data (https://doi.org/10.6084/m9.figshare.29371550).
The formation of reaction products was monitored using 1H NMR spectra, obtained with a BRUKER AVANCE 400 and a JEOL ECZ600R NMR spectrometer. The samples were prepared in NMR tubes by dissolving them in 600 µL of a 9:1 mixture of H2O and D2O. NMR experiments were carried out at 25 °C in potassium phosphate buffer 20 mM at pH 7.9 to prevent interference from signals of organic buffers in the 1H NMR spectra. Water suppression was achieved both via excitation sculpting (BRUKER AVANCE 400) and presaturation (JEOL ECZ600R), 1D spectra were recorded with an acquisition time of 2 s, a recycle delay of 5 s and 32 to 64 transient. 2D COSY experiment was acquired with 8 scans, 128 points in the indirect dimension and 1.5 s as relaxation delay. All the spectra were processed and analyzed using MestReNova software, version 15.0.0.
C11orf54 reaction was monitored in UHPLC-HRMS in negative polarity. We applied UHPLC coupled to electrospray ionization quadrupole time-of-flight high-resolution mass spectrometry (ESI–qTOF-HRMS) to identify the L-xylulose formation by C11orf54. Samples were analyzed by ESI–qTOF-HRMS on a X500B QTOF system (SCIEX) equipped with the Twin Sprayer ESI probe coupled to an ExionLC system (SCIEX) controlled by SCIEX OS software 3.0.0. Injection volume was 10 µL. Chromatographic separation was carried out with a Luna Omega SUGAR (150 mm length × 4.6 mm diameter, 3 µm particle size; Phenomenex). The mobile phase consisted of water A) and acetonitrile B) both including 0.1% (v/v) NH4OH using an isocratic elution of 75% B. Flow rate was set at 0.3 mL min−1. MS detection parameters were set as follows: curtain gas 30 psi, ion source gas 1 45 psi, ion source gas 2 55 psi, and temperature 350 °C. The full scan range of m/z 50 to 2,000 was monitored in negative mode, with spray voltage of −4,500 V, declustering potential of −60 V, and collision energy of −10 V. Mass calibration was performed with the ESI Negative Calibration solution for the SCIEX X500 system (SCIEX) before experiments. XIC were extracted with the theoretical mass of the compound ± 0.0001 Da and Gaussian smoothed with 2 points.
Supplementary Material
Appendix 01 (PDF)
Dataset S01 (XLSX)
Acknowledgments
We thank Andrea Fuso, Massimo Golinelli, Giulia Mori, and Davide Cavazzini for their assistance and Alessio Peracchi for discussion. We thank Centro Grandi Strumenti of the University of Pavia and Barbara Mannucci for expert technical assistance with Ultra-High-Performance Liquid Chromatography–High-Resolution Mass Spectrometry (UHPLC-HRMS) analyses. This work benefited from the equipment and framework of the COMP-HUB and COMP-R initiatives, funded by the “Departments of Excellence” program of the Italian Ministry for University and Research (MIUR, 2018-2022 and MUR, 2023-2027), and from the High Performance Computing facility of the University of Parma, Italy. This work was supported by the Italian Ministry for Education, University and Research Progetti di Ricerca di Rilevante Interesse Nazionale (Research Projects of National Interest) Grant 2017483NH8 to R.P.
Author contributions
M.M. and R.P. designed research; M.M., C.D.R., F.G., G.M., D.D., and G.Q. performed research; M.M., C.D.R., G.M., F.S., and R.P. analyzed data; and M.M. and R.P. wrote the paper.
Competing interests
The authors declare no competing interest.
Footnotes
This article is a PNAS Direct Submission.
Contributor Information
Marco Malatesta, Email: marco.malatesta@unipr.it.
Riccardo Percudani, Email: riccardo.percudani@unipr.it.
Data, Materials, and Software Availability
Data generated in this study are available at the Figshare repository (https://doi.org/10.6084/m9.figshare.29371550) (70).
Supporting Information
References
- 1.Benkovic S. J., Hammes-Schiffer S., A perspective on enzyme catalysis. Science 301, 1196–1202 (2003). [DOI] [PubMed] [Google Scholar]
- 2.Bar-Even A., et al. , The moderately efficient enzyme: Evolutionary and physicochemical trends shaping enzyme parameters. Biochemistry 50, 4402–4410 (2011). [DOI] [PubMed] [Google Scholar]
- 3.Ribeiro A. J. M., Riziotis I. G., Borkakoti N., Thornton J. M., Enzyme function and evolution through the lens of bioinformatics. Biochem. J. 480, 1845–1863 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 4.Kanehisa M., Furumichi M., Sato Y., Matsuura Y., Ishiguro-Watanabe M., KEGG: Biological systems database as a model of the real world. Nucleic Acids Res. 53, D672–D677 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 5.Caspi R., et al. , The MetaCyc database of metabolic pathways and enzymes and the BioCyc collection of pathway/genome databases. Nucleic Acids Res. 42, D459–D471 (2014). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 6.Milacic M., et al. , The reactome pathway knowledgebase 2024. Nucleic Acids Res. 52, D672–D678 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 7.Lespinet O., Labedan B., Orphan enzymes?. Science 307, 42–42 (2005). [DOI] [PubMed] [Google Scholar]
- 8.Hanson A. D., Pribat A., Waller J. C., de Crécy-Lagard V., ‘Unknown’ proteins and ‘orphan’ enzymes: The missing half of the engineering parts list – and how to find it. Biochem. J. 425, 1–11 (2010). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 9.Green M. L., Karp P. D., A bayesian method for identifying missing enzymes in predicted metabolic pathway databases. BMC Bioinformatics 5, 76 (2004). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 10.Osterman A., Missing genes in metabolic pathways: A comparative genomics approach. Curr. Opin. Chem. Biol. 7, 238–251 (2003). [DOI] [PubMed] [Google Scholar]
- 11.Schläpfer P., et al. , Genome-wide prediction of metabolic enzymes, pathways, and gene clusters in plants. Plant Physiol. 173, 2041–2059 (2017). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 12.Hadadi N., MohammadiPeyhani H., Miskovic L., Seijo M., Hatzimanikatis V., Enzyme annotation for orphan and novel reactions using knowledge of substrate reactive sites. Proc. Natl. Acad. Sci. 116, 7298–7307 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 13.Malatesta M., et al. , One substrate many enzymes virtual screening uncovers missing genes of carnitine biosynthesis in human and mouse. Nat. Commun. 15, 3199 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 14.Tatusov R. L., Koonin E. V., Lipman D. J., A genomic perspective on protein families. Science 278, 631–637 (1997). [DOI] [PubMed] [Google Scholar]
- 15.Pellegrini M., Marcotte E. M., Thompson M. J., Eisenberg D., Yeates T. O., Assigning protein functions by comparative genome analysis: Protein phylogenetic profiles. Proc. Natl. Acad. Sci. 96, 4285–4288 (1999). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 16.Barker D., Pagel M., Predicting functional gene links from phylogenetic-statistical analyses of whole genomes. PLoS Comput. Biol. 1, e3 (2005). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 17.Dey G., Jaimovich A., Collins S. R., Seki A., Meyer T., Systematic discovery of human gene function and principles of modular organization through phylogenetic profiling. Cell Rep. 10, 993–1006 (2015). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 18.Tsaban T., et al. , CladeOScope: Functional interactions through the prism of clade-wise co-evolution. NAR Genomics Bioinforma. 3, lqab024 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 19.Moi D., Dessimoz C., Phylogenetic profiling in eukaryotes comes of age. Proc. Natl. Acad. Sci. U. S. A. 120, e2305013120 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 20.Dembech E., et al. , Identification of hidden associations among eukaryotic genes through statistical analysis of coevolutionary transitions. Proc. Natl. Acad. Sci. U. S. A. 120, e2218329120 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 21.Smiley J. D., Ashwell G., Purification and properties of β-l-hydroxy acid dehydrogenase. J. Biol. Chem. 236, 357–364 (1961). [Google Scholar]
- 22.Winkelman J., Ashwell G., Enzymic formation of l-xylulose from β-keto-l-gulonic acid. Biochim. Biophys. Acta. 52, 170–175 (1961). [DOI] [PubMed] [Google Scholar]
- 23.Linster C. L., Van Schaftingen E., Vitamin C: Biosynthesis, recycling and degradation in mammals. FEBS J. 274, 1–22 (2007). [DOI] [PubMed] [Google Scholar]
- 24.Kuznetsov D., et al. , OrthoDB v11: Annotation of orthologs in the widest sampling of organismal diversity. Nucleic Acids Res. 51, D445–D451 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 25.Colabroy K. L., Begley T. P., The pyridine ring of NAD is formed by a nonenzymatic pericyclic reaction. J. Am. Chem. Soc. 127, 840–841 (2005). [DOI] [PubMed] [Google Scholar]
- 26.Vogel H. J., Davis B. D., Glutamic γ-semialdehyde and Δ1 -pyrroline-5-carboxylic acid, intermediates in the biosynthesis of proline 1, 2. J. Am. Chem. Soc. 74, 109–112 (1952). [Google Scholar]
- 27.Borsook H., Dubnoff J. W., The hydrolysis of phosphocreatine and the origin of urinary creatinine. J. Biol. Chem. 168, 493–510 (1947). [PubMed] [Google Scholar]
- 28.Holick M. F., et al. , Photosynthesis of previtamin D3 in human skin and the physiologic consequences. Science 210, 203–205 (1980). [DOI] [PubMed] [Google Scholar]
- 29.Nicoll C. R., et al. , In vitro construction of the COQ metabolon unveils the molecular determinants of coenzyme Q biosynthesis. Nat. Catal. 7, 148–160 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 30.Pelosi L., et al. , COQ4 is required for the oxidative decarboxylation of the C1 carbon of coenzyme Q in eukaryotic cells. Mol. Cell 84, 981–989.e7 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 31.Lagier-Tourenne C., et al. , ADCK3, an ancestral kinase, is mutated in a form of recessive ataxia associated with coenzyme Q10 deficiency. Am. J. Hum. Genet. 82, 661–672 (2008). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 32.Mollet J., et al. , CABC1 gene mutations cause ubiquinone deficiency with cerebellar ataxia and seizures. Am. J. Hum. Genet. 82, 623–630 (2008). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 33.Szklarczyk D., et al. , The STRING database in 2023: Protein–protein association networks and functional enrichment analyses for any sequenced genome of interest. Nucleic Acids Res. 51, D638–D646 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 34.Manjasetty B. A., et al. , Crystal structure of Homo sapiens PTD012 reveals a zinc-containing hydrolase fold. Protein Sci. 15, 914–920 (2006). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 35.Kempen M., et al. , Fast and accurate protein structure search with Foldseek. Nat. Biotechnol. 42, 243–246 (2023). 10.1038/s41587-023-01773-0. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 36.Zimmermann L., et al. , A completely reimplemented MPI bioinformatics toolkit with a new hhpred server at its core. J. Mol. Biol. 430, 2237–2243 (2018). [DOI] [PubMed] [Google Scholar]
- 37.Marlow V. A., Rea D., Najmudin S., Wills M., Fülöp V., Structure and mechanism of acetolactate decarboxylase. ACS Chem. Biol. 8, 2339–2344 (2013). [DOI] [PubMed] [Google Scholar]
- 38.Blum M., et al. , Interpro: The protein sequence classification resource in 2025. Nucleic Acids Res. 53, gkae1082 (2024). 10.1093/nar/gkae1082. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 39.Henry C. S., et al. , Systematic identification and analysis of frequent gene fusion events in metabolic pathways. BMC Genomics 17, 473 (2016). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 40.Abdoul-Zabar J., et al. , Thermostable transketolase from Geobacillus stearothermophilus: Characterization and catalytic properties. Adv. Synth. Catal. 355, 116–128 (2013). [Google Scholar]
- 41.Eberhardt J., Santos-Martins D., Tillack A. F., Forli S., AutoDock vina 1.2.0: New docking methods, expanded force field, and python bindings. J. Chem. Inf. Model. 61, 3891–3898 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 42.Hankes L. V., Politzer W. M., Touster O., Anderson L., Myo-inositol catabolism in human pentosurics: The predominant role of the glucuronate-xylulose-pentose phosphate pathway. Ann. N. Y. Acad. Sci. 165, 564–576 (1970). [PubMed] [Google Scholar]
- 43.Thorsell A.-G., et al. , Structural and biophysical characterization of human myo-inositol oxygenase. J. Biol. Chem. 283, 15209–15216 (2008). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 44.Ray J., Wu Y., Aguirre G. D., Characterization of ß-glucuronidase in the retinal pigment epithelium. Curr. Eye Res. 16, 131–143 (1997). [DOI] [PubMed] [Google Scholar]
- 45.Barski O. A., Gabbay K. H., Grimshaw C. E., Bohren K. M., Mechanism of human aldehyde reductase: Characterization of the active site pocket. Biochemistry 34, 11264–11275 (1995). [DOI] [PubMed] [Google Scholar]
- 46.Ishikura S., Structural and functional characterization of rabbit and human l-gulonate 3-dehydrogenase. J. Biochem. (Tokyo) 137, 303–314 (2005). [DOI] [PubMed] [Google Scholar]
- 47.Aizawa S., et al. , Structural basis of the γ-lactone-ring formation in ascorbic acid biosynthesis by the senescence marker protein-30/gluconolactonase. PLoS ONE 8, e53706 (2013). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 48.Grollman A. P., Lehninger A. L., Enzymic synthesis of l-ascorbic acid in different animal species. Arch. Biochem. Biophys. 69, 458–467 (1957). [DOI] [PubMed] [Google Scholar]
- 49.Adams M. J., Ellis G. H., Gover S., Naylor C. E., Phillips C., Crystallographic study of coenzyme, coenzyme analogue and substrate binding in 6-phosphogluconate dehydrogenase: Implications for NADP specificity and the enzyme mechanism. Structure 2, 651–668 (1994). [DOI] [PubMed] [Google Scholar]
- 50.Pierce S. B., et al. , Garrod’s fourth inborn error of metabolism solved by the identification of mutations causing pentosuria. Proc. Natl. Acad. Sci. U. S. A. 108, 18313–18317 (2011). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 51.Garrod A. E., The croonian lectures on inborn errors of metabolism. Lancet 172, 1–7 (1908). [Google Scholar]
- 52.Chiang C., Knight S. G., A new pathway of pentose metabolism. Biochem. Biophys. Res. Commun. 3, 554–559 (1960). [DOI] [PubMed] [Google Scholar]
- 53.Baron J. H., Sailors’ scurvy before and after James Lind - A reassessment. Nutr. Rev. 67, 315–332 (2009). [DOI] [PubMed] [Google Scholar]
- 54.Drouin G., Godin J.-R., Page B., The genetics of vitamin C loss in vertebrates. Curr. Genomics 12, 371–378 (2011). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 55.Massie H. R., Shumway M. E., Whitney S. J. P., Sternick S. M., Aiello V. R., Ascorbic acid in Drosophila and changes during aging. Exp. Gerontol. 26, 487–494 (1991). [DOI] [PubMed] [Google Scholar]
- 56.Patananan A. N., et al. , The invertebrate Caenorhabditis elegans biosynthesizes ascorbate. Arch. Biochem. Biophys. 569, 32–44 (2015). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 57.Kondo Y., et al. , Senescence marker protein 30 functions as gluconolactonase in L-ascorbic acid biosynthesis, and its knockout mice are prone to scurvy. Proc. Natl. Acad. Sci. U. S. A. 103, 5723–5728 (2006). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 58.Boverio A., et al. , Structure, mechanism, and evolution of the last step in vitamin C biosynthesis. Nat. Commun. 15, 4158 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 59.Tan J., et al. , C11orf54 promotes DNA repair via blocking CMA-mediated degradation of HIF1A. Commun. Biol. 6, 606 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 60.Pereira M. T., Brock K., Meep a novel regulator of insulin signaling, supports development and insulin sensitivity via maintenance of protein homeostasis in Drosophila melanogaster. G3 10, 4399–4410 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 61.Goode D., Lewis M. E., Crabbe M. J. C., Accumulation of xylitol in the mammalian lens is related to glucuronate metabolism. FEBS Lett. 395, 174–178 (1996). [DOI] [PubMed] [Google Scholar]
- 62.Thomas P. D., et al. , PANTHER: Making genome-scale phylogenetics accessible to all. Protein Sci. 31, 8–22 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 63.Kuivanen J., Richard P., NADPH-dependent 5-keto-D-gluconate reductase is a part of the fungal pathway for D-glucuronate catabolism. FEBS Lett. 592, 71–77 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 64.Kuivanen J., Sugai-Guérios M. H., Arvas M., Richard P., A novel pathway for fungal D-glucuronate catabolism contains an L-idonate forming 2-keto-L-gulonate reductase. Sci. Rep. 6, 26329 (2016). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 65.Kresse H., Schönherr E., Proteoglycans of the extracellular matrix and growth control. J. Cell. Physiol. 189, 266–274 (2001). [DOI] [PubMed] [Google Scholar]
- 66.Matsubayashi Y., Dynamic movement and turnover of extracellular matrices during tissue development and maintenance. Fly (Austin) 16, 248–274 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 67.Baquer N. Z., Hothersall J. S., Mclean P., “Function and Regulation of the Pentose Phosphate Pathway in Brain. in Current Topics in Cellular Regulation, (Elsevier, 1988), vol. 29, pp. 265–289. [PubMed] [Google Scholar]
- 68.MacDougall A., et al. , UniRule: A unified rule resource for automatic annotation in the UniProt Knowledgebase. Bioinformatics 36, 4643–4648 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 69.Santos-Martins D., Forli S., Ramos M. J., Olson A. J., AutoDock4(Zn): An improved AutoDock force field for small-molecule docking to zinc metalloproteins. J. Chem. Inf. Model. 54, 2371–2379 (2014). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 70.Malatesta M., Percudani R., C11orf54 catalyzes L-xylulose formation in human metabolism. Figshare. 10.6084/m9.figshare.29371550.v3. Deposited 23 June 2025. [DOI] [PMC free article] [PubMed]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Appendix 01 (PDF)
Dataset S01 (XLSX)
Data Availability Statement
Data generated in this study are available at the Figshare repository (https://doi.org/10.6084/m9.figshare.29371550) (70).
