Skip to main content
Journal of Genetic Engineering & Biotechnology logoLink to Journal of Genetic Engineering & Biotechnology
. 2026 Jun 23;24(3):100731. doi: 10.1016/j.jgeb.2026.100731

In silico genome mining and characterization of putative horse feces–derived bacterial phytases as potential monogastric animal feed additive candidates

Olyad Erba Urgessa a,b, Mesfin Tafesse Gemeda c, Hunduma Dinka b, Firayad Girma Abdi a, Ketema Tafess b,d,⁎
PMCID: PMC13316621  PMID: 42749412

Abstract

Phytic acid exerts a significant antinutritional effect in poultry, swine, and fish, which can be mitigated by supplementing monogastric feeds with efficient microbial phytases. Accordingly, mining bacterial genomes for novel phytases represents a strategic computational approach to identifying candidates for improving monogastric animal nutrition. In this study, 162 bacterial genomes associated with horse feces were systematically mined using an in silico pipeline to identify and characterize putative phytases.A total of 69 non-redundant sequences were identified and classified as histidine acid phytase (HAPhy) or protein tyrosine phosphatase-like phytase (PTPLPhy). HAPhys were detected in the genomes of Escherichia coli, Klebsiella pneumoniae, Salmonella enterica, Acinetobacter baumannii, and Cutibacterium equinum, whereas PTPLPhys were found in K. pneumoniae, Limosilactobacillus reuteri, Pediococcus acidilactici, Bifidobacterium pseudolongum, and Prescottella equi. Principal component analysis identified glucose-1-phosphatase (CAJ1242485.1) and bifunctional acid phosphatase (NHR17779.1) as the HAPhy candidates exhibiting the most favorable predicted physicochemical properties for potential feed applications. Similarly, among the PTPLPhys, protein tyrosine phosphatase (UNQ40438.1) and a hypothetical protein (CAJ1246072.1) showed the most favorable computational profiles. Biosafety analysis identified potential virulence factors, indicating that sources should be screened prior to feed application. High-quality AlphaFold2 models were obtained for these phytases (90.9–97.2). Molecular docking analysis showed that NHR17779.1 exhibited the strongest binding to phytic acid, whereas CAJ1246072.1 demonstrated the weakest interaction. Overall, this study identifies the horse fecal microbiota as a diverse source of putative phytases that may serve as promising targets for genetic and protein engineering; however, further in vitro and in vivo studies are essential to validate the enzymatic activity and industrial efficacy of these computational candidates.

Keywords: Bacterial phytase, Genome, Histidine acid phytases, Horse feces, Poultry, Physicochemical properties

Graphical abstract

Unlabelled Image

1. Introduction

Phytases are key enzymes used in monogastric animal nutrition to hydrolyze phytic acid. Because poultry and other monogastric species lack sufficient endogenous phytase, supplementation with microbial phytases has become essential for improving phosphorus availability, reducing feed costs, and minimizing environmental pollution.1 Phytic acid, also known as myo-inositol hexakisphosphate, is widely recognized as one of the most significant antinutritional factors in monogastric animal nutrition.2 It occurs in various plant tissues but is especially concentrated in storage organs where phosphorus is deposited in the form of phytate. Seeds contain the highest concentrations of phytate because these tissues function as nutrient reserves for germination. Other plant parts such as roots, tubers, fruits, and leafy vegetables contain much lower levels of phytic acid and therefore contribute less to dietary phytate intake.3

The antinutritional effect of phytic acid is primarily due to its strong ability to bind essential minerals including calcium, zinc, iron, and magnesium. It also forms complexes with dietary proteins and digestive enzymes, which reduces protein digestibility and limits nutrient utilization. As a result, diets with high phytate content can negatively affect growth performance, mineral absorption, and feed efficiency in monogastric animals.2, 4

Monogastric animals such as poultry, pigs, and fish have a single chambered stomach and depend mainly on their own digestive enzymes for nutrient breakdown. Unlike ruminants, they do not possess an extensive microbial fermentation system in the upper digestive tract that can degrade phytate effectively.5 Their natural phytase activity is extremely low, which limits their ability to hydrolyze phytic acid and release the phosphorus bound within it.1 As a result, a large portion of dietary phytate passes through the digestive system unmetabolized, leading to reduced nutrient availability and increased phosphorus excretion. The excretion of undigested phytate phosphorus contributes to environmental problems such as soil phosphorus accumulation and eutrophication of water bodies.6

Although commercial phytases are already used in animal feed, the discovery of new microbial phytases with improved biochemical stability, catalytic efficiency, and structural robustness remains an important research priority.7 Microorganisms from diverse ecological niches, including those associated with animal digestive systems, represent promising sources of novel phytase enzymes.8, 9 Horse (Equus caballus) feces-associated bacteria, in particular, provide a largely unexplored microbial community with potential for phytase production.

Horses are monogastric herbivores that consume a variety of plant materials, including feed ingredients such as wheat bran that are naturally rich in phytic acid. They rely on a large and diverse intestinal microbial community capable of degrading fibrous plant components.5, 10As hindgut fermenters, horses possess an extensive microbial ecosystem in the cecum and colon, where fermentation of complex plant substrates occurs. This environment encourages microbial adaptation to phytic acid, a primary storage form of phosphorus in feed ingredients like wheat bran, creating a natural selective pressure for the development of specialized phytase-producing bacteria.11

The equine hindgut is a highly efficient bioreactor where natural microbiota degrade over 93% of dietary phytate.11, 12 This rate is significant when compared to modern commercial phytases used in other monogastric animals; even with high-dose “superdosing” (>1500 FTU/kg), maximum phytate degradation reaches only 87.6% in broilers and 81.9–85% in pigs.13, 14 Such high natural degradation suggests that the equine gut harbors efficient enzymes that are functionally adapted to the complex, physicochemical environment of the monogastric animal gut. This specialized niche is further intensified by traditional feeding practices in developing tropical and sub-tropical areas of Asia and Africa, where working horses are frequently maintained on a basal diet of cereal crop residues supplemented with wheat bran which contains high concentrations of phytate (3.1–5.8%).15 We hypothesize that this environment selects for microbial communities with phytase activities that may be superior in stability or substrate affinity to existing commercial enzymes. Consequently, exploring these microorganisms holds the potential to uncover novel phytase genes and enzymes with properties uniquely suited for broader monogastric animal feed applications.

The search for efficient phytases increasingly relies on computational approaches that enable rapid discovery and characterization of promising enzymes.16 Genome mining has emerged as a powerful in silico strategy for identifying phytase-encoding genes within microbial genomes using sequence-based algorithms and curated protein databases.17 This approach allows rapid detection of both known and novel phytase genes, prediction of their potential functions, and assessment of their evolutionary relationships without the need for initial laboratory experimentation. Genome mining is particularly valuable for screening diverse microbial sources, including environmental isolates, to identify candidates with desirable properties for animal feed applications.18

Once candidate genes are identified, physicochemical analysis of the encoded proteins provides critical insights into their suitability for industrial use. Parameters such as molecular weight, theoretical isoelectric point (pI), instability index (II), aliphatic index (AI), extinction coefficient (EC), and grand average of hydropathicity (GRAVY) help predict enzyme behavior under processing and gastrointestinal conditions. For instance, higher aliphatic index and lower instability index values indicate better thermostability, an essential requirement for surviving feed pelleting processes.9, 15 Beyond physicochemical properties, structural features analysis such as modeling of secondary and tertiary structures, identification of active-site residues, and evaluation of substrate-binding features strengthens functional predictions and supports enzyme selection.20 Molecular docking further reveals the enzyme's binding affinity and selectivity toward phytate, offering a theoretical assessment of catalytic efficiency. These structural insights complement experimental studies and help prioritize phytases with enhanced performance characteristics.9

To date, no studies, either experimental or computational, have investigated phytases from horse-feces associated bacteria. In the absence of wet-lab data, in silico analysis offers a practical and efficient first step, minimizing the time and cost of experimental workflows while guiding more targeted laboratory studies.16 Thus, the aim of this study was to identify phytase genes and their translated enzymes from genomes of horse feces-associated bacteria and to evaluate their potential for use in poultry, pig, and fish feed. This was achieved by analyzing their biochemical properties, modeling their three-dimensional structures, and assessing substrate-binding interactions through molecular docking.

2. Methodology

2.1. Retrieving horse feces-associated bacterial genomes

Horse feces-sourced BioSamples were retrieved from the NCBI BioSample database using the query (horse OR equine OR Equus caballus) AND (feces OR feces). On the results page, the organism filter was restricted to bacteria. BioSample categories such as Microbe, Pathogen: clinical (Pathogen.cl), Pathogen: environmental (Pathogen.env), Minimal Information about a Genome Sequence: cultured bacteria/archaea (MIGS.ba), and environmental packages were filtered sequentially. BioSamples annotated as Pathogen.env were excluded, as they represent environmental sources not directly associated with the equine gastrointestinal tract. Among BioSamples assigned to the host-associated environmental package (n = 115), 110 records corresponded to metagenome-assembled genomes (MIMAGs) and were therefore excluded because they do not represent cultured isolates. In total, 148 BioSamples comprising 105 Microbe, 37 Pathogen.cl, and 6 MIGS.ba records with explicit links to the Nucleotide database were retained for downstream analysis.

Additional BioSamples were retrieved from the NCBI BioProject database using above search term to ensure comprehensive coverage of horse feces-associated bacterial genomes. During this search, the organism group filter was restricted to bacteria, and project data filters for Nucleotide or Assembly were applied. Further BioProject search was conducted using the query (Equus caballus AND Bacteria), which returned two BioProject records. One of these BioProjects corresponded to bacterial isolates derived from equine fecal samples.

Taxonomic names or BioSample accession numbers obtained from BioSample metadata or BioProject assembly records were used to retrieve the corresponding genome metadata from NCBI datasets. Relevant information, including accession identifiers, organism names, and assembly details, was compiled in a Microsoft Excel spreadsheet, and duplicate entries were manually removed providing 162 genome metadata (Supplementary Table 1).

2.2. The genome and proteome sequence collection

Using genome GenBank accessions (GCA), a total of 162 genome assemblies which include genomic (genomic.gff) and proteomic (protein.faa) files were retrieved using the NCBI Datasets command-line tool.21 FASTA headers of protein sequences retrieved from the EMBL-EBI database (PRJEB65070) were manually reformatted to the standard >gnl|EMBL|accession format. All proteomes were then merged into a single FASTA file, and a local BLAST protein database was constructed from this file using the standalone BLAST+ suite (v2.17.0) via command-line execution. The commands were generated using ChatGPT and executed in the Windows Anaconda PowerShell environment (Supplementary File 1). Datasets of proteomes from the genomes of horse feces-associated bacteria were available at https://doi.org/10.5281/zenodo.18100064.22

2.3. Query sequence identification

Query phytases were identified through a systematic search of published literature.23, 24, 25, 26, 27 and the NCBI databases. Searches were conducted using genus-level taxonomic names corresponding to horse feces–associated bacterial isolates identified in this study. To ensure representative coverage, query sequences were selected from the four major phytase classes: histidine acid phytases (HAPhys), β-propeller phytases (BPPhys), purple acid phytases (PAPhys), and protein tyrosine phosphatase–like phytases (PTPLPhys). The selected reference sequences included MBI1447871.1, WP_080550623.1, ACJ51391.1, NP_415500.1, and AAM23271.1 (HAPhys); AKS25181.2 and AAQ13669.1 (PTPLPhys); CAB91845.1 (BPPhy); and WDQ15211.1 (PAPhy).

2.4. Identification of phytase enzymes and their coding DNA sequences (CDS)

To identify phytase enzymes from the 162 proteomes, BLASTp searches were performed against the locally built protein database via command-line execution (Supplementary File 1). An E-value threshold of ≤1 × 10−6 and a minimum query coverage of 70% were applied. BLASTp hits with a percent identity of ≥30% were retained.9 All homologous phytases were retrieved in FASTA format from the local BLAST protein database. To avoid redundancy, sequences were clustered at 100% identity using PowerShell scripts generated by ChatGPT (Supplementary File 1). The representative sequence for each cluster was defined as the lexicographically first accession identifier within that cluster, and the corresponding phytase sequence was then extracted from the local BLAST protein database. The methods for retrieving phytases from the local BLAST protein database was published on Zenodo22 and available at https://doi.org/10.5281/zenodo.18100064.

To incorporate the identified phytases with their corresponding CDS entries, CDS features, including gene names or locus tags, genomic coordinates, and protein accession numbers (protein_id), were first extracted from genomic.gff file using custom PowerShell scripts generated by ChatGPT. Next, subject sequence identifiers in the BLASTp output were normalized by removing database-specific prefixes (gnl|EMBL|, gb|, tpg|), retaining only the protein accession numbers. These accession numbers were then used to link BLASTp results with CDS metadata via command-line execution generated by ChatGPT (Supplementary File 1).

2.5. Confirmation of phytases using domain, family and motif predictions

Representative phytases were further validated by conserved domain analysis to confirm the presence of phytase-specific catalytic motifs and to exclude non-specific phosphatases. The domain and family analysis was conducted using NCBI Batch CD-search (https://www.ncbi.nlm.nih.gov/Structure/bwrpsb/bwrpsb.cgi).28 Multiple sequence alignment (MSA) was conducted using (https://www.ebi.ac.uk/jdispatcher/msa/muscle) MUSCLE29 and viewed using JALVIEW to show conserved motif signature.

2.6. Prediction of biochemical properties of phytase enzymes

Prediction of signal peptides and their corresponding cleavage sites was performed using SignalP version 6.0 (https://services.healthtech.dtu.dk/service.php?SignalP-6.0) with default parameter.30 Signal peptides were removed in batch from phytase sequences using a custom PowerShell script that generated using ChatGPT (Supplementary File 1). The physicochemical properties of each mature phytase were determined using the Multiple Protein Profiler (MPP) version 1.0, which is freely available online at https://mproteinprofiler.microbiologyandimmunology.dal.ca/. Protein sequences in FASTA format were uploaded to the server, and the resulting data were generated in tabular format. Output files were downloaded in CSV format for downstream analysis.

2.7. Biosafety and regulatory analysis of sources

The regulatory standing of the source bacteria of candidate phytases was verified through international safety databases. The FDA Search Portal was utilized, employing the “Advanced Search” function. Scientific names of the source bacteria (e.g., Escherichia coli) were entered to identify relevant safety dossiers and “No Questions” letters issued by the FDA. In addition, the taxonomic status of each species was cross-referenced with the European Food Safety Authority (EFSA) Qualified Presumption of Safety (QPS) list to confirm their suitability for use in food and feed applications based on established safety tracks. Furthermore, the safety profiles of the bacterial sources were evaluated by screening for potential pathogenicity and resistance factors using a comparative bioinformatic approach. Initially, the PC1 of 69 phytases were ranked in descending order along with their sources. To ensure high performance while maintaining species diversity, the strain with the highest PC1 value from each unique bacterial source species was selected as a representative query proteome. For the safety assessment, subject datasets were retrieved from the Virulence Factor Database (VFDB) core dataset to identify potential toxins, while antibiotic resistance genes (ARGs) datasets were retrieved from the Comprehensive Antibiotic Resistance Database (CARD) version 4.0.1. Homology searches were conducted using NCBI BLAST+ blastp (Galaxy Version 2.16.0 + galaxy0) at default parameters.31 To ensure scientific rigor and minimize false positives, the outputs were manually filtered with a stringent E-value threshold of 1 × 10−5 and a sequence identity cutoff of ≥35% over a 70% query coverage, adhering to standard regulatory guidelines for the identification of homologous functional proteins.32, 33

2.8. Structural modeling and molecular docking

Prior to structural modeling, the mature protein tyrosine phosphatase (UNQ40438.1) from Prescottella equi strain P2117036, a hypothetical protein (CAJ1246072.1) and glucose-1-phosphatase (CAJ1242485.1) from Klebsiella pneumoniae JRT4AKP isolate, and a bifunctional acid phosphatase/4-phytase (NHR17779.1) from Escherichia coli strain W65.1 were selected using principal component analysis (PCA). Higher values of the first principal component (PC1) indicated greater suitability of phytases for monogastric animal feed applications. PCA was performed in the Windows Anaconda Python environment using scripts generated by ChatGPT (Supplementary File 1). Prior to PCA, biochemical properties reflecting the general gut physiological conditions of poultry, pigs, and fish, including pI >6, lower molecular weight, instability index <40, higher aliphatic index, and the presence of a signal peptide, were defined as selection criteria. These properties were directionally normalized to reflect enzyme desirability and standardized using Z-score normalization.

Structural models of the UNQ40438.1, CAJ1246072.1, CAJ1242485.1 and NHR17779.1 were predicted using AlphaFold2 implemented in the ColabFold environment on Google Colaboratory (https://colab.research.google.com/github/sokrypton/ColabFold/blob/main/AlphaFold2.ipynb).34The top-ranked models, based on Predicted Local Distance Difference Test (pLDDT) scores, are available in ModelArchive (modelarchive.org) with the accession codes ma-va5fj, ma-3zzby, ma-w0001 and ma-8mcqy. The models were visually inspected using UCSF Chimera to analyze secondary structure organization and characteristic catalytic motifs, and were subsequently subjected to molecular docking.

Molecular docking of the phytic acid (ligand) to the phytase models (receptors) was conducted using the AutoDock Vina algorithm in SwissDock web server (https://www.swissdock.ch/).35 The receptor structures were prepared in UCSF Chimera and AutoDockTools version 1.5.7 and exported in PDBQT format. For ligand preparation, the molecular structure of phytic acid was obtained from a published source in image format.4 The image was converted into a machine-readable structure using OSRA (Optical Structure Recognition Application) (https://cactus.nci.nih.gov/cgi-bin/osra/index.cgi). Any structural distortions were corrected, and the molecule was saved in SDF format. The 3D coordinates were then generated, and the geometry was optimized using Avogadro. The optimized 3D conformation was saved as a MOL2 file, converted to PDBQT format using Open Babel GUI.36 and visualized in Jmol. The ligand PDBQT file and receptor PDBQT files were uploaded to SwissDock web server for docking. The docking grid was calculated from coordinates of the last atoms of active site residues (Supplementary File 2).

Docking results were downloaded and visualized using UCSF Chimera. For each target phytase-phytic acid pair, AutoDock Vina generated many docking conformations ranked by binding affinity. Only the top-ranked conformation (Model 1), which exhibited the lowest binding energy (kcal/mol), was selected for downstream analysis. Consequently, the lower-bound and upper-bound Root-Mean-Square Deviation (RMSD) values for these primary baseline poses are reported as 0.000 Å in the results, as they serve as the reference coordinates for the remaining generated models.

3. Results

3.1. Genome-wide identification of phytases in horse feces–associated bacteria

The mined horse feces–associated bacterial genomes for phytase enzymes are provided in Supplementary Table 1. A total of 273 phytase enzymes were identified (Supplementary Table 2) and clustered into 69 groups of 100% identity (Supplementary Table 3). The non-redundant, lexicographically selected phytases are distributed across nine9 bacterial species. E. coli was the most frequent source, accounting for approximately 62% (43 out of 69) of the identified sequences. Several strains were found to harbor multiple phytase-encoding genes and the phytases are located on diverse genomic scaffolds (Table 1).

Table 1.

Non-redundant phytases identified from genomes of horse feces–associated bacteria.

No Microbial source Scaffold Start End Phytase Id
1 Acinetobacter baumannii IHIT26023 DAKHOD010000037.1 20,833 22,401 HCV3120289.1
2 A.baumannii IHIT26669 DAKHOC010000020.1 20,836 22,404 HCV3111965.1
3 Bifidobacterium pseudolongum subsp. globosum 1577B SBKZ01000007.1 19,046 19,816 RYQ77489.1
4 Cutibacterium equinum CBA3108 CP115668.1 1,602,308 1,603,870 WCC79340.1
5 Escherichia coli 109.1 JAQRGT010000002.1 1,592,603 1,593,844 MDC9085637.1
6 E. coli B18_0220 DANILX010000006.1 53,173 54,483 HDK8886165.1
7 E. coli B18_0220 DANILX010000006.1 69,091 70,332 HDK8886181.1
8 E. coli Culture DAFKQR010000016.1 46,350 47,660 HBM7503837.1
9 E. coli Culture DAFKQR010000016.1 61,198 62,439 HBM7503852.1
10 E. coli Culture DAFNSO010000043.1 27,243 28,484 HBN7144586.1
11 E. coli Culture DAFNVT010000017.1 42,763 44,067 HBN7539472.1
12 E. coli Culture DAFUPI010000019.1 42,155 43,453 HBQ4166663.1
13 E. coli JMT78AEC CAUJMT010000008.1 259,480 260,721 CAJ1310421.1
14 E.coli JMT78AEC CAUJMT010000008.1 274,149 275,459 CAJ1310448.1
15 E. coli JRT13AEC CAUJMS010000001.1 1,954,381 1,955,622 CAJ1264477.1
16 E.coli JRT13AEC CAUJMS010000001.1 1,968,957 1,970,267 CAJ1264671.1
17 E.coli JRT27AEC CAUJMK010000001.1 2,851,729 2,852,970 CAJ1276500.1
18 E. coli JRT28BEC CAUJMO010000001.1 1,190,055 1,191,353 CAJ1251675.1
19 E. coli JRT48AEC CAUJMG010000001.1 1,688,042 1,689,340 CAJ1258273.1
20 E. coli JRT48AEC CAUJMG010000001.1 1,712,585 1,713,826 CAJ1258624.1
21 E. coli JRT73AEC CAUJMJ010000001.1 2,748,820 2,750,061 CAJ1275302.1
22 E.coli JRT77AEC CAUJMH010000001.1 1,095,775 1,097,016 CAJ1250103.1
23 E. coli JRT77AEC CAUJMH010000001.1 1,119,849 1,121,147 CAJ1250430.1
24 E. coli JRT83AEC CAUJML010000001.1 851,089 852,330 CAJ1247167.1
25 E. coli JRT83AEC CAUJML010000001.1 875,532 876,830 CAJ1247589.1
26 E. coli RKI6020 DADRTA010000021.1 51,294 52,592 HBA3798890.1
27 E. coli RKI6037 DADRSP010000041.1 27,367 28,608 HBA3746836.1
28 E. coli RKI6097 DADWKQ010000005.1 171,217 172,458 HBB8409509.1
29 E. coli RKI6097 DADWKQ010000005.1 187,012 188,322 HBB8409525.1
30 E. coli RKI6154 DADRSQ010000001.1 322,287 323,585 HBA3748170.1
31 E. coli RKI6154 DADRSQ010000001.1 346,474 347,715 HBA3748192.1
32 E. coli RKI6236 DADRSX010000008.1 135,643 136,953 HBA3783171.1
33 E. coli RKI6236 DADRSX010000035.1 29,777 31,018 HBA3785253.1
34 E. coli RKI6237 DADRSV010000001.1 27,318 28,559 HBA3772245.1
35 E. coli RKI6237 DADRSV010000001.1 51,399 52,697 HBA3772268.1
36 E. coli RKI6248 DADRST010000002.1 394,378 395,619 HBA3763449.1
37 E. coli RKI6248 DADRST010000002.1 417,559 418,863 HBA3763471.1
38 E. coli RKI6252 DADRTB010000012.1 33,333 34,574 HBA3804600.1
39 E. coli RKI6252 DADRTB010000012.1 57,402 58,700 HBA3804623.1
40 E. coli RKI6283 DADVZF010000005.1 125,209 126,519 HBB7008401.1
41 E. coli RKI6283 DADVZF010000005.1 141,127 142,368 HBB7008417.1
42 E. coli RKI6296 DADRVL010000104.1 2152 3450 HBA4096616.1
43 E. coli RKI6323 DADRUG010000012.1 115,436 116,734 HBA3944573.1
44 E. coli RKI6323 DADRUG010000012.1 139,568 140,809 HBA3944596.1
45 E. coli RKI6498 DADRTT010000030.1 27,817 29,058 HBA3884431.1
46 E. coli RKI6624 DADRTG010000023.1 41,100 42,398 HBA3816383.1
47 E. coli RKI6632 DADRUM010000062.1 25,533 26,774 HBA3974158.1
48 E. coli W65.1 JAAOWO010000019.1 51,324 52,622 NHR17779.1
49 Klebsiella pneumonia JMT68BKP CAUJMP010000001.1 3,666,013 3,666,798 CAJ1289980.1
50 K. pneumoniae JMT68BKP CAUJMP010000001.1 3,890,993 3,892,261 CAJ1293159.1
51 K. pneumoniae JRT4AKP CAUJMR010000001.1 805,091 806,359 CAJ1242485.1
52 K. pneumoniae JRT4AKP CAUJMR010000001.1 1,031,164 1,031,949 CAJ1246072.1
53 K. pneumoniae JRT4AKP CAUJMR010000001.1 1,248,923 1,250,170 CAJ1249150.1
54 Limosilactobacillus reuteri LR17 QGHQ01000008.1 51,074 51,868 PWT44668.1
55 L. reuteri LR17 QGHQ01000004.1 68,638 69,420 PWT45088.1
56 L. reuteri LR18 QGHP01000006.1 15,360 16,154 PWT69600.1
57 L. reuteri LR19 QGHO01000006.1 30,873 31,667 PWT31569.1
58 Pediococcus acidilactici pll CP067392.1 272,132 272,920 QQP83584.1
59 P. acidilactici pll CP067392.1 329,061 329,840 QQP83630.1
60 P. acidilactici pll CP067392.1 330,123 330,914 QQP83631.1
61 Prescottella equi P2117036 CP093553.1 3,719,009 3,719,824 UNQ38504.1
62 P. equi P2117036 CP093553.1 847,061 847,888 UNQ40438.1
63 P. equi P2117036 CP093553.1 2,441,799 2,442,713 UNQ41831.1
64 P. equi P2117036 CP093553.1 2,442,857 2,443,771 UNQ41832.1
65 P. equi P2120831 CP093554.1 837,841 838,668 UNQ35784.1
66 P. equi P2120831 CP093554.1 2,423,196 2,424,110 UNQ37129.1
67 P. equi P2120831 CP093554.1 2,424,254 2,425,168 UNQ37130.1
68 Salmonella enterica subsp. enterica serovar 4, PNCS015432 AAPSYM010000020.1 66,023 67,264 EEH4304300.1
69 S. enterica subsp. enterica serovar Anatum 15,461 JAKEYR010000003.1 124,599 125,840 MCF1779828.1

Homology search complemented with conserved domain analysis revealed two mechanistically distinct protein families (Supplementary Table 2 and Supplementary Table 4). The histidine acid phosphatase (HAP) superfamily (cl11399) dominated the phytase landscape, with a smaller contribution from protein tyrosine phosphatase–like phytases. Histidine acid phytases (HAPhys) (∼410–436 amino acids) included AppA family phytases, acid phosphatases/4-phytases, and bifunctional glucose-1-phosphatase/inositol phosphatases (Agp). These proteins exhibited nearly complete domain coverage with highly significant matches (E-values often 0, bit scores >800), indicating strong functional conservation. HAPhys were abundant in E. coli and S. enterica. In E. coli, AppA and Agp often co-occurred at distinct loci, reflecting dual 3- and 4-phytase activity. Beyond Enterobacteriaceae, HAPhys were detected in A. baumannii and C. equinum, showing lower sequence identity but conserved catalytic domains. A few sequences were assigned only at the superfamily level, suggesting potential divergence.

PTP-like phytases (PTPLPhys) (∼250–310 amino acids) were characterized by Oca4/COG2365 or Y_phosphatase3 (pfam13350) domains within the cl25999 and cl48179 superfamilies. These enzymes were present in L. reuteri, P.acidilactici, B. pseudolongum, and P. equi, and displayed moderate sequence identity (∼30–47%) but strong statistical support (E-values 10−21–10−81), consistent with a cysteine-based catalytic mechanism.

Detailed inspection of the MSA revealed both strong conservation and lineage-specific variation within canonical HAPhy motifs (Fig. 1). All sequences contained the characteristic RHGXRXP-type catalytic motif near the N-terminus; however, distinct motif variants were observed among different phytase groups. Bifunctional glucose-1-phosphatases predominantly exhibited a RHNXRXP motif, whereas phytases from C. equinum and Acinetobacter spp. displayed a RHGXRXL motif. Despite these substitutions, the catalytic histidine residue was invariant across all sequences, indicating preservation of the nucleophilic center essential for phytate hydrolysis.

Fig. 1.

Fig. 1

Multiple sequence alignment of histidine acid phytases identified from genomes of horse feces–associated bacteria. Conserved catalytic motifs characteristic of HAPhys are highlighted. The N-terminal catalytic signature occurs as RHGXRXP, RHNXRXP in bifunctional glucose-1-phosphatases, and RHGXRXL in Cutibacterium and Acinetobacter phytases. The second conserved motif shows lineage-specific variants, including HDT in AppA-type phytases, HDS in bifunctional glucose-1-phosphatases, and HAE in Cutibacterium and Acinetobacter phytases. Invariant catalytic histidine residues are conserved across all sequences, whereas sequence divergence is largely confined to non-catalytic regions.

Variation was also observed in the second conserved HAP motif located in near the C-terminus. Phytases from Cutibacterium and Acinetobacter consistently harbored an HAE motif, while bifunctional glucose-1-phosphatases contained an HDS motif. In contrast, the majority of AppA-type phytases and related acid phosphatases retained the canonical HDT motif. These motif substitutions occurred at positions known to contribute to proton transfer and substrate coordination,37 suggesting subtle lineage-specific adaptations in catalytic chemistry rather than loss of function. Outside these conserved catalytic motifs, sequence variability was largely restricted to insertions, deletions, and low-complexity regions, particularly in the N-terminal signal peptides.

Multiple sequence alignment of PTPLPhys from Bifidobacterium, Klebsiella, Pediococcus, Limosilactobacillus, and Prescottella species revealed a highly conserved catalytic architecture characteristic of protein PTPLPhys (Fig. 2). All sequences harbor the hallmark PTP-loop motif, HC(X)₅R, unambiguously conserved across the genome (e.g., HCTAGKDRT, HCSAGKDRT, HCAVGKDRT). The catalytic cysteine (Cys) and invariant arginine (Arg) residues are strictly conserved, supporting a shared cysteine-dependent phospho-enzyme intermediate mechanism.30, 31 Closely related strains (e.g., L. reuteri LR17–LR19 and P. acidilactici paralogs) exhibit near-identical sequences, consistent with recent divergence or gene duplication, whereas inter-genus comparisons reveal greater divergence outside conserved motifs. Nonetheless, the preservation of catalytic residues across all taxa indicates strong purifying selection acting on enzymatic function.

Fig. 2.

Fig. 2

Multiple sequence alignment of protein tyrosine phosphatase–like phytases identified from genomes of horse feces–associated bacteria. Alignment of PTPLPhys sequences highlights strong conservation of the catalytic core despite taxonomic diversity. The hallmark PTP-loop motif HC(X)₅R is strictly conserved across all sequences, indicating a shared cysteine-dependent catalytic mechanism. Sequence variability is largely restricted to the N-terminal regions, whereas residues flanking the active site remain highly conserved.

3.2. Signal peptide of horse feces-associated bacterial phytases

Comparative signal peptide prediction revealed a fundamental divergence in secretion strategies between HAPhys and PTPLPhys associated with horse feces (Supplementary Table 5). Among HAPhys, signal peptides were detected in nearly all sequences, with overwhelming support for Sec/SPI-type and consistently high posterior probabilities (typically ≥0.97). Predicted cleavage sites were highly conserved, occurring predominantly at positions 22–23, with minor variation27, 28, 29 in a subset of K. pneumoniae sequences. Acinetobacter HAPhys carried lipoprotein (Sec/SPII) signal peptides with cleavage sites at positions 18–19, indicating membrane anchoring. Only a single HAPhy of C. equinum (WCC79340.1) exhibited ambiguous classification with reduced confidence, suggesting a non-canonical signal peptide. Collectively, these results demonstrate that HAPhys are predominantly secretory enzymes, consistent with extracellular or periplasmic phytate hydrolysis.25, 39

In contrast, signal peptide prediction for PTPLPhys (n = 18) showed that 12 sequences were confidently classified as OTHER with posterior probabilities of 1.0, indicating a complete lack of signal peptides and exclusion from known secretion pathways. These non-secretory PTPLPhys were distributed across Limosilactobacillus, Pediococcus, Bifidobacterium, and Klebsiella species, supporting a conserved intracellular localization for this phytase family. Only 6 PTPLPhys, all derived from P. equi, were predicted to contain Sec/SPI signal peptides, with cleavage sites between positions 33–36 and confidence (Pr = 0.61–0.97).

From a functional and applied perspective, this sharp contrast underscores the predicted suitability of HAPhys for monogastric animal feed applications. Extracellular phytases are easier to produce, recover, and formulate as commercial feed additives, thereby lowering production costs and facilitating large-scale industrial application.40 Conversely, the predominantly intracellular nature of PTPLPhys suggests a primary role in microbial phosphorus metabolism rather than direct interaction with luminal phytate.

3.3. Physicochemical properties of mature phytases from horse feces–associated bacteria

Comprehensive physicochemical characterization of 73 bacterial phytases, comprising 69 horse feces–associated isolates and four reference strains, revealed that they segregate into two distinct molecular classes: the larger HAPhys and the smaller PTPLPhys. Due to the high degree of similarity among certain isolates, these 73 enzymes are summarized into 20 representative groups in Table 2 based on their type and source organism. Detailed physicochemical profiles for each of the 73 individual phytases are provided in Supplementary Table 6.

Table 2.

Condensed physicochemical characterization of horse feces–associated bacterial phytases.

No Type Name Source organisms MW (kDa) pI II AI GRAVY
1 HAPhy AppA family phytase/histidine-type acid phosphatase E.coli 44.62–45.12 5.45–6.11 34.7–37.89 88.62–90.68 −0.21 to −0.25
2 HAPhy Bifunctional glucose-1-phosphatase/inositol phosphatase E.coli 43.51–43.63 5.29–5.6 42.6–44.83 76.06–78.06 −0.40 to −0.44
3 HAPhy Bifunctional glucose-1-phosphatase/inositol phosphatase S. enterica subsp. enterica serovar Anatum15461 43.5 6.22 41.73 78.34 −0.39
4 HAPhy 3-phytase (AAM23271.1) K. pneumoniae ARS1 43.4 8.7 38.96 85.51 −0.34
5 HAPhy 3-phytase/Glucose-1-phosphatase E.coli 43.51–43.63 5.38–5.60 44.27–44.99 77.06 −0.41 to −0.43
6 HAPhy Acid phosphatase E.coli 44.58–45.19 5.95–6.34 36.24–37.75 88.04–89.98 0.21–0.26
7 HAPhy bifunctional acid phosphatase/4-phytase E.coli 45.15–44.90 5.68–6.17 35.78–37.75 89.02–90.93 −0.22 to −0.25
8 HAPhy bifunctional glucose-1-phosphatase/inositol phosphatase S. enterica subsp. enterica serovar 4,PNCS015432 43.45 6.27 42.8 78.34 −0.38
9 HAPhy bifunctional glucose-1-phosphatase/inositol phosphatase E.coli 109.1 43.46 5.38 45.26 77.06 −0.41
10 HAPhy Glucose-1-phosphatase K.pneumoniae 43.19–43.39 5.95–9.12 39.41–42.01 77.46–86.27 −0.28 to −0.39
11 HAPhy histidine-type phosphatase (MBI1447871.1) Acinetobacter sp AC12 56.56 5.17 34.15 75.58 −0.53
12 HAPhy histidine-type phosphatase A.baumannii IHIT26669 56.43–56.56 5.32–5.33 32.95–34.95 76.90–76.92 −0.49 to −0.51
13 HAPhy histidine-type phosphatase C.equinum CBA3108 54.06 5.34 26.8 67.99 −0.54
14 HAPhy Multiple myoinositol polyphosphate phosphatase (NP_415500.1) E.coli strain K12 substrain MG1655 44.69 6.11 37.25 89.73 −0.24
15 PTPLPhy Protein tyrosine phosphatase like phytase (AKS25181.2) L.fermentum 29.04 5.28 28.94 89.55 −0.2
16 PTPLPhy hypothetical protein K.pneumoniae 28.51 5.71–5.89 39.05 99.92 −0.09
17 PTPLPhy protein-tyrosine-phosphatase L.reuteri 29.47–30.4 5.86–9.09 21.98–28.44 86.86–89.77 −0.34 to −0.48
18 PTPLPhy protein-tyrosine-phosphatase P.acidilactici pll 29.87–30.38 5.57–9.35 35.51–46.60 84.89–96.76 −0.46 to −0.54
19 PTPLPhy protein-tyrosine-phosphatase B.pseudolongum subsp.globosum 1577B 28.91 5.01 44.09 99.49 −0.16
20 PTPLPhy protein-tyrosine-phosphatase P.equi 26.19–28.79 4.64–5.09 10.34–32.28 86.15–94.02 −0.07 to −0.19

Values represent the ranges (min–max) or specific values for 73 studied enzymes (69 isolates and 4 references) grouped by source organism and enzyme type. Bold texts indicate reference phytases with experimentally determined molecular and biochemical properties. MW: Molecular Weight; pI: Isoelectric Point; II: Instability Index; AI: Aliphatic Index; GRAVY: Grand Average of Hydropathicity; HAPhy: Histidine acid type phytase family; PTPLPhy: protein-tyrosine-phosphatase-like phytase family. Full individual data for all 73 enzymes is available in Supplementary Table 6.

3.3.1. Molecular weight distribution

HAPhys predominantly exhibited molecular weights in the range of 43.19–45.19 kDa (Table 2), falling within the typical range (37–55 kDa) reported for bacterial phytases.41 The HAPhys of Acinetobacter and Cutibacterium species extended to 54.06–56.56 kDa, consistent with additional structural elements and extended loop regions present in HAP family members.24 In contrast, PTPLPhys were markedly smaller, clustering tightly between 26.19 and 30.40 kDa, indicative of compact catalytic cores typical of low–molecular-weight PTP-like enzyme.42 This size reduction suggests reduced domain complexity and a mechanistic strategy distinct from that of HAPhys.

Experimentally determined MWs of 42 kDa39 and 44.71 kDa23 were reported for NP_415500.1, which closely match the in silico determined value, 44.69 kDa. For the reference K. pneumoniae ARS1 HAPhy (AAM23271.1), the enzyme was found to be 43.40 kDa in this study, which is comparable with the experimental value of 42 kDa.25 The slight deviation observed is fall within the SDS-PAGE–based MW estimation error. SDS–PAGE provides only an approximate MW estimate, with reported inaccuracies typically in the range of ±5–10%, due to factors such as gel composition, protein charge, and SDS binding variability.37, 38 Thus, a protein predicted to be ∼42 kDa may migrate on SDS–PAGE anywhere between approximately 38 and 46 kDa.

Knowing the MW of phytase plays a critical role in its application as a feed additive for two main reasons. First, the MW guides enzyme production and purification strategies, as techniques such as ultrafiltration, chromatography, and electrophoresis are often optimized based on the size of the protein.43 The knowledge of MW ensures efficient separation from other proteins and contaminants, improving yield and purity. Second, MW is essential in evaluating enzyme kinetics and commercial cost. By converting mass-based measurements to molar units using the MW, producers can accurately calculate catalytic parameters such as Vmax, kcat and catalytic efficiency (kcat /km),19, 32 which reflect the efficiency of the enzyme in hydrolyzing phytate. Higher catalytic efficiency can reduce the amount of enzyme required in feed formulations, thereby lowering production costs. In this way, understanding the MW not only informs biochemical characterization but also has direct implications for the economic feasibility of phytase as a monogastric animal feed additive.

3.3.2. Isoelectric point (pI)

The predicted isoelectric points spanned a broad range (pI 4.64–9.35). HAPhys exhibited slightly acidic to alkaline pI values (5.19–9.12) (Table 2). Notably, a subset of Klebsiella-derived HAPhy showed alkaline pI values (8.70–9.12). The E. coli HAPhy (NP_415500.1) exhibited a predicted pI of 6.11, closely matching the experimentally determined pIs of its two isoforms (6.50 and 6.30).23 The experimentally determined pI of 8.7 for K. pneumoniae subsp. pneumoniae XY-545 validates the in silico–predicted pI of the K. pneumoniae ASR1 HAPhy. PTPLPhys exhibited substantial pI heterogeneity. Phytases from Limosilactobacillus and Pediococcus predominantly showed alkaline isoelectric points (pI 8.71–9.35), whereas those from Bifidobacterium and Prescottella were distinctly acidic (pI 4.64–5.21). Klebsiella PTPLPhys also clustered within the acidic range (pI 5.71–5.89). This variability likely reflects niche-specific adaptation and differences in substrate accessibility or cellular localization.

The isoelectric point (pI) represents the pH at which a protein carries no net electrical charge.45 The relationship between pI and environmental pH therefore directly determines the net charge of phytase enzymes: at pH values below the pI, phytases carry a net positive charge, whereas at pH values above the pI they become negatively charged. Consequently, phytases with acidic pIs are predicted to be positively charged and function optimally in acidic environments, such as the gastrointestinal tract or acidic fermentation niches, while phytases with alkaline pIs may retain favorable electrostatic properties under neutral to alkaline conditions. Accordingly, the observed pI range suggests that horse feces–associated bacterial phytases are likely to acquire a positive net charge under the strongly acidic conditions of the monogastric stomach, which is recognized as the primary site of phytase activity.4, 20, 46 In contrast, phytases with pI ≥8.7 are expected to remain positively charged throughout much of the gastrointestinal tract. Consistent with this interpretation, stomach pH in poultry ranges from 5.2 to 5.8 in the crop4 and decreases to 2.5–3.5 in the proventriculus and gizzard.20 Similarly, the porcine stomach maintains an acidic pH of 2.0–3.5,4 while fish species exhibit stomach pH values ranging from 3 to 6.47

Experimental studies validated that optimum activity of phytases occurs at reaction medium pH less than theoretical or experimental pI of phytases. NP_415500.1 (pI = 6.11 or 6.50 or 6.30) demonstrated optimal pH 4.5.23 Similarly, AAM23271.1 (pI = 8.7) displayed optimum pH 5.0.25 PTPLPhys of L. fermentum NKN51 (AKS25181.2) (pI = 5.28) showed optimum pH 5.0.26 The phytases of P. acidilactici SMVDUDB2 showed optimum activity of pH 5.5 which supports PTPLPhys from the same species (QQP83631.1, pI = 8.73 and QQP83584.1, pI = 9.35).

3.3.3. Phytase stability (instability index)

The instability index (II) values indicated that the majority of phytases are predicted to be stable (II < 40) (Table 2). HAPhys exhibited instability indices (26.8–45.26), where all bifunctional glucose-1-phosphatase variants approached the threshold of instability (II = 41.73–45.26), suggesting increased conformational flexibility possibly linked to multifunctionality. PTPLPhys demonstrated greater variability, ranging from stable enzymes (II = 10.34–39.05) to unstable P. acidilactici pII phytases (II = 41.64 QQP83584.1 and II = 46.6 QQP83630.1). The exceptionally low instability indices observed in several PTPLPhys imply enhanced structural rigidity, potentially supporting function under fluctuating environmental conditions.

The II is a predictive measure of protein stability based on amino acid composition, designed to estimate whether a protein is likely to remain stable under experimental conditions (in vitro, in vivo or physiology).9, 25 Proteins with an II value below 40 are considered stable, indicating lower susceptibility to degradation, higher structural integrity, and a greater likelihood of maintaining proper folding. In contrast, proteins with II values above 40 are predicted to be less stable and more prone to unfolding or rapid breakdown. Accordingly, a lower II of the horse feces–associated bacterial phytases suggests that these proteins are more likely to maintain structural integrity and enzymatic activity during storage, feed processing, and exposure to harsh conditions such as heat, moisture, or the acidic environment of the stomach.

3.3.4. Aliphatic index and thermostability

The aliphatic index (AI) was consistently high across all phytases, indicating adaptation to function over broad temperature ranges. Stable HAPhys typically exhibited AI values of 67.69–90.93, with E. coli HAPhys clustering at the upper end of this range (88.04–90.93), consistent with their well-documented thermal robustness. In contrast, stable PTPLPhys frequently displayed very high AI values (84.98–99.92), particularly among P. equi strains and K. pneumoniae isolates, many of which exceeded AI values of 90. These elevated indices reflect enrichment in aliphatic residues and likely confer enhanced thermostability despite the reduced protein size of PTPLPhys.

The AI, which reflects the relative abundance of aliphatic amino acids (alanine, valine, isoleucine, and leucine), is widely used as an indicator of protein thermostability, with higher values generally associated with increased resistance to heat-induced denaturation.48 The interpretive value of AI is supported by experimental thermostability data for reference phytases. The E.coli HAPhy (NP_415500.1) exhibited a high AI (89.73) and displays optimal activity at 60 °C, although it loses approximately 80% of its activity after 30 min at 85 °C.23, 39, 49 In comparison, the K. pneumoniae HAPhy (AAM23271.1), with a lower AI (85.51), shows optimal activity at 50 °C but undergoes a rapid loss of activity at 55–60 °C and becomes undetectable after 15 min at 65 °C.25 These differences in thermal tolerance are consistent with the higher AI of the E. coli phytase. Similarly, the L. fermentum PTPLPhy exhibits a high AI (89.55) and an optimal temperature of 60 °C.26 In P. acidilactici, phytase activity peaks at 37 °C and maintains 64.7% after 3 h of incubation at 60 °C.50 These observations support the moderate thermostability inferred from the AI value (84.98) of the corresponding P. acidilactici PTPLPhy (QQP83631.1).

Animal feed pelleting involves severe thermal conditions, with conditioning temperatures of 65–90 °C and die friction further increasing temperatures to around 90 °C or higher,51 which most native phytases are unlikely to withstand without protein engineering49, 52 and the enhanced coating technology.53 In contrast, the gastrointestinal tract of monogastric animals operates at much lower physiological temperatures. Pigs typically maintain body temperatures of 37.8–40 °C,54 while poultry exhibit higher physiological temperatures of 41–42 °C,55 In ectothermic fish, GIT temperature generally reflects that of the surrounding aquatic environment.56 These temperature ranges fall within the functional window deduced from the AI and supported by the experimental value of the reference phytases. Consequently, horse feces-associated bacterial phytases are more suitable for practical feed applications when applied post-pelleting (e.g., via liquid coating or top-dressing), where they can retain catalytic activity and function effectively under normal gastrointestinal conditions.7

3.3.5. Hydropathicity (GRAVY)

All phytases exhibited negative GRAVY (grand average of hydropathicity) values, indicating an overall hydrophilic character and compatibility with aqueous cellular or extracellular environments. Stable HAPhys showed GRAVY scores ranging from −0.21 to −0.54, while stable PTPLPhys displayed slightly broader but similarly negative values (−0.07 to −0.48) (Table 2). The GRAVY index reflects the average hydropathy of all amino acids in a protein, with negative values denoting hydrophilicity and positive values indicating hydrophobicity.57 Accordingly, horse feces–associated bacterial phytases, together with their reference counterparts, are predominantly hydrophilic, a property that favors solubility in aqueous environments. This characteristic is particularly advantageous for feed applications, as enhanced solubility in gastrointestinal fluids of monogastric animals is expected to improve substrate accessibility and enzymatic efficiency.

3.4. Biosafety and regulatory analysis of sources

Analysis of bacterial sources of the 69 phytase candidates via the FDA Search Portal and EFSA QPS list categorized the sources by taxonomic safety and functional relevance to animal nutrition. Results identified L. reuteri and P. acidilactici strains as GRAS, while B. pseudolongum remains unclassified. Conversely, Salmonella enterica subsp. enterica, K. pneumoniae, and A. baumannii were designated as UNSAFE, with E. coli strains split between clinical (unsafe) and engineered (GRAS) isolates. Data for C. equinum and P. equi were unavailable. While VFDB and CARD homology searches confirmed predicted virulence factors across all selected strains, no antimicrobial resistance genes were detected. The virulence factors with the highest percent identity for each selected bacterial strain are provided in Table 3, and comprehensive details are published on Zenodo.org (https://doi.org/10.5281/zenodo.19731840).

Table 3.

Identified virulence factors from the selected horse feces–associated and phytases source bacteria.

Bacterial strains Query/Subject Accessions Virulence factor Identity E -value Query coverage
K. pneumoniae JRT4AKP gnl|EMBL|CAJ1239145.1/VFG049144 Multidrug efflux pump subunit AcrB/acriflavine resistance protein B 100% 0 100%
E. coli W65.1 NHR17722.1/VFG045791 transcriptional regulator CsgD 100% 0 100%
S. enterica subsp. enterica serovar Anatum15461 MCF1778825.1/VFG000515 SctN family type III secretion system ATPase SsaN 100% 0 100%
A. baumannii IHIT26023 HCV3118926.1/VFG050428 Hypothetical protein 100% 0 100%
C. equinum CBA3108 WCC79552.1/VFG046465 Elongation factor Tu 71% 0 100%
P. equi P2120831 UNQ36973.1/VFG001405 RNA polymerase sigma factor/ sigA 93.0% 0 75.2%
B. pseudolongum subsp. globosum 1577B RYQ77937.1/VFG046465 Elongation factor Tu 67.4% 0 99.7%
L. reuteri LR17 PWT44508.1/ VFG000964 UTP–glucose-1-phosphate uridylyltransferase hasC 70.4% 1.8910−158 99.7%
P. acidilactici pll QQP83744.1/VFG000077 ATP-dependent Clp endopeptidase proteolytic subunit ClpP 71.6% 6.5510−105 100%

Of the many virulence factors, the one with the highest percent identity for each bacterial source is shown. Virulence hits selected based on the ≥35% identity, an E-value = 1х10−5 and query coverage ≥70% are published on Zenodo.org (https://doi.org/10.5281/zenodo.19731840).

3.5. Structural modeling and docking of horse feces-associated bacterial phytases

3.5.1. Phytase selection strategy for structural modeling and molecular docking

Principal component analysis (PCA) revealed that most of the physicochemical variation along with signal peptides among horse feces–associated bacterial phytases can be effectively summarized in two principal components. Among HAPhys, CAJ1242485.1 (PC1 = 2.15) and NHR17779.1 (PC1 = 1.97) exhibited the most favorable physicochemical profiles. In the PTPLPhys, UNQ40438.1 (PC1 = 3.38) and CAJ1246072.1 (PC1 = 0.67) were identified as the most suitable candidates (Supplementary Table 7).

3.5.2. Alphafold model of horse feces-associated bacterial phytases

The overall AlphaFold2 model of HAPhy is monomer composed of α-helices, β-sheets, and connecting coil regions, organized into a larger α/β-domain and a smaller α-domain. The α/β-domain of E. coli phytase is composed of 6 β-sheets and 8 α-helices, and the α-domain consists of 8 α-helices and 2 β-sheets (Fig. 3 A). The α/β-domain of K. pneuomiae phytase is composed of 8 β-sheets and 8 α-helices, and the α-domain consists of 8 α-helices and 2 β-sheets (Fig. 3 B). The active site includes RHGxRxP, RHGVRAP for E. coli HAPhy and RHGIRPP for K.pneuomiae HAPhy were exist at similar regions of the 3D models. The overall AlphaFold2 model of PTPLPhys is also monomer composed of a central 4 stranded β-sheets and α-helices. The central sheet of P. equi phytase is flanked by 2 α-helices on one side and 8 α-helix bundle on the opposite side (Fig. 3 C), whereas by 2 α-helices on one side and 11 α -helix bundle on the opposite in K. pneuomiae phytases (Fig. 3 D). The topology of phosphate-binding loop (P-loop), HCSAGKDR for P. equi phytase and HCAVGKDR for K.pneuomiae phytases were found in different.

Fig. 3.

Fig. 3

Alphafold models of horse-feces-derived bacterial phytase. (A) Model of glucose-1-phosphatase of K. pneumoniae JRT4AKP (CAJ1242485.1) colored by secondary structures; (B) Model of bifunctional acid phosphatase/4-phytase of E. coli W651 (NHR17779.1) colored by secondary structures; (C) Model of protein tyrosine phosphate of P. equi P2117036 (UNQ40438.1) colored by secondary structures; (D) Model of hypothetical protein of K. pneumoniae JRT4AKP (CAJ1246072.1) colored by secondary structures. Helices (slate gray), sheet (firebrick) and coil (dark khaki). Active sites are shown by sticks. (For interpretation of the references to colour in this figure legend, the reader is referred to the web version of this article.)

In terms of model quality, all AlphaFold2 predictions showed high confidence and near-experimental accuracy, with pLDDT scores of 97.2, 96, 95.2, and 90.9 for CAJ1242485.1, NHR17779.1, UNQ40438.1, and CAJ1246072.1, respectively. In addition, 91% of residues had pLDDT scores ≥70, indicating that the models were of sufficient quality for subsequent molecular docking analysis (data not shown).

Further model validation was performed by structural superimposition using the UCSF Chimera MatchMaker tool (Supplementary Fig. 1). The superimposition of the AlphaFold2-predicted bifunctional acid phosphatase/4-phytase model from E. coli with the experimentally resolved structure 7Z2S demonstrated an exceptionally low Root Mean Square Deviation (RMSD) of 0.549 Å over 394 pruned atom pairs (1.129 Å across all atoms). Such a sub-angstrom RMSD indicates near-identical spatial positioning of backbone atoms following optimal structural alignment. The minimal deviation observed suggests that the AlphaFold2 model accurately recapitulates the experimentally determined fold, including preservation of secondary structural elements and overall domain architecture. The slightly higher RMSD when considering all atoms likely reflects minor differences in flexible loop regions or side-chain conformations.58 Similarly, alignment of the experimental structure 2WNH with the AlphaFold2-predicted glucose-1-phosphatase model from K. pneumoniae yielded an even lower RMSD of 0.485 Å over 391 pruned atom pairs (0.621 Å across all atoms). The consistency between backbone and all-atom RMSD values indicates not only accurate backbone tracing but also reliable side-chain positioning. Overall, RMSD values below 1 Å strongly demonstrate that AlphaFold2 predictions for these phosphatases achieve near-experimental accuracy.58 This high level of structural agreement supports the use of these predicted models for downstream structural analyses, including substrate docking, and protein engineering applications aimed at predicting catalytic efficiency or stability under gastrointestinal conditions relevant to monogastric nutrition.

In contrast, superimposition of the experimental structure 2B4P (chain A) with the AlphaFold2-predicted protein tyrosine phosphatase from P. equi yielded an RMSD of 1.098 Å for pruned atom pairs. This value indicates strong structural agreement within the aligned core region, suggesting that the conserved catalytic scaffold of the PTPLPhy is reliably predicted. However, when all atom pairs were included in the calculation, the RMSD increased markedly to 17.314 Å. Such a substantial increase indicates major structural divergence outside the conserved core.58 This pattern suggests that while the catalytic core architecture is preserved, peripheral regions may exhibit considerable conformational variability or domain rearrangement. A comparable trend was observed for the hypothetical K. pneumoniae protein, where alignment with 2B4P produced an RMSD of 1.068 Å for pruned atoms but increased dramatically to 17.991 Å when all atoms were considered. Collectively, these results indicate that the AlphaFold2 models accurately capture the conserved catalytic framework characteristic of protein tyrosine phosphatase–like enzymes, but substantial divergence exists in variable surface or accessory regions.

3.5.3. Molecular docking of horse feces-associated bacterial phytase with phytic acid

Docking analysis revealed differential binding affinities and interaction profiles of phytic acid with the evaluated phytases with clear differences between HAPhy and PTPLPhys. The top-ranked docking conformations (Model 1) selected for final evaluation exhibited baseline RMSD values of 0.0 Å (Table 4), as they represent the highest-affinity reference structures used by the software for each respective ligand.

Table 4.

AutoDock Vina docking results of phytic acid with histidine acid and protein tyrosine phosphatase-like phytases.

Phytases Score RMSD l.b.a RMSDa u.b. H-bonds Interacting residues
CAJ1242485.1 -5.42 0.00 0.000 5 R17 (10), T20 (5), G22 (4), N23 (8), R89 (4), K187 (1), N198 (6), N242 (5), H279 (1), D280 (3), T281 (12)
2wnh chain A -6.97 0.000 0.000 17 R24 (6), H25 (9), R28 (10), T31 (12), N34 (12), L98 (4), R100 (16), Q132 (2), Y249 (4), N253 (10), H290 (15), D291 (9), T292 (4), N293 (6)
NHR177791 -6.36 0.000 0.000 8 H17 (8), R20 (13), T23 (6), K24 (9), R92 (12), N126 (1), M216 (1), E219 (2), Q253 (1), F254 (1), H303 (8), D304 (8), T305 (5)
7Z2S chain A -6.85 0.000 0.000 9 R16 (12), H17 (4), R20 (24), T23 (11), K24 (3), R92 (7), S215 (13), M216 (6), H250 (3), N253 (9), F254 (7), N258 (5), R267 (1), D304 (4), T305 (3)
UNQ404381 -5.15 0.000 0.000 4 N65 (8), Q68 (4), V87 (1), V88 (6), G120 (3), N121 (19), G121 (3), M128 (4), C157 (2), S158 (4), A159 (2), D162 (2), R163 (27)
CAJ1246072 -2.89 0.000 0.000 3 G16 (6), I18 (10), A44 (16), N46 (10), R47 (45), T76 (14), T103 (5), A153 (6), V154 (5), Q190 (4), V191 (24), W194 (16)
2B4P -5.77 0.000 0.000 8 R57 (4), K189 (13), N223 (5), H224 (11), C252 (1), V256 (10), G257 (17), F289 (2), K305 (22), Y309 (13), K312 (3)

Abbreviations: 2WNH chain A: reference of CAJ1242485.1; 7Z2S chain A: reference of NHR177791; 2B4P: reference of UNQ404381 and CAJ1246072. Numbers in brackets indicate interaction counts. Bold font shows catalytically important residues. The catalytic cysteine residue (C252) was not part of the interacting residues in CAJ1246072.

2WNH chain A: reference of CAJ1242485.1; 7Z2S chain A: reference of NHR177791; 2B4P: reference of UNQ404381 and CAJ1246072); numbers in brackets indicate interaction counts; bold font shows catalytically important residues; catalytic cysteine residue (C252) was not part of interacting residues in CAJ1246072.

a

RMSD values (RMSD/lb. and RMSD/ub) are reported as 0.000 Å because Table 4 strictly presents the metrics of the single lowest-energy pose (Model 1), which AutoDock Vina uses as the structural reference origin for that specific docking run.

Among HAPhys (CAJ1242485.1, 2WNH chain A, NHR177791, and 7Z2S chain A), binding affinities were consistently strong, ranging from −5.422 to −6.967 kcal/mol, and were supported by extensive hydrogen-bond networks (5–17 hydrogen bonds). CAJ1242485.1 exhibited moderate binding (−5.422 kcal/mol) with five hydrogen bonds involving conserved Arg, Asn, and Thr residues. Its reference structure, 2WNH chain A, showed the strongest interaction overall (−6.967 kcal/mol) and the highest number of hydrogen bonds,17 with dense contacts dominated by Arg and His residues typical of the HAPhy active site. Similarly, NHR177791 demonstrated favorable binding (−6.356 kcal/mol) with eight hydrogen bonds, engaging key catalytic residues, while its reference 7Z2S chain A displayed slightly stronger binding (−6.846 kcal/mol) and a broader interaction network, indicating conserved ligand recognition between predicted and experimental HAPhy structures. The score difference among HAPhys were ≤ 1 kcal/mol. The score differences ≤2.85 kcal/mol is within the typical scoring error of AutoDock Vina59 and community protocol considers differences ≤1 kcal/mol as negligible.

In contrast, PTPLPhys (UNQ404381 and CAJ1246072) showed weaker phytic acid binding, with docking scores of −5.152 and − 2.889 kcal/mol and only three to four hydrogen bonds, suggesting reduced electrostatic complementarity with the highly negatively charged substrate. Although multiple residue contacts were observed, interactions with positively charged residues were less frequent and less organized. The PTPLPhys reference structure 2B4P exhibited improved binding (−5.773 kcal/mol) and eight hydrogen bonds relative to its predicted counterparts; however, its interaction profile remained less extensive than those of horse-feces associated HAPhys.

4. Discussion

While poultry, swine, and fish lack sufficient endogenous phytase to unlock phosphorus from plant-based phytic acid,4 the equine fecal microbiome may present solution. Our in silico analysis reveals a diverse bacterial reservoir of putative phytases that could bridge the gap between poor nutrient utilization and environmental sustainability (Table 1). As hindgut fermenters, horses consume diets rich in plant-based phytic acid, which reaches the cecum and colon largely intact due to the lack of endogenous equine phytases.5, 11This creates a selective niche for microorganisms capable of phytate hydrolysis, explaining the high frequency of these enzymes in our dataset. Consistent with this ecological adaptation, gene redundancy was observed in several taxa; nearly all E. coli isolates encoded both AppA phytases and bifunctional glucose-1-phosphatases, while K. pneumoniae, L. reuteri, P. acidilactici, and P. equi harbored multiple phytase gene copies (Supplementary Table 2). While such redundancy suggests functional importance in the gut environment, it also offers a diverse genetic resources for protein engineering and biotechnology.

The majority of identified phytase sequences belonged to the histidine acid phosphatase (HAPhy) family, while a smaller proportion were classified as protein tyrosine phosphatase-like phytases (PTPLPhy) (Fig. 1, Fig. 2, and Supplementary Table 4). This distribution aligns with previous reports identifying HAPhys as the most prevalent and industrially exploited phytase class, largely due to their high specific activity and established efficacy as feed additives.60 Notably, while commercial phytases are predominantly HAPhys, PTPLPhys possess distinct biochemical features—such as unique substrate specificity and narrow pH optima—that may offer complementary advantages under specific gastrointestinal conditions.26 By categorizing the horse-derived sequences into these two distinct families, we provided a systematic structural filter to prioritize candidates. For instance, the identification of HAPhys from E. coli allows for a direct comparison with known industrial benchmarks, while the PTPLPhys from probiotic sources offer novel, non-pathogenic templates for potential “whole-cell” additive development. This structural classification is a critical first step in narrowing down the 69 identified sequences to the most biologically and industrially viable candidates.

To evaluate the putative suitability of horse feces-associated phytases for monogastric feed applications, signal peptide prediction and physicochemical characterization were integrated via principal component analysis (PCA). The presence of Sec/SPI-type signal peptides in several candidates suggests potential secretion into the extracellular milieu—a critical attribute for effective phytate hydrolysis within the digestive tract.40 Key parameters, including molecular weight, isoelectric point, instability index, aliphatic index, and GRAVY, were assessed as computational indicators of enzyme stability and functionality under gastrointestinal conditions.9, 45 PCA effectively summarized this variation, with PC1 serving as a suitability score for ranking these in silico results (Supplementary File 3). Among HAPhys, glucose-1-phosphatase (CAJ1242485.1) from K. pneumoniae JRT4AKP (PC1 = 2.15) and bifunctional acid phosphatase/4-phytase (NHR17779.1) from E. coli W65.1 (PC1 = 1.97) emerged as the highest-ranked candidates based on their predicted physicochemical profiles. Within PTPLPhys, protein tyrosine phosphatase (UNQ40438.1) from P. equi P2117036 (PC1 = 3.38) demonstrated promising structural characteristics. Conversely, CAJ1246072.1 from K. pneumoniae JRT4AKP (PC1 = 0.67)—despite moderate physicochemical scores—lacked a predicted signal peptide, highlighting the importance of multi-parameter filtering to identify sequences with the highest probability of extracellular functionality. These rankings provide a systematic basis for selecting the most viable sequences for future recombinant expression and biochemical validation.

The transition from sequence discovery to commercial feed-additive application is frequently bottlenecked by the safety profiles and regulatory requirements of the host bacterial sources. To streamline this process, we propose a template based framework that harmonizes the biological potential of novel enzymes with monogastric feed standard by categorizing them into three distinct groups: Well-Characterized Industrial Sources, Established Probiotic Sources, and Pathogenic Sources. This template strategy provides a systematic approach to navigating the regulatory landscape for equine-derived phytases.

Several phytase-encoding genes were identified in equine gut–associated E. coli (Table 1). Although some E. coli strains are pathogenic, the appA gene from this species serves as the gold standard for major commercial enzymes, including Quantum®, OptiPhos®, and Finase® EC.4, 7 Our findings suggest the equine microbiome harbors novel variants of this established enzyme class, making them ideal candidates for high-yield recombinant protein production. Alternatively, phytase candidates from L. reuteri, P. acidilactici, and B. pseudolongum represent the most direct route to feed application. These taxa are Generally Recognized as Safe (GRAS) and possess an extensive history as probiotics.27, 50, 61 Their presence in the equine gut may supports the feasibility of utilizing whole-cell or minimally processed additives, which bypasses the need for complex purification—a significant regulatory advantage. However, despite the recognized probiotic status at the species level, strain-specific safety data for these specific sources remain undefined. Consequently, our framework necessitates rigorous screening for predicted virulence factors (Table 3) to ensure absolute safety for the animal host, human consumers, and the environment.

Pathogenic sources identified in this study, including K. pneumoniae, S. enterica, A. baumannii, and P. equi, present significant biosafety concerns in their native form.42, 62, 63, 64, 65 The information on phytases from these taxa remains limited with the exception of K. pneumoniae.25, 66 Modern biotechnology views these sequences primarily as molecular templates. For monogastric animal feed application, the specific phytase gene would be synthesized and expressed within a non-pathogenic, well-characterized industrial host. For example, the 6-phytase gene, sourced from Buttiauxella sp., is heterologously expressed in a Trichoderma reesei production host to yield a recombinant enzyme that optimizes growth performance, bone mineralization, and nutrient bioavailability in nursery and growing swine.67 This approach effectively decouples the enzyme from the donor's virulence factors, a transition essential for securing EFSA or FDA regulatory approval. Consequently, our identification of these phytases serves as a discovery of genetic blueprints for future protein engineering, rather than a recommendation for the use of the native bacteria.

Following phytase classification, PCA analysis, and biosafety filtering, molecular docking was employed to predict the substrate-binding potential of prioritized candidates toward phytic acid (Table 4). Among the modeled structures, NHR17779.1 displayed the strongest predicted binding affinity, followed by CAJ1242485.1 and UNQ40438.1. Conversely, CAJ1246072.1 exhibited a substantially weaker interaction, indicating a potentially limited substrate-binding capacity compared to the other prioritized candidates. However, the affinity differences among the top candidates and their respective reference enzymes fell within the recognized scoring error range of AutoDock Vina59; thus, these results suggest comparable substrate-binding potential rather than definitive catalytic superiority. When integrated with the PCA findings, CAJ1242485.1 and UNQ40438.1 emerge as the most computationally balanced candidates, as they combine favorable predicted physicochemical properties and secretory signals with robust structural interaction profiles. These integrated in silico data provide a rational prioritization for selecting sequences for future enzyme kinetics and activity assays.

Although this study is entirely computational, the integration of genome mining, comparative sequence analysis, physicochemical profiling, biosafety analysis, structural modeling, and molecular docking provides a systematic framework for the rational prioritization of enzyme candidates. The identification of putative phytases from horse feces–associated bacteria significantly expands the known repertoire of enzymes and highlights the equine gut microbiome as a valuable reservoir of biotechnologically relevant genetic diversity. Nevertheless, we emphasize that these in silico predictions serve as a necessary precursor to, and not a substitute for, laboratory validation. Future studies employing recombinant expression, biochemical characterization, and feeding trials are essential to confirm catalytic efficiency, pH and thermal stability, and overall performance in monogastric production systems. Such validation will be critical to translate these genomic discoveries into functional, safe, and stable feed additives within the scope of sustainable biotechnology.

5. Conclusion

This study systematically mined 162 horse feces–associated bacterial genomes, identifying 69 putative phytases for monogastric feed applications. By integrating physicochemical profiling, biosafety analysis, and AlphaFold2-based molecular docking, we established a computational pipeline to prioritize two lead candidates: UNQ40438.1 (P. equi) and CAJ1242485.1 (K. pneumoniae). These sequences exhibited optimal predicted attributes, including favorable secretory signals and robust substrate-binding profiles. While these in silico findings provide a rational framework for selection, experimental validation remains essential. Ultimately, this work highlights the equine gut microbiome as a rich reservoir of genetic diversity, offering a high-priority library for future protein engineering. Our identification of these enzymes serves as a discovery of “genetic blueprints” for sustainable feed biotechnology. While probiotic-derived variants offer potential as whole-cell additives, this framework primarily emphasizes the recombinant potential of these sequences rather than a recommendation for the use of native bacteria in their natural state.

The following are the supplementary data related to this article.

Supplementary Fig. S1.

IMAGE 1

Models horse feces-associated phytases superimposed with experimental 3D structure counterparts. (A) Alphafold2 model of glucose-1-phosphatase of K. pneumoniae JRT4AKP (CAJ1242485.1) superimposed with 2WNH chain A; (B) Alphafold2 model of bifunctional acid phosphatase/4-phytase of E. coli 651 (NHR17779.1) superimposed with 7Z2S structure. Alphafold2 model are cyan and experimental models are gold. (C) Alphafold2 model of protein tyrosine phosphate of P. equi P2117036 (UNQ40438.1) and Alphafold2 model of hypothetical protein of K. pneumoniae JRT4AKP (CAJ1246072.1) superimposed with 2B4P chain. Model of UNQ40438.1 is cyan, model of CAJ1246072.1 is blue and experimental model (2B4P) is gold.

Supplementary material 2

Calculation for docking box center from active site residues.

mmc2.xlsx (15.2KB, xlsx)
Supplementary material 3

Supplementary File 2. Calculation for docking box center from active site residues.

mmc3.zip (6.2KB, zip)
Supplementary material 4

Supplementary Table 1. Metadata of horse feces-associated bacterial genomes mined for phytase.

mmc4.xlsx (25.7KB, xlsx)
Supplementary material 5

Supplementary Table 2. Total identified phytases from genomes of horse feces-associated bacteria.

mmc5.xlsx (34.6KB, xlsx)
Supplementary material 6

Supplementary Table 3. Cluster of 100% identiy of phytases identified from genomes of horse feces-associated bacteria

mmc6.xlsx (19.7KB, xlsx)
Supplementary material 7

Supplementary Table 4. Conserved domains and famlies of phytases identified from genomes of horse feces-associated bacteria.

mmc7.xlsx (15.6KB, xlsx)
Supplementary material 8

Supplementary Table 5: Predicted signal peptide of horse-feces-associated bacterial phytase

mmc8.xlsx (15.7KB, xlsx)
Supplementary material 9

Supplementary Table 6: Physicochemical properties of horse feces-associated bacterial phytases.

mmc9.xlsx (14.8KB, xlsx)
Supplementary material 10

Supplementary Table 7: The result of principal component analysis of histidine acid phytases (HAPhy) and protein tyrosine like phytases (PTPLPhys) identified from genomes of horse feces-associated bacteria.

mmc10.xlsx (16.9KB, xlsx)

Declaration of generative AI and AI-assisted technologies in the manuscript preparation process

During the preparation of this work the author(s) used ChatGPT in order to get required command or script for data downloading, combining, and extracting, and bioinformatics analysis. After using this tool/service, the author(s) reviewed and edited the content as needed and take(s) full responsibility for the content of the published article.

CRediT authorship contribution statement

Olyad Erba Urgessa: Writing – original draft, Methodology, Formal analysis, Data curation, Conceptualization. Mesfin Tafesse Gemeda: Writing – review & editing, Methodology, Data curation, Conceptualization. Hunduma Dinka: Writing – review & editing, Methodology, Data curation, Conceptualization. Firayad Girma Abdi: Formal analysis, Data curation, Conceptualization. Ketema Tafess: Writing – review & editing, Validation, Methodology, Data curation, Conceptualization.

Declaration of competing interest

The authors declare that they have no known competing financial interests or personal relationships that could have appeared to influence the work reported in this paper.

Acknowledgement

The authors acknowledge the Haramaya, Adama Science and Technology, and Addis Ababa Science and Technology Universities for providing the research facilities.

The authors declare no conflict of interest.

References

  • 1.Outchkourov N., Petkov S. Industrial Enzyme Applications. John Wiley & Sons Ltd; 2019. Phytases for feed applications; pp. 255–285. [DOI] [Google Scholar]
  • 2.Humer E., Schwarz C., Schedle K. Phytate in pig and poultry nutrition. J Anim Physiol Anim Nutr (Berl) 2015;99(4):605–625. doi: 10.1111/jpn.12258. [DOI] [PubMed] [Google Scholar]
  • 3.Ravindran V., Ravindran G., Sivalogan S. Total and phytate phosphorus contents of various foods and feedstuffs of plant origin. Food Chem. 1994;50(2):133–136. doi: 10.1016/0308-8146(94)90109-0. [DOI] [Google Scholar]
  • 4.Dersjant-Li Y., Awati A., Schulze H., Partridge G. Phytase in non-ruminant animal nutrition: a critical review on phytase activities in the gastrointestinal tract and influencing factors. J Sci Food Agric. 2015;95:878–896. doi: 10.1002/jsfa.6998. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 5.Tardiolo G., La Fauci D., Riggio V., et al. Gut microbiota of ruminants and monogastric livestock: an overview. Animals. 2025;15(5):1–28. doi: 10.3390/ani15050758. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6.Abbasi F., Fakhur-un-Nisa T., Liu J., Luo X., Hussain I., Abbasi R. Low digestibility of phytate phosphorus, their impacts on the environment, and phytase opportunity in the poultry industry. Environ Sci Pollut Res. 2019;26:9469–9479. doi: 10.1007/s11356-018-4000-0. [DOI] [PubMed] [Google Scholar]
  • 7.Urgessa O.E., Koyamo R., Dinka H., Tefese K., Gemeda M.T. Review on Desirable Microbial Phytases as a Poultry Feed Additive : Their Sources, Production, Enzymatic Evaluation, Market Size, and Regulation. Int J Microbiol. 2024;2024 doi: 10.1155/2024/9400374. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 8.Qin Tan W., Yee P.C., Chin S.C., et al. Cloning of a novel phytase from an anaerobic rumen bacterium, Mitsuokella jalaludinii, and its expression in Escherichia coli. J Integr Agric. 2015;14(9):1816–1826. doi: 10.1016/S2095-3119(14)60960-6. [DOI] [Google Scholar]
  • 9.Urgessa O.E., Tulu K.T., Gemeda M.T., Dinka H. In silico analysis of insect-associated bacterial phytases reveals optimal biochemical properties and function in poultry gut condition. Bioinforma Adv. 2025;5(1) doi: 10.1093/bioadv/vbaf256. vbaf256. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 10.Kara K., Altınsoy A. Comparison of forages’ digestion levels for different in vitro digestion techniques in horses. Vet Med Sci. 2024;10(2):1–13. doi: 10.1002/vms3.1373. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 11.Brinkley-Bissinger K., Cersosimo L.M., Sullivan K.E., et al. Phytate supplementation has minimal impact on mineral digestibility in horses. J Anim Sci. 2020;98 [Google Scholar]
  • 12.Lavin T.E., Nielsen B.D., Zingsheim J.N. Effects of phytase supplementation in mature horses fed alfalfa hay and pelleted concentrate diets. J Anim Sci. 2013;91:1719–1727. doi: 10.2527/jas2012-5081. [DOI] [PubMed] [Google Scholar]
  • 13.Dersjant-Li Y., Davin R., Christensen T., Kwakernaak C. Effect of two phytases at two doses on performance and phytate degradation in broilers during 1-21 days of age. PLoS One. 2021;16:1–12. doi: 10.1371/journal.pone.0247420. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 14.Rosenfelder-Kuon P., Klein N., Zegowitz B., et al. Phytate degradation cascade in pigs as affected by phytase supplementation and rapeseed cake inclusion in corn-soybean meal-based diets. J Anim Sci. 2020;98(3) doi: 10.1093/jas/skaa053. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15.Shah A.M., Zhang H., Shahid M., et al. The vital roles of agricultural crop residues and Agro-industrial by-products to support sustainable livestock productivity in subtropical regions. Animals. 2025;15(8):1184. doi: 10.3390/ani15081184. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16.Ugbaja S.C., Mushebenge A.G., aimé, Kumalo H, Ngcobo M, Gqaleni N. Potential benefits of in silico methods : a promising alternative in natural compound ’ s drug discovery and repurposing for HBV therapy. Pharmaceuticals. 2025;8(3) doi: 10.3390/ph18030419. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 17.Yu B. In: Clinical Bioinformatics. 141st ed. Trent R.J.A., editor. Humana Press; 2008. In silico gene discovery; pp. 1–22. [DOI] [Google Scholar]
  • 18.Ferreira R.C., Tavares M.P., Morgan T., et al. Genome-scale characterization of fungal phytases and a comparative study between Beta-propeller phytases and histidine acid phosphatases. Appl Biochem Biotechnol. 2020;192(1):296–312. doi: 10.1007/s12010-020-03309-7. [DOI] [PubMed] [Google Scholar]
  • 19.Pramanik K., Kundu S., Banerjee S., Kumar P., Tushar G., Maiti K. Computational-based structural, functional and phylogenetic analysis of Enterobacter phytases. 3 Biotech. 2018;8(6):1–12. doi: 10.1007/s13205-018-1287-y. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 20.Nezhad N.G., Noor R., Raja Z., et al. Integrative structural and computational biology of phytases for the animal feed industry. Catalysts. 2020;10(844):1–24. [Google Scholar]
  • 21.Leary N.A.O., Cox E., Holmes J.B., et al. Exploring and retrieving sequence and metadata for species across the tree of life with NCBI datasets. Sci Data. 2024;11(1):732. doi: 10.1038/s41597-024-03571-y. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22.Urgessa O.E., Tulu K.T., Gemeda M.T., Dinka H., Abdi F.G. Zenodo; 2026. Annotated proteomes from 162 horse feces-associated bacteria (v.0.0) [Data set] Published online 2025. [DOI] [Google Scholar]
  • 23.Golovan S., Wang G., Zhang J., Forsberg C.W. Characterization and overproduction of the Escherichia coli appA encoded bifunctional enzyme that exhibits both phytase and acid phosphatase activities. Can J Microbiol. 2000;46(1):59–71. doi: 10.1139/cjm-46-1-59. [DOI] [PubMed] [Google Scholar]
  • 24.Rix G.D., Sprigg C., Whitfield H., Id A.M.H., Todd D., Id C.A.B. Characterisation of a soil MINPP phytase with remarkable long-term stability and activity from Acinetobacter sp. PLoS One. 2022;17(8):1–24. doi: 10.1371/journal.pone.0272015. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 25.Sajidan A., Farouk A., Greiner R., Jungblut P., Müller E.C., Borriss R. Molecular and physiological characterisation of a 3-phytase from soil bacterium Klebsiella sp. ASR1. Appl Microbiol Biotechnol. 2004;65:110–118. doi: 10.1007/s00253-003-1530-1. [DOI] [PubMed] [Google Scholar]
  • 26.Sharma R., Kumar P., Kaushal V., Das R., Kumar Navani N. A novel protein tyrosine phosphatase like phytase from lactobacillus fermentum NKN51: cloning, characterization and application in mineral release for food technology applications. Bioresour Technol. 2018;249:1000–1008. doi: 10.1016/j.biortech.2017.10.106. [DOI] [PubMed] [Google Scholar]
  • 27.Tamayo-Ramos J.A., Sanz-Penella J.M., Yebra M.J., Monedero V., Haros M. Novel Phytases from Bifidobacterium pseudocatenulatum ATCC 27919 and Bifidobacterium longum subsp. infantis ATCC 15697. Appl Environ Microbiol. 2012;78(14):5013–5015. doi: 10.1128/AEM.00782-12. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 28.Wang J., Chitsaz F., Derbyshire M.K., et al. The conserved domain database in 2023. Nucleic Acids Res. 2023;(51):384–388. doi: 10.1093/nar/gkac1096. December 2022. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 29.Madeira F., Madhusoodanan N., Lee J., et al. The EMBL-EBI job dispatcher sequence analysis tools framework in 2024. Nucleic Acids Res. 2024;52(W1):W521–W525. doi: 10.1093/nar/gkae241. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 30.Teufel F., Almagro Armenteros J.J., Johansen A.R., et al. SignalP 6.0 predicts all five types of signal peptides using protein language models. Nat Biotechnol. 2022;40(7):1023–1025. doi: 10.1038/s41587-021-01156-3. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 31.Community T.G. The galaxy platform for accessible, reproducible, and collaborative data analyses: 2024 update. Nucleic Acids Res. 2024;52(W1):W83–W94. doi: 10.1093/nar/gkae410. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 32.Choudhuri SBTB for B . Ioinformatics for Beginners. Academic Press; 2014. Sequence Alignment and Similarity Searching in Genomic Databases: BLAST and FASTA; pp. 133–155. [DOI] [Google Scholar]
  • 33.Silvanovich A., Bannon G., McClain S. The use of E-scores to determine the quality of protein alignments. Regul Toxicol Pharmacol. 2009;54(3 Suppl):S26–S31. doi: 10.1016/j.yrtph.2009.02.004. [DOI] [PubMed] [Google Scholar]
  • 34.Mirdita M., Schütze K., Moriwaki Y., Heo L. ColabFold : making protein folding accessible to all. Nat Methods. 2022;19:679–682. doi: 10.1038/s41592-022-01488-1. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 35.Bugnon M., Röhrig U.F., Goullieux M., et al. SwissDock 2024: major enhancements for small-molecule docking with attracting cavities and AutoDock Vina. Nucleic Acids Res. 2024;52(W1):W324–W332. doi: 10.1093/nar/gkae300. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 36.O’Boyle N.M., Banck M., James C.A., Morley C., Vandermeersch T., Hutchison G.R. Open babel: an open chemical toolbox. Aust J Chem. 2011;3:33. doi: 10.1186/1758-2946-3-33. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 37.Acquistapace I.M., Thompson E.J., Kühn I., Bedford M.R., Brearley C.A., Hemmings A.M. Insights to the structural basis for the stereospecificity of the Escherichia coli phytase. AppA Int J Molecilar Sci. 2022;23 doi: 10.3390/ijms23116346. doi: 10.3390/ijms23116346 Academic. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 38.Bruder L.M., Gruninger R.J., Cleland C.P., Mosimann S.C. cro Bacterial PhyA protein-tyrosine phosphatase-like myo -inositol phosphatases in complex with the Ins (1, 3, 4, 5) P 4 and Ins (1, 4, 5) P 3 second messengers. J Biol Chem. 2017;292(42):17302–17311. doi: 10.1074/jbc.M117.787853. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 39.Greiner R., Konietzny U., Jany K.D. Purification and characterization of two phytases from Escherichia coli. Arch Biochem Biophys. 1993;303(1):107–113. doi: 10.1006/abbi.1993.1261. [DOI] [PubMed] [Google Scholar]
  • 40.Navone L., Vogl T., Luangthongkam P., et al. Synergistic optimisation of expression, folding, and secretion improves E. coli AppA phytase production in Pichia pastoris. Microb Cell. 2021:1–14. doi: 10.1186/s12934-020-01499-7. Fact Published online. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 41.Jain J., Sapna Singh B. Characteristics and biotechnological applications of bacterial phytases. Process Biochem. 2016;51(2):159–169. doi: 10.1016/j.procbio.2015.12.004. [DOI] [Google Scholar]
  • 42.Villamizar G.A.C., Nacke H., Griese L., Tabernero L., Funkner K., Daniel R. Characteristics of the first Protein tyrosine phosphatase with phytase activity from a. Genes (Basel) 2019:10. doi: 10.3390/genes10020101. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 43.Lee C.H. A simple outline of methods for Protein isolation and purification. Endocrinol Metab (Seoul, Korea) 2017;32(1):18–22. doi: 10.3803/EnM.2017.32.1.18. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 44.Wang X., Upatham S., Panbangred W., et al. Purification, characterization, gene cloning and sequence analysis of a phytase from Klebsiella pneumoniae subsp. pneumoniae XY-5. Sci. Asia. 2004;30:383–390. [Google Scholar]
  • 45.Tomii K. In: Encyclopedia of Bioinformatics and Computational Biology. Ranganathan S., Gribskov M., Nakai K., Schönbach CBTE of B and CB, editors. Academic Press; 2019. Protein Properties; pp. 28–33. [DOI] [Google Scholar]
  • 46.Greiner R. Limitations of an in vitro model of the poultry digestive tract on the evaluation of the catalytic performance of phytases. J Sci Food Agric. 2020;101:2519–2524. doi: 10.1002/jsfa.10878. [DOI] [PubMed] [Google Scholar]
  • 47.Solovyev M.M. Dependence of pH values in the digestive tract of freshwater fishes on some abiotic and biotic factors. Hydrobiologia. 2017 doi: 10.1007/s10750-017-3383-0. Published online. [DOI] [Google Scholar]
  • 48.Ikai A. Thermostability and aliphatic index of globular proteins. J Biochem. 1980;88(6):1895–1898. doi: 10.1093/oxfordjournals.jbchem.a133168. [DOI] [PubMed] [Google Scholar]
  • 49.Yao M.Z., Wang X., Wang W., Fu Y.J., Liang A.H. Improving the thermostability of Escherichia coli phytase, appA, by enhancement of glycosylation. Biotechnol Lett. 2013;35:1669–1676. doi: 10.1007/s10529-013-1255-x. [DOI] [PubMed] [Google Scholar]
  • 50.Bhagat D., Raina N., Ku A., Katoch M., Khajuria Y. Probiotic properties of a phytase producing Pediococcus acidilactici strain SMVDUDB2 isolated from traditional fermented cheese product. Kalarei Sci Rep. 2020;10:1926–1937. doi: 10.1038/s41598-020-58676-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 51.Iravani S., Aziz-Aliabadi F., Vakili R. Feed processing: a review of the impacts of conditioning time and temperature on feed quality and broilers performance. Worlds Poult Sci J. 2024;80(3):657–679. doi: 10.1080/00439339.2024.2341276. [DOI] [Google Scholar]
  • 52.Wu T.H., Chen C.C., Cheng Y.S., et al. Improving specific activity and thermostability of Escherichia coli phytase by structure-based rational design. J Biotechnol. 2014;175:1–6. doi: 10.1016/j.jbiotec.2014.01.034. [DOI] [PubMed] [Google Scholar]
  • 53.Wang Y., Zhang H., Cao N., Qi B., Zhao F., Xie J. Effects of coating on recovery of Escherichia coli-derived phytase under different steam pelleting conditions. Transl Anim Sci. 2025;9 doi: 10.1093/tas/txaf035. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 54.Ingram D.L., Legge K.F. Variations in deep body temperature in the young unrestrained pig over the 24 hour period. J Physiol. 1970;210(4):989–998. doi: 10.1113/jphysiol.1970.sp009253. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 55.Troxell B., Petri N., Daron C., et al. Poultry body temperature contributes to invasion control through reduced expression of Salmonella Pathogenicity Island genes in Salmonella enterica serovars typhimurium and enteritidis. Appl Environ Microbiol. 2015;81(23):8192–8201. doi: 10.1128/AEM.02622-15.Editor. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 56.Volkoff H., Rønnestad I. Effects of temperature on feeding and digestive processes in fish. Temperature. 2020;7(4):307–320. doi: 10.1080/23328940.2020.1765950. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 57.Kyte J., Doolittle R.F. A simple method for displaying the hydropathic character of a protein. J Mol Biol. 1982;157(1):105–132. doi: 10.1016/0022-2836(82)90515-0. [DOI] [PubMed] [Google Scholar]
  • 58.Carugo O., Pongor S. A normalized root-mean-square distance for comparing protein three-dimensional structures. Protein Sci. 2001;10(7):1470–1473. doi: 10.1110/ps.690101. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 59.Trott O., Olson A.J. AutoDock Vina: improving the speed and accuracy of docking with a new scoring function, efficient optimization, and multithreading. J Comput Chem. 2010;31(2):455–461. doi: 10.1002/jcc.21334. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 60.Kumar V., Singh G., Verma A.K., Agrawal S. In silico characterization of histidine acid phytase sequences. Enzyme Res. 2012;2012 doi: 10.1155/2012/845465. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 61.Ali S., Lee E. Bee. Lim S kyung, Suk K. solation and Identification of Limosilactobacillus reuteri PSC102 and Evaluation of Its Potential Probiotic, Antioxidant, and Antibacterial Properties. Antioxidant. 2023;12(2) doi: 10.3390/antiox12020238. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 62.De Cesare A. In: Biological Emerging Risks in Foods. Rodríguez-Lázaro DBTA in F and NR, editor. Vol. 86. Academic Press; 2018. Salmonella in foods: A reemerging problem; pp. 137–179. [DOI] [PubMed] [Google Scholar]
  • 63.Howard A., O’Donoghue M., Feeney A., Sleator R.D. Acinetobacter baumannii: an emerging opportunistic pathogen. Virulence. 2012;3(3):243–250. doi: 10.4161/viru.19700. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 64.Jin S.S., Wang W.Q., Jiang Y.H., Yu Y.T., Wang R.L. A comprehensive overview of Klebsiella pneumoniae: resistance dynamics, clinical manifestations, and therapeutic options. Infect Drug Resist. 2025;18(March):1611–1628. doi: 10.2147/IDR.S502175. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 65.Vanniasinkam T., Giles C., Ndi S., Heuzenroeder M., Barton M. Vaccine Development for Prescottella Equi. Procedia Vaccinol. 2014;8:50–57. doi: 10.1016/j.provac.2014.07.009. [DOI] [Google Scholar]
  • 66.Houshyar M., Saki A.A., Alikhani M.Y., Bedford M.R., Soleimani M., Kamarehei F. Approaches to determine the effi ciency of novel 3-phytase from Klebsiella pneumoniae and commercial phytase in broilers from 1 to 14 d of age. Poult Sci. 2022;102(11) doi: 10.1016/j.psj.2023.103014. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 67.Jlali M., Hincelin C., Manceaux C., Ozbek S. Efficacy of a new biosynthetic bacterial 6-phytase on growth performance, bone mineralization, and nutrient digestibility in nursery and growing pigs. Front Vet Sci. 2025;12 doi: 10.3389/fvets.2025.1591214. [DOI] [PMC free article] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Supplementary material 2

Calculation for docking box center from active site residues.

mmc2.xlsx (15.2KB, xlsx)
Supplementary material 3

Supplementary File 2. Calculation for docking box center from active site residues.

mmc3.zip (6.2KB, zip)
Supplementary material 4

Supplementary Table 1. Metadata of horse feces-associated bacterial genomes mined for phytase.

mmc4.xlsx (25.7KB, xlsx)
Supplementary material 5

Supplementary Table 2. Total identified phytases from genomes of horse feces-associated bacteria.

mmc5.xlsx (34.6KB, xlsx)
Supplementary material 6

Supplementary Table 3. Cluster of 100% identiy of phytases identified from genomes of horse feces-associated bacteria

mmc6.xlsx (19.7KB, xlsx)
Supplementary material 7

Supplementary Table 4. Conserved domains and famlies of phytases identified from genomes of horse feces-associated bacteria.

mmc7.xlsx (15.6KB, xlsx)
Supplementary material 8

Supplementary Table 5: Predicted signal peptide of horse-feces-associated bacterial phytase

mmc8.xlsx (15.7KB, xlsx)
Supplementary material 9

Supplementary Table 6: Physicochemical properties of horse feces-associated bacterial phytases.

mmc9.xlsx (14.8KB, xlsx)
Supplementary material 10

Supplementary Table 7: The result of principal component analysis of histidine acid phytases (HAPhy) and protein tyrosine like phytases (PTPLPhys) identified from genomes of horse feces-associated bacteria.

mmc10.xlsx (16.9KB, xlsx)

Articles from Journal of Genetic Engineering & Biotechnology are provided here courtesy of Academy of Scientific Research and Technology, Egypt

RESOURCES