Skip to main content
ACS Omega logoLink to ACS Omega
. 2026 Sep 6;11(37):54936–54949. doi: 10.1021/acsomega.5c12875

Machine Learning Guided Discovery of Microbiome Metabolites That Inhibit HDAC

Hong A Chung †, Emily M Eshleman ‡,§, James Carter †, Theresa Alenghat ‡,§,*, Daniel Reker †,#,∥,*
PMCID: PMC13613474  PMID: 42798903

Abstract

Histone deacetylase (HDAC) is a family of key epigenetic regulator implicated in inflammation, metabolism, and cancer. Microbiome metabolites can modulate host histone acetylation, yet systematic identification of metabolites that target HDAC remains limited. Here we present a computationally driven discovery pipeline that combines training data composition optimization with experimental validation to prioritize microbiome-derived HDAC inhibitors. Starting from a large, public HDAC3 screening data set (314,129 compounds; 485 actives), we developed an automated iterative sampling strategy that balances active and inactive compounds while enriching the inactive class for metabolite-like chemistry. Models trained on the balanced metabolite-enriched subsets achieved substantially higher sensitivity and balanced accuracy than models trained on the full data set. Consensus predictions from multiple optimized runs were applied to a curated microbiome metabolite database to prioritize candidates for testing. Two top candidates, 5-(hydroxymethyl)­furoic acid (HMFA) and d-glucuronolactone (DGL), were evaluated using an HDAC3-specific biochemical assay, followed by broader HDAC activity assessment. Both metabolites exhibited mild inhibitory activity overall. Docking simulations suggested that HMFA may interact with the HDAC3 catalytic site. Together, we show that targeted training set composition can improve machine learning-assisted discovery of microbiome-derived small molecule inhibitors and identified HMFA as a microbiome-associated metabolite with measurable in vitro HDAC inhibitory activity.


graphic file with name ao5c12875_0008.webp


graphic file with name ao5c12875_0006.webp

Introduction

Histone deacetylase (HDAC) is a family of epigenetic enzyme that removes acetyl groups from lysine residues on histones and nonhistone proteins, promoting chromatin compaction, transcriptional repression, and regulation of gene expression. , The human HDAC family comprises 18 enzymes divided into four classes. Classes I (HDAC1–3, 8), II (HDAC4–7, 9, 10), and IV (HDAC11) are zinc-dependent classical HDACs, while Class III sirtuins (SIRT1–7) are NAD+-dependent deacetylases with distinct catalytic mechanisms. Abnormal HDAC activity has been linked to various diseases, including cancer, neurodegeneration, metabolic disorders, and inflammatory conditions. − Therefore, HDAC inhibition has been recognized as a promising therapeutic approach for modulating epigenomic regulation in multiple pathological contexts. ,

Several small-molecule HDAC inhibitors have received FDA approval, including vorinostat, romidepsin, belinostat, and panobinostat. , While these compounds demonstrated clinical benefits for certain hematologic malignancies, their therapeutic use remains limited because they are associated with substantial systemic toxicities such as fatigue, gastrointestinal toxicity, thrombocytopenia, and cardiac arrhythmias. , This motivates the identification of alternative HDAC inhibitors that could potentially provide superior safety profiles and thereby could be applied more broadly across diseases.

Microbiome-derived metabolites may offer a compelling starting point for the development of alternative HDAC modulators. Many microbial metabolites are predicted to have substantially lower toxicity compared with synthetic drugs, as shown in our prior work using large-scale toxicity prediction across microbiome metabolites. Many are restricted to the gut or display tissue-biased biodistributions, which may reduce systemic side effects. In addition, microbiome-derived metabolites provide mechanistic insights into host–microbiome communication because epigenetic regulation represents a major interface between diet, microbial metabolism, and host physiology.

Emerging evidence suggests that the microbiome can influence host epigenomic regulation through its metabolites. − Several microbial metabolites have been shown to modulate histone acetylation and chromatin remodeling. For example, short-chain fatty acids such as butyrate, propionate, and acetate act as natural HDAC modulators and contribute to intestinal homeostasis and immune regulation. − Among these, butyrate is one of the most well-characterized microbial HDAC inhibitors and is frequently used as a positive control in HDAC inhibition assays because of its robust and reproducible effects on histone acetylation. Although butyrate is a relatively weak biochemical HDAC inhibitor in vitro, its mild, localized, and physiologically tuned epigenetic modulation in vivo has been found to exert beneficial effects, supporting intestinal barrier maintenance, anti-inflammatory signaling, and colorectal cancer protection. ,

Recent work additionally suggests that certain microbiome-derived metabolites might also act as activators of HDACs activity. For example, inositol-1,4,5-trisphosphate (InsP3), a product of bacterial phytate metabolism, was shown to activate HDAC3 and promote intestinal epithelial repair. These findings highlight the potential bidirectional influence of specific microbial metabolites on HDAC function and emphasize the microbiome’s capacity to fine-tune host epigenetic regulation. However, the full range of microbiome-derived metabolites that modulate HDAC remains unknown. Systematic experimental testing of thousands of metabolites is infeasible, particularly because many metabolites are difficult to isolate or synthesize. This motivates a computational approach to prioritize candidates, enabling efficient identification of microbiome metabolites with predicted HDAC-modulating potential for targeted experimental validation.

In this study, we aimed to discover microbiome metabolites that can modulate HDAC activity by developing a machine learning framework tailored specifically to microbiome metabolite chemistry. Identifying such metabolites could reveal mechanistic links between microbial composition and host epigenetic regulation, offering new insights into microbiome-driven phenotypes. Moreover, these metabolites may serve as valuable starting points for future metabolite-inspired drug development targeting epigenetic pathways. While HDAC isoform-selective data are limited, a large, high-quality HDAC3 biochemical screening data set (n = 314,129) is available from a single high-throughput screening campaign, providing an attractive resource for supervised learning. We here aimed to combine machine learning on this data set with HDAC3-specific and broader HDAC activity profiling in vitro to identify microbial metabolites with HDAC modulating function.

To do so, machine learning models need to be fine-tuned to be able to extrapolate from synthetic drug-like chemicals toward metabolite phenotypes that often differ in their chemical makeup. , To this end, we hypothesized that adjusting the training data composition to better reflect the structural chemistry of microbial metabolites would improve prediction accuracy and lead to the identification of novel bioactive candidates. As a proof-of-concept for such a workflow, we rebalanced the training data set and enriched it with compounds that share higher chemical Tanimoto similarity with known metabolite structures. Known microbiome metabolites identified in the data set were held out as a validation set, all of which were inactive, to test the model’s generalizability.

Building on the success of this rebalancing approach, we then applied an automated data composition optimization method. Through iterative addition of compounds and continuous performance tracking, the workflow revealed an early optimal region (approximately round 257, total n = 515) at which predictive power stabilized, achieving performance comparable to or better than the manually rebalanced metabolite enriched model while requiring fewer training samples. Two candidate metabolites, 5-(hydroxymethyl)­furoic acid (HMFA) and d-glucuronolactone (DGL), were selected, and both were tested in vitro. Both compounds were first evaluated in an HDAC3-specific biochemical assay using recombinant HDAC3–NCoR complex and subsequently tested in a broader HDAC activity assay using HeLa nuclear extracts. Both metabolites exhibited mild inhibitory activity overall. Molecular docking further supported the potential for HMFA to fit the HDAC3 active site.

Overall, this work demonstrates that tailoring machine learning training data sets to microbiome metabolite chemical space can improve predictive accuracy and facilitate the discovery of novel bioactive microbial metabolites. Combining computational enrichment, predictive modeling, and docking provides an efficient framework for exploring microbiome metabolites that may regulate epigenomic targets such as HDAC, providing a blueprint for the development of more effective computer-assisted workflows to delineate biological effects of metabolite structures.

Results

Manual Subsampling of the Original Train Set

The original public HDAC3 high-throughput screening (HTS) data set comprised 314,129 compounds, including 485 actives and 313,644 inactives (0.15% actives in total), demonstrating a strong class imbalance typical of large-scale biochemical screening data (Figure A). To assess whether the training data set contained any overlap with known microbiome metabolites, we cross-referenced the compound library with our curated microbiome metabolite database. This comparison identified 162 overlapping compounds representing known microbiome metabolites that were assessed for their HDAC3 activity, all of which were labeled as inactive in the screening data. These known microbiome metabolites were reserved as part of the validation set to evaluate model generalizability toward metabolite-like chemistry. Because no active microbiome metabolites were available in the data set, a small subset of active nonmetabolite compounds was included in the validation set to enable evaluation of model sensitivity and accuracy. Specifically, the validation set comprised 162 inactive metabolites, supplemented with 10% of the active nonmetabolite compounds (n = 48) and an equal number of inactive nonmetabolite compounds (n = 48), resulting in a total validation set of 258 molecules.

1.

1

Manual subsampling of the HDAC3 training data set and model performance comparison. (A) Schematic of the data set composition and manual subsampling strategy. The original HDAC3 screening data set contained over 312,000 inactive and 485 active compounds, including a subset of metabolite-like molecules identified by Tanimoto similarity. Known microbiome metabolites (all inactive) were used as a validation set, supplemented with 10% of active nonmetabolite compounds and an equal number of inactive nonmetabolites to balance the set. Multiple training subsets were generated to test the effect of data set composition, including balanced and metabolite-enriched combinations. (B) Radar plots showing the MLP classifier performance across six training set compositions, evaluated using the metabolite validation set. Metrics include F1 score, PR-AUC, ROC-AUC, balanced accuracy, sensitivity, and MCC. The balanced metabolite-enriched subsample outperformed other configurations across most metrics, showing particularly strong gains in sensitivity.

The fact that all 162 overlapping microbiome metabolites were labeled inactive in the HTS further underscores the challenge of identifying microbiome-derived metabolites with measurable HDAC3 activity, as few have been experimentally characterized to date and many tested metabolites show no measurable in vitro activity. Although a small number of microbiome metabolites have been reported to modulate HDAC activity, , these reports used different assays and commonly describe very weak activity that may be considered as “inactive” in typical drug discovery screens.

The remaining compounds were used as the training pool. Initial baseline modelsmultilayer perceptron (MLP), random forest (RF), and XGBoost (XGB)were trained using the full data set. These baseline models exhibited poor performance across all metrics, particularly sensitivity (MLP = 0.12, RF = 0.10, XGB = 0.07) and MCC (MLP = 0.29, RF = 0.27, XGB = 0.23), indicating strong bias toward the majority inactive class (Figure S1). ROC-AUC values were moderate (MLP = 0.74, RF = 0.78, XGB = 0.86), suggesting that while the models could broadly rank molecules by their likelihood of being active or inactive, they failed to correctly classify active compounds due to the severe class imbalance.

To address this issue, we investigated whether subsampling strategies could improve model performance. Subsampling is a well-established approach for handling class imbalance, where the data set composition is adjusted to achieve a more balanced representation between active and inactive compounds. , Common strategies include random undersampling of the majority class and oversampling of the minority class, the latter often implemented using techniques such as the synthetic minority oversampling technique (SMOTE). Importantly, optimal training data sets for molecular machine learning often depend strongly on the data set’s chemical makeup and distribution. , However, random undersampling of inactives might discard a large number of compounds including ones that are structurally relevant for microbiome-related discovery. We hypothesized that balancing the data set by including all actives while more selectively undersampling inactives could enhance sensitivity if using subsampling approaches that are aware of metabolite chemical structures to generate more information-rich training data sets.

To identify training data molecules most similar to metabolites and therefore potentially most informative for a metabolite prediction tool, we identify metabolite-like molecules by computing chemical similarity using the Tanimoto measure on Morgan fingerprints (radius = 2, bit length = 2048). Compounds with a maximum similarity score ≥ 0.6 to any microbiome metabolite were classified as “metabolite-like”. This threshold was chosen because it is commonly employed to capture moderate structural similarity for chemical space enrichment without restricting selection to close analogs. , This process identified 1003 inactive metabolite-like compounds, with no active compounds exceeding the 0.6 threshold. These “metabolite-like” molecules were used to construct six distinct subsampling strategies (Figure A):

  • 1.

    Balanced metabolite-enrichedincluded all actives (n = 437) and an equal number of inactive metabolite-like compounds (n = 437). This strategy ensured both class balance and chemical enrichment toward metabolite-like structures.

  • 2.

    Balanced randomincluded all actives (n = 437) and an equal number of randomly selected inactive compounds (n = 437), serving as a control for balance without structural enrichment.

  • 3.

    Metabolite-enrichedincluded all actives (n = 437) and all inactive metabolite-like compounds (n = 1003 ± 1), thereby maximizing exposure to metabolite-like chemistry while maintaining all known actives. The goal was not to force class balance but to maximize representation of metabolite-like chemistry in the training pool so the model could learn features relevant to microbiome-derived molecules.

  • 4.

    Mixedincluded all actives (n = 437), all inactive metabolite-like compounds (n = 1003 ± 1), and an equal number of nonmetabolite-like inactives (n = 1003). This strategy was termed “mixed” because it intentionally combined metabolite-like and nonmetabolite-like compounds to preserve chemical space coverage while maintaining metabolite representation.

  • 5.

    Random mixedincluded all actives (n = 437) and a randomly selected inactive set (n = 2006) of the same total size as the mixed subset. This strategy controlled for total size effects while lacking metabolite-based enrichment.

  • 6.

    Randomincluded all actives (n = 437) and a larger random subset of inactives (n = 2443), serving as a size-matched but fully random baseline.

For each of the five repeats, the validation set was independently generated by randomly removing 48 inactives regardless of metabolite-likeness. Consequently, the total training set size varied slightly between repeats (typically ±1 compound), while the number of actives remained constant.

Model evaluation results (Figures B and S1) showed substantial improvements after subsampling, particularly in sensitivity. Because our primary objective was to prioritize the identification of potentially active microbiome-derived metabolites from a limited candidate space, we emphasized model sensitivity to maximize recall of true positives, accepting an increased risk of false positives at this exploratory stage. The balanced metabolite-enriched configuration achieved the highest sensitivity across all three algorithms, with the MLP reaching 0.95, RF 0.94, and XGB 0.93. Sensitivity improvement is particularly valuable for inhibitor discovery, as it increases the likelihood of identifying true actives. While the random forest model performed comparably, the MLP achieved the highest overall sensitivity and balanced accuracy, leading us to select it as the primary model for subsequent optimization. The MLP’s advantage likely stems from its capacity to capture complex, nonlinear structure–activity relationships.

Although the balanced metabolite-enriched subset did not always yield the top scores in every metric, it consistently ranked among the highest across ROC-AUC, PR-AUC, MCC, and F1 (Figure S1). For PR-AUC, differences among the top three strategies were not statistically significant, suggesting that the main benefit of the metabolite-enriched training set was improved recall rather than precision.

The superior performance of the metabolite-enriched strategy supports the hypothesis that including structurally metabolite-like inactives helps the model better generalize to microbiome chemical space. Even though the active compounds were not structurally similar to known metabolites, exposure to metabolite-like molecular features during training likely improved the model’s ability to recognize relevant structural motifs associated with microbiome-derived inhibitors.

Automated Data Composition Optimization

Following the manual subsampling results, in which a fixed number of data points was selected based on a fixed threshold, we sought to determine whether an optimal “sweet spot” in training set composition could be identified through an automated, iterative sampling process. This approach systematically explored how incremental inclusion of compounds affected model performance under different selection strategies. Unlike the manual subsampling, which relied on a fixed Tanimoto similarity threshold and intuitive balancing of actives and inactives, the automated procedure optimized both the number and structural multiplicity of included compounds. In particular, we sought to test whether adaptively prioritizing metabolite-like structures, rather than applying a fixed similarity cutoff, could further enhance predictive performance.

The automated data composition optimization was designed as a cyclic process consisting of iterative training, evaluation, and resampling steps (Figure A). Each cycle began with a small “warm start” training set of 10 compounds (5 actives and 5 inactives), drawn from the same pool used in the manual subsampling step. Two additional molecules were added in each round, after which the model was retrained and evaluated on the same fixed validation set used in the manual subsampling experiments. This validation set included all inactive microbiome metabolites and a balanced selection of active and inactive nonmetabolites. This process was repeated until all molecules in the pool set had been used, with performance metrics tracked across rounds.

2.

2

Automated optimization of HDAC3 training data set composition. (A) Overview of the automated data composition optimization workflow. The validation set used here is identical to that in the manual subsampling, while the remaining molecules were assigned to the pool set. Each optimization iteration began with a warm start containing five active and five inactive molecules. Two additional compounds were added per round, followed by model retraining and evaluation on the validation set. This iterative process continued until the best-performing round was identified. (B) Illustration of the top three of six total sampling strategies used for molecule selection. All balanced strategies maintained class balance by adding one active and one inactive molecule per round. The balanced metabolite-enriched strategy prioritized metabolite-like compounds first, the balanced most-metabolite-like strategy sorted candidates by Tanimoto similarity to metabolite-like structures, and the balanced random strategy selected molecules randomly. The iteration points around 430 and 750 indicate depletion of active and metabolite-like compounds from the pool set.

To investigate how selection order influenced model behavior, six molecule-selection strategies were implemented (Figure B):

  • 1.

    Balanced most-metabolite-likeeach round began with equal numbers of active and inactive compounds, prioritizing the inclusion of the most metabolite-like molecules based on decreasing Tanimoto similarity scores. This approach maintained class balance while progressively incorporating the most relevant chemical space for microbiome metabolites.

  • 2.

    Balanced metabolite-enrichedmaintained equal numbers of active and inactive compounds, where inactives were randomly selected from the full set of metabolite-like compounds (similarity ≥0.6). This strategy ensured class balance while maintaining chemical enrichment through random sampling within the metabolite-like pool instead of focusing on the most metabolite-like compounds first.

  • 3.

    Balanced randommaintained class balance between actives and inactives, with inactive compounds added randomly at each iteration without considering their metabolite-likeness. This served as a baseline control to evaluate the effect of metabolite-based prioritization.

  • 4.

    Most-metabolite-likesequentially added compounds with the highest similarity to microbiome metabolites, without enforcing class balance. This maximized focus on metabolite-like chemistry while allowing class imbalance to follow background training data distributions.

  • 5.

    Metabolite-enrichedsimilar to the most-metabolite-like approach but included metabolite-like compounds selected randomly from the full set (similarity ≥0.6). This approach aimed to test whether including a broader, randomly sampled range of metabolite-like molecules could enhance model generalizability despite causing a strong class imbalance.

  • 6.

    Randomadded inactive compounds randomly in each iteration without regard to structural similarity or class balance, serving as an unbiased baseline for assessing the effect of enrichment and balancing strategies.

All balanced strategies maintained class balance by adding one active and one inactive molecule per round until all actives are depleted after which only inactives are added. The balanced metabolite-enriched strategy prioritized metabolite-like inactives first, followed by nonmetabolite-like inactives once the metabolite-like inactive pool was exhausted. The balanced most-metabolite-like strategy selected inactives sorted by decreasing Tanimoto similarity to microbiome metabolites. The balanced random strategy added inactives in random order without structural prioritization. In contrast, the metabolite-enriched, most-metabolite-like, and random strategies followed the natural class distribution of the full training data set, adding molecules without enforcing class balance.

Results of this iterative optimization revealed that the balanced strategies consistently outperformed the unbalanced ones across nearly all metrics (Figures , S2 and S3). In particular, sensitivity peaked early for both balanced metabolite-enriched and balanced most-metabolite-like strategies, indicating that models incorporating metabolite-like chemistry were better able to recognize true actives. The performance trajectory could be divided into three distinct phases: phase 1, corresponding to balanced sampling when both actives and metabolite-like inactives were plentiful; phase 2, marked by emerging imbalance as active compounds became depleted but metabolite-like inactives remained; and phase 3, when both actives and metabolite-like compounds were exhausted and all sampling strategies simply select inactive nonmetabolite-like structures for the remaining sampling process.

3.

3

Automated optimization of HDAC3 training data set composition. Six sampling strategies and three evaluation phases are shown for both panels. Background colors indicate the same phases: phase 1 (balanced sampling), phase 2 (early imbalance with metabolite-like compounds still present), and phase 3 (increased imbalance as metabolite-like compounds deplete). Shaded areas around each line represent standard deviation across five repeats. (A) Model sensitivity during automated optimization across 1000 rounds. The black dashed line represents the full training result. Balanced strategies quickly outperform the full-train model, with the balanced most-met-like and balanced metabolite-enriched strategies achieving comparable performance and converging around rounds 430 and 730, respectively. (B) Sensitivity snapshots across incremental rounds on a logarithmic scale. All strategies converge to the same performance level as the full-train model by the final round since the same final data set composition is used.

To supplement the continuous performance traces that were generated until 1000 data points had been added, we included incremental snapshots of model performance at log-spaced rounds (101, 102, 103, 104, and full) to visualize the long-term behavior of each sampling strategy when sampling beyond 1000 data points (Figures B and S2). These snapshots provide further evidence that balanced strategies rapidly achieve high sensitivity and balanced accuracy during phase 1, while unbalanced strategies exhibit early instability and markedly poorer positive-class performance.

To identify the most data-efficient stopping point, exponential moving average smoothing was applied to detect the earliest top performance across repeats while smoothing out potential noisy performance peaks. The optimal rounds varied across sampling strategies from 112 to 434, with an average of 258 for the balanced metabolite-enriched configuration and 386 for the balanced most-metabolite-like configuration. This variability among runs indicates some stochasticity in the sampling order due to randomness of sampling and the distinct validation set compositions, but consistently shows that both strategies reached optimal performance well before the manual subsampling equivalent (round 432). Models trained at these early rounds achieved comparable or slightly higher sensitivity than those trained on the complete data set, confirming that similar predictive performance can be achieved with substantially fewer compounds. Each model’s earliest best round was therefore used to define its final training subset for subsequent analyses and experimental validation.

Between the two metabolite-enriched strategies, the balanced metabolite-enriched model achieved a higher level of stability and peaked earlier than the balanced most-metabolite-like model. Although the peak sensitivities of these two models were not significantly different (0.97 for balanced metabolite-enriched and 0.98 for balanced most-metabolite-like), the balanced metabolite-enriched approach performed more consistently across other metrics and reached comparable results in fewer iterations. In the balanced metabolite-enriched model, sensitivity remained stable throughout phase 1 and began to decline only after the transition into phase 2, when the class balance started to break due to the depletion of active compounds. For metrics such as ROC-AUC, PR-AUC, F1, and MCC, the most-metabolite-like strategy showed a delayed improvement, suggesting that an overly narrow focus on the most similar structures may limit chemical scope and reduce early model generalization. In contrast, including a broader range of metabolite-like compounds in the balanced metabolite-enriched model may provide a more diverse and informative representation of the chemical space, resulting in faster convergence and greater stability.

While a Tanimoto similarity threshold of 0.6 was used to define metabolite-like compounds in both the manual subsampling and the balanced metabolite-enriched automated strategy, we explicitly compared this threshold-based enrichment with a continuous similarity-ranking approach in the balanced most-metabolite-like strategy. Interestingly, the threshold-based balanced metabolite-enriched configuration achieved earlier convergence and greater stability across repeats. This suggests that defining a moderately similar chemical neighborhood may provide a more diverse and informative representation of microbiome chemical space than focusing exclusively on the highest-similarity structures. Nevertheless, systematic evaluation of alternative similarity thresholds could further refine the identification of optimal enrichment boundaries and represents an avenue for future work.

Although the iterative retraining process is computationally intensive due to the large number of model retraining steps, this approach provides several advantages. It facilitates transparent tracking of performance trends, enables easier data management by identifying the most informative subsets, and reduces redundancy by highlighting compounds that contribute meaningfully to prediction accuracy. Moreover, these smaller, compositionally optimized data sets are easier to share and reproduce across studies, providing practical benefits that may outweigh the additional computational cost. In larger data sets, where noise and data diversity are higher, the iterative retraining approach could also improve the reliability of identifying stable performance peaks, as demonstrated in previous large-scale data optimization studies.

While the balanced most-metabolite-like strategy began to show slightly higher sensitivity during phase 3, this difference emerged only after the depletion of most active and metabolite-like compounds and did not translate into broader performance improvements across other metrics. Notably, this late gain is potentially due to the continued use of similarity-based sorting even beyond the 0.6 metabolite-likeness threshold, which may still prioritize molecules that are closer to metabolite chemical space than a purely random selection. This observation suggests that ranking compounds by decreasing metabolite similarity even for nonmetabolite-like compounds can provide additional relevance for metabolite-focused prediction tasks, even after the pool of explicitly metabolite-like compounds has been exhausted. Nonetheless, both balanced strategies demonstrated comparable consistency across repeats, but the balanced metabolite-enriched configuration achieved high, stable performance earlier and required fewer iterations. Given its faster convergence and similar robustness, the balanced metabolite-enriched strategy was selected for downstream validation experiments.

Overall, the automated optimization confirmed that balanced, metabolite-enriched compositions produced the most robust and sensitive models for predicting microbiome-relevant HDAC inhibitors while offering clear benefits in data efficiency, interpretability, and generalizability compared with manual or random sampling approaches. However, it is important to acknowledge that these benefits were based on retrospective evaluations and might not capture the true benefits of the model. Therefore, we next set out to perform prospective validations by predicting HDAC inhibiting microbiome metabolites and assessing their activity in vitro.

Prospective Validation of Machine Learning Model Predictions

Following model optimization, the automated balanced metabolite-enriched subsampling strategy was selected as the final configuration for prospective experimental testing. For each optimization repeat, the training set corresponding to the earliest best round for sensitivity (as determined by exponential moving average analysis) was used to train an independent model. The resulting models were then applied to the curated microbiome metabolite database, and the predicted probabilities from all runs were averaged to generate a consensus prediction score for each metabolite.

Based on these consensus predictions, together with compound availability, solubility, and microbiome association, two microbiome-derived metabolites were prioritized for in vitro validation: 5-(hydroxymethyl)­furoic acid (HMFA) with a predicted probability of 0.62 and d-glucuronolactone (DGL) with a predicted probability of 0.55. Both compounds were commercially available, water-soluble, and structurally distinct, representing practical test cases for model-predicted HDAC3 inhibitory activity. HMFA can be produced by bacterial oxidation of HMF, and several Bacillus speciesincluding species known to inhabit the human guthave been shown in vitro to catalyze the conversion of HMF to HMFA. Thus, HMFA is a plausible microbiome-derived metabolite (Table S1). − HMFA can also be produced through host metabolism, as mammalian aldehyde dehydrogenases oxidize HMF to its corresponding carboxylic acid, indicating that both host and microbial pathways may contribute to HMFA formation. It has been detected in human biofluids as a marker of microbial or dietary metabolism and has demonstrated bioactivity, such as nematicidal effects, although its role in host epigenetic regulation has not been previously reported. In contrast, DGL is a microbiome-influenced cometabolite, primarily derived from host glucuronic acid metabolism. Serum levels of multiple glucuronide-related metabolites have been shown to correlate with shifts in gut microbial composition, including taxa within the Bacteroidetes, Firmicutes, and Proteobacteria phyla. The gut microbiome plays a well-established role in glucuronide turnover through β-glucuronidase-mediated deconjugation, a process that can recycle glucuronides and modulate the availability of downstream metabolites. , The selection of these two structurally distinct metabolites thus provided an opportunity to assess whether the machine-learning framework could identify microbiome-associated metabolites with measurable HDAC inhibitory activity.

To assess HDAC inhibition, we tested compound activity in two experimental contexts: an HDAC3-specific assay and a broader HDAC activity assay. Because the machine learning model was trained specifically on a high-quality HDAC3 biochemical screening data set, we first evaluated compound activity using a recombinant HDAC3–NCoR complex to directly assess HDAC3-specific inhibition (Figure A). At 10 mM, HMFA and d-glucuronolactone inhibited HDAC3 activity by 29.12% and 13.28%, respectively, while the known metabolite control butyrate inhibited HDAC3 activity by 37.04%. Both HMFA and DGL showed mild but statistically significant HDAC3 inhibition compared with vehicle, and the inhibitory effect of HMFA was not statistically significantly different from that of butyrate at the same concentration. Importantly, 10 mM falls within the range of luminal butyrate concentrations reported in vivo, providing a physiologically relevant context for this comparison. Although HMFA showed numerically greater inhibition than DGL in the HDAC3-specific assay, this difference was not statistically significant.

4.

4

Experimental validation of top metabolite candidates in the HDAC inhibition assay. (A) HDAC3-specific activity measured using recombinant HDAC3–NCoR complex in the presence of the top predicted metabolites, HMFA and DGL, with TSA and butyrate included as reference inhibitor controls (N = 1, n = 3). (B) Unspecific HDAC activity measured in HeLa nuclear extracts at the indicated concentrations. TSA and butyrate served as positive inhibitor controls (N = 2, n = 1–3). Statistical analysis was performed using unpaired t-test compared with vehicle control (*p < 0.05, **p < 0.01, ***p < 0.001, ****p < 0.0001). Hash symbols indicate comparisons versus 1 M HMFA (#p < 0.05, ##p < 0.01, ###p < 0.001, ####p < 0.0001).

We next assessed broader HDAC inhibition using HeLa cell extracts containing endogenous HDAC activity (Figure B). Compounds were tested at 10 mM and at an elevated concentration of 1 M, and HDAC activity was quantified relative to vehicle and known inhibitors. In the HeLa extract assay, both HMFA and DGL reduced HDAC activity in a concentration-dependent manner. At 1 M, HMFA inhibited HDAC activity by approximately 97.1%, whereas DGL showed 30.8% inhibition, and the inhibitory effect of HMFA was significantly greater than that of DGL at this concentration. At 10 mM, HMFA and DGL exhibited 14.7% and 17.8% inhibition, respectively. For comparison, the known HDAC inhibiting metabolite butyrate inhibited HDAC by 89.6% at 1 M and 47.7% at 10 mM in this experiment, whereas the synthetic reference inhibitor trichostatin A (TSA) at 1 mM nearly abolished HDAC activity (90.4% inhibition). Because the recombinant HDAC3 assay and HeLa extract assay represent distinct experimental systems, inhibition levels between the two assays are not directly comparable, but serve as separate confirmations of a mild inhibitory activity of both predicted compounds.

Although the absolute inhibition levels for HMFA and DGL were modest relative to TSA, their effects were within the same range as butyrate, a physiologically relevant microbial HDAC inhibitor. Previous studies have shown that butyrate’s relatively weak inhibition in cell-free assays contrasts with its pronounced in vivo activity. − However, it is important to acknowledge that butyrate’s physiological relevance is driven by local accumulation, and it is not known whether the here investigated metabolites would show the same localized effects in vivo. Nevertheless, these results confirm that the machine-learning model can identify metabolites with measurable in vitro HDAC inhibitory activity to potentially accelerate future investigations of microbial–host interactions.

Docking Simulation Supports the Experimental Results

To gain structural insight into how the top two predicted metabolites might be able to interact with HDACs, we performed molecular docking simulations for HMFA and DGL using the GNINA deep learning docking engine. The HDAC3 crystal structure (PDB ID: 4A69) was used as the receptor model, which includes the inositol-1,4,5,6-tetrakisphosphate (IP4) molecule and the SMRT deacetylase-activating domain (DAD) as part of the native corepressor complex (Figure ). This configuration was selected to maintain the enzyme’s catalytically active conformation and to evaluate how the candidate metabolites might engage the zinc-dependent catalytic pocket.

5.

5

Docking simulation results for HMFA and DGL with HDAC3. (A,D) GNINA docking simulations showing predicted binding poses of 5-(hydroxymethyl)­furoic acid (HMFA) and d-glucuronolactone (DGL) within the HDAC3 active site. Each row shows three visualization levels: the protein backbone structure, the structure with side chains added, and the surface representation. (A) HMFA (purple) is predicted to bind more tightly within the catalytic pocket. HDAC3 is shown in gray, the zinc ion as a red sphere, and IP4, the reference ligand from the HDAC3 crystal structure, is shown as orange and red sticks. (B) Distance between the zinc ion and the predicted binding position of HMFA, calculated as 2.21 Å. (C) Predicted atomic interactions of HMFA within the HDAC3 binding pocket, including hydrogen bonding, π–π stacking, metal coordination, and salt bridge formation. (D) DGL is predicted to bind farther from the zinc-binding site and exhibits weaker predicted interactions compared with HMFA.

Docking analysis predicted that HMFA might bind within the HDAC3 catalytic channel near the zinc ion (Figure A), with a calculated zinc–ligand distance of 2.21 Å (Figure B). Predicted atomic interactions of HMFA within the HDAC3 binding pocket included hydrogen bonding, π–π stacking, metal coordination, and salt bridge formation, suggesting the potential for multiple stabilizing contacts that could contribute to modest inhibitory binding (Figure C). In contrast, DGL was predicted to be positioned farther from the zinc center and predicted to form fewer potentially stabilizing interactions within the pocket (Figure D).

GNINA incorporates deep learning-based scoring functions, and the CNN pose score has been shown to outperform classical Vina-style energy terms in reproducing crystallographic binding modes. This metric evaluates whether a predicted pose resembles true protein–ligand interactions learned from thousands of experimentally solved complexes, making it more reliable for ranking small and chemically diverse molecules. For this reason, we used the CNN pose score as the primary indicator of docking confidence. Both HMFA and DGL achieved high CNN pose scores (0.82 and 0.68; Table S1).

To assess whether these trends extended beyond HDAC3, we performed exploratory docking against two additional class I HDAC isoforms, HDAC2 (PDB ID: 7ZZO) and HDAC8 (PDB ID: 1W22) (Table S1). Across both isoforms, HMFA again exhibited higher CNN pose confidence and comparable or stronger predicted affinities than DGL. These supplementary analyses support the overall pattern that HMFA engages HDAC active sites more reliably than DGL. However, these docking analyses are intended to be hypothesis-generating only and do not provide evidence of binding or isoform selectivity. Experimental validation using purified HDAC isoforms will be required to determine whether these predicted interactions translate into biochemical activity.

Discussion

Model Design and Optimization Insights

Here, we developed a machine learning-based framework to identify in vitro biological effects of microbiome-derived metabolites. Using machine learning-approaches for the identification of such hidden biological effects is heavily investigated area of research given the potential biological impact. By systematically evaluating how training data composition affects predictive performance, we demonstrated that balancing active and inactive compounds while enriching the inactive class with metabolite-like molecules substantially improved model sensitivity and overall accuracy. This work extends conventional subsampling approaches, which are commonly used to address class imbalance but often ignore the molecular characteristics of the selected data points. In contrast, our metabolite-enriched strategy ensures that training compounds reflect microbiome chemical space, allowing models to better capture the subtle structure–activity relationships that characterize natural products. Importantly, these improvements in predictive performance translated into experimentally measurable inhibitory activity in an HDAC3-specific biochemical assay, consistent with the isoform-resolved nature of the training data.

Both manual and automated optimization experiments revealed that training data composition critically shapes model behavior. The threshold-based subsampling approach already produced strong and stable models by maintaining a balanced ratio between actives and metabolite-like inactives, demonstrating that even classic rebalancing with molecular-aware composition adjustments can enhance predictive robustness. The automated iterative optimization extended these insights by dynamically adjusting data set composition. By incrementally introducing new compounds and tracking model performance, this method identified an optimal “sweet spot” where predictive power plateaued early (approximately round 257, resulting total n = 515), achieving comparable or superior performance to the best-performing-manually rebalancing model (balanced-metabolite-enriched, total n = 874) while using only about half the training data.

Three characteristic learning phases characterized the automated process: an early balanced phase where actives and metabolite-like inactives were equally represented, a transitional phase where actives were depleted but metabolite-like inactive molecules continue to be added to the data, and a late phase where both categories were exhausted and the model samples nonmetabolite-like inactive molecules. Model sensitivity consistently peaked at the transition between the first and second phases, highlighting that early exposure to metabolite-like structures is key to maximizing performance. The balanced metabolite-enriched configuration reached this peak faster and required fewer data points than the manually rebalancing approach, indicating that more specific data selection can yield efficient, high-performing models even with smaller training sets.

Experimental Validation and Biological Implications

Experimental testing confirmed that the machine learning model successfully identified microbiome-derived metabolites with modest ability of inhibiting HDAC in vitro. Because the biochemical data used for model training were generated from an HDAC3-targeted assay, experimental validation was designed to first assess HDAC3-specific inhibition using a recombinant HDAC3–NCoR complex, followed by broader HDAC activity profiling in HeLa nuclear extracts. The two top-ranked candidates, HMFA and DGL, both showed measurable inhibition of HDAC in vitro.

Although the inhibitory levels were admittedly modest, they aligned in vitro with the behavior of the known microbiome-derived HDAC inhibitor butyrate, which exhibits weak in vitro biochemical inhibition but significant physiological epigenetic relevance in vivo. Butyrate is known to reach high local concentrations in the gut, where it modulates HDAC3-dependent transcriptional programs involved in epithelial differentiation, immune regulation, and inflammation. While concentrations of HMFA and DGL remain unknown, these metabolites could contribute to HDAC3 modulation in the intestine. Even modest inhibitory activity may have meaningful effects, particularly in combination with other synergistic metabolites. Determining whether HMFA and DGL reach local concentrations sufficient to modulate HDAC activity, individually or in combination, will require future targeted metabolomic and in vivo studies.

It is also important to note that the HDAC3-specific biochemical assay provides a simplified view of enzyme inhibition, and that biological effects in vivo may differ due to local metabolite accumulation, cellular context, and indirect regulatory mechanisms. In addition, because the unspecific HDAC assay reflects inhibition across multiple HDAC isoforms, the relative contributions of individual enzymes cannot be resolved. Future studies profiling these metabolites against additional purified HDAC isoforms and in more complex assay systems will be required to determine whether they act as weak broad-spectrum HDAC modulators or preferentially influence specific HDAC family members.

Limitations and Mechanistic Interpretation

While docking simulations suggested that HMFA may directly coordinate with the catalytic zinc center of HDAC3, these results are predictive rather than conclusive. Experimental confirmation through biophysical techniques such as surface plasmon resonance or isothermal titration calorimetry would be required to verify direct binding. At the same time, because broader HDAC activity was assessed using HeLa nuclear extracts that contain multiple HDAC isoforms, inhibitory effects observed in this context may reflect contributions from enzymes beyond HDAC3. The observed inhibitory differences between HMFA and DGL may arise from distinct interactions across the broader HDAC family rather than HDAC3 alone. Exploratory docking to HDAC2 and HDAC8 revealed similar trends in which HMFA consistently exhibited higher pose confidence than DGL. However, these docking results are hypothesis-generating and will require biochemical validation using purified HDAC isoforms to determine isoform-specific contributions.

More broadly, the weak inhibition observed reflects a fundamental challenge in identifying potent HDAC inhibitors from nonsynthetic chemical space. Classical HDACs, including HDAC3, often require corepressor complexes such as SMRT/NCoR for full catalytic activity. Nevertheless, weak modulation of HDAC activity by microbial metabolites adds to the body of knowledge of HDAC modulators as well as providing further evidence of our incomplete understanding of metabolite activities.

Another important limitation of this study is the quality of the available biochemical screening data used for model training. The HDAC3 data set was generated using a single-point, low-micromolar screening format, which likely results in systematic false negatives for weak but biologically relevant inhibitors that act at higher concentrations. In addition, the trypsin-coupled assay format may introduce false positives due to off-target effects unrelated to HDAC inhibition. These sources of label noise constrain achievable model performance and limit the interpretability of absolute activity predictions.

Future work could address these limitations by integrating multiconcentration screening data, orthogonal biochemical assays, or isoform-resolved activity measurements, which would improve label quality and enable more precise modeling of HDAC inhibition. As higher-quality data sets become available, the framework presented here can be readily adapted to incorporate improved annotations, enabling more accurate prioritization of microbiome-derived metabolites with epigenetic regulatory potential.

Future Outlook and Broader Significance

This study illustrates how tailored data curation for machine learning models can enhance their ability to predict complex natural chemotypes. Future work should expand this framework to other epigenetic and immunoregulatory targets and incorporate multiobjective learning that balances sensitivity, specificity, and toxicity prediction. Including structural or transcriptomic readouts could further strengthen biological interpretation. In addition, chemical derivatization of identified metabolites may provide a path to explore structure–activity relationships, enhance potency, and clarify isoform preferences.

Methods

Data Set Preparation

The HDAC3 bioassay data set was obtained from PubChem BioAssay (AID 2718), which contains experimental screening results for over 312,000 compounds. The activity column was used for binary labeling of classes, with compounds marked as active or inactive based on the reported assay outcome. Records with missing values, inconclusive activity, or duplicate entries were removed prior to analysis. To identify overlap with known microbiome metabolites, canonical SMILES structures from the HDAC3 data set were cross-referenced with the curated microbiome metabolite database (n = 2645). Molecules with exact structural matches were labeled as metabolites. Compounds structurally similar to known microbiome metabolites were annotated as “metabolite-like” when their maximum Tanimoto similarity, calculated against all metabolites in the microbiome metabolite database, exceeded 0.6. All data processing, filtering, and annotation steps were performed in Python (v3.12) using the pandas, RDKit, and scikit-learn libraries.

Model Training and Evaluation

Chemical features for all molecules were generated as binary Morgan fingerprints with a radius of 2 and 2048 bits using RDKit. Because these fingerprints represent molecular substructure presence or absence, no additional feature scaling or normalization was applied. The resulting fingerprints were used directly as model input features for three supervised machine learning algorithms: multilayer perceptron (MLP), random forest (RF), and XGBoost (XGB). Model hyperparameters were standardized across experiments for consistency. The MLP classifier consisted of a single hidden layer with 10 neurons, ReLU activation, the Adam solver, and a maximum of 100 training iterations. The random forest classifier used 100 estimators with bootstrap aggregation and default Gini impurity criteria. The XGBoost classifier employed 100 estimators, log loss as the evaluation metric, and suppressed label encoding for compatibility. All models were trained using a fixed random seed (42) and executed with parallel computation where applicable (n_jobs = −1). Model performance was evaluated across five random repeats using sensitivity, ROC-AUC, PR-AUC, MCC, balanced accuracy, and F1 score as performance metrics. Statistical comparisons between subsampling strategies were conducted using paired t tests to assess significance.

Subsampling and Automated Data Composition Optimization

Before model training, the data set was split into two parts: a fixed validation set and a remaining pool set used for all subsampling and training procedures. The validation set was defined by holding out all known microbiome metabolites (n = 162, all inactive), supplemented with 10% of the active nonmetabolite compounds (n = 48) and an equal number of inactive nonmetabolite compounds (n = 48) to maintain class balance. This resulted in a validation set of 258 compounds. The remaining data were designated as the training pool for model development.

To ensure robustness, the validation/pool split was repeated five times using different random seeds. In each iteration, approximately 10% of the active nonmetabolite compounds and an equal proportion of inactive nonmetabolites were randomly assigned to the validation set, while the remaining compounds were allocated to the pool set. Because the sampling was randomized, the specific compounds included in each validation split differed across repeats, ensuring broad coverage of chemical space. Each split was generated independently to ensure no data leakage between the validation and pool sets, while maintaining reproducibility within each repeat. All subsequent subsampling and training operations were performed exclusively on the pool data, and the validation set remained completely unseen during model fitting and parameter optimization.

For manual subsampling, six training set configurations were tested, each defined by distinct combinations of active, inactive, and metabolite-like molecules:

  • 1.

    Balanced metabolite-enrichedall actives (n = 437) and an equal number of inactive metabolite-like compounds (n = 437).

  • 2.

    Balanced randomall actives (n = 437) and an equal number of randomly selected inactive compounds (n = 437).

  • 3.

    Metabolite-enrichedall actives (n = 437) and all inactive metabolite-like compounds (n = 1003 ± 1).

  • 4.

    Mixedall actives (n = 437), all inactive metabolite-like compounds (n = 1003 ± 1), and an equal number of nonmetabolite-like inactives (n = 1003 ± 1).

  • 5.

    Random mixedall actives (n = 437) and a randomly selected inactive set (n = 2006) of the same total size as the mixed subset.

  • 6.

    Randomall actives (n = 437) and a larger random subset of inactives (n = 2443).

Each configuration was generated five times per data split. As a result, some subsampled sets (metabolite-enriched and mixed) exhibited minor size variations (±1 molecule). This occurred because the random selection of inactives into the validation set slightly altered the number of metabolite-like inactives remaining in the pool for subsampling.

For automated optimization, an iterative selection procedure was implemented to explore how training data composition affects model performance dynamically. Each iteration began with a warm start of 5 active and 5 inactive molecules, followed by the addition of two new compounds per round-one active and one inactive for balanced strategies, or two randomly chosen molecules for unbalanced ones:

  • 1.

    Balanced most-metabolite-likeeach round began with equal numbers of active and inactive compounds, prioritizing the inclusion of the most metabolite-like molecules based on decreasing Tanimoto similarity scores.

  • 2.

    Balanced metabolite-enrichedmaintained equal numbers of active and inactive compounds, where inactives were randomly selected from the full set of metabolite-like compounds (similarity ≥0.6).

  • 3.

    Balanced randommaintained class balance between actives and inactives, with inactive compounds added randomly at each iteration.

  • 4.

    Most-metabolite-likesequentially added compounds with the highest similarity to microbiome metabolites, without enforcing class balance.

  • 5.

    Metabolite-enrichedsimilar to the most-metabolite-like approach but included metabolite-like compounds selected randomly from the full set (similarity ≥0.6).

  • 6.

    Randomadded inactive compounds randomly in each iteration without regard to structural similarity or class balance.

The model was retrained and evaluated on the fixed validation set at each step, and this process was repeated up to 1000 rounds. The selection strategies included balanced metabolite-enriched, balanced most-metabolite-like, balanced random, and corresponding unbalanced variants. Model metrics such as sensitivity, ROC-AUC, MCC, and balanced accuracy were tracked across iterations to monitor convergence.

All subsampling and optimization workflows were implemented in Python using pandas and scikit-learn, with randomization controlled through distinct seeds across repeats to ensure both reproducibility and independence of splits.

HDAC Inhibition Assay

Experimental validation of predicted HDAC inhibitors was performed using the Histone Deacetylase (HDAC) Activity Assay Kit (Fluorometric, ab156064, Abcam), following the manufacturer’s two-step protocol. The assay measures HDAC enzymatic activity based on the deacetylation of an acetylated substrate, followed by developer-mediated fluorescence generation.

Unspecific HDAC activity was measured using HeLa nuclear extracts provided with the assay kit as the enzyme source, following the manufacturer’s instructions. Reactions were assembled in microplate format, incubated at 37 °C for 1 h, and terminated by addition of the developer reagent. HDAC3-specific activity was assessed using the same assay chemistry with recombinant HDAC3–NCoR protein (100 nM; Enzo Life Sciences) substituted as the enzyme source. Briefly, 20 μL of recombinant HDAC3–NCoR complex was combined with 28 μL of assay buffer, substrate, and water mixture, followed by addition of 2 μL of inhibitor or vehicle (DMSO).

All compounds were dissolved in DMSO and diluted in assay buffer to the desired working concentrations (final concentration of 2% DMSO). The predicted metabolites 5-(hydroxymethyl)­furoic acid (HMFA, MedChemExpress HY-W005241) and d-glucuronolactone (DGL, MedChemExpress HY-41982) were tested following the ab156064 protocol, along with 1 mM trichostatin A (TSA) and sodium butyrate as positive controls. Reactions were assembled in microplate format, incubated at 37 °C for 1 h, and terminated by addition of the developer reagent. Fluorescence intensity was measured using a microplate reader (BioTek) according to the manufacturer’s specifications (excitation/emission = 350/450 nm). Relative fluorescence units (RFUs) were normalized to vehicle-treated controls to quantify HDAC inhibition. Each condition was tested in two biological replicates with one to three technical replicates per measurement to ensure reproducibility.

Molecular Docking Simulation

Molecular docking simulations were performed using GNINA version 1.1, a convolutional neural network-enhanced docking software built upon AutoDock Vina. The docking was carried out on a system equipped with an NVIDIA Tesla T4 GPU (CUDA version 12.2, driver version 535.104.05). The HDAC3 crystal structure in complex with the SMRT deacetylase-activating domain and inositol-1,4,5,6-tetrakisphosphate (PDB ID: 4A69) was used as the receptor model. The receptor was prepared by removing water molecules and other heteroatoms not directly involved in catalysis, while retaining the catalytic zinc ion and the IP4 cofactor for structural reference.

Ligands were prepared using RDKit by generating 3D conformers, assigning protonation states, and minimizing energy before conversion to PDBQT format. Docking grids were centered on the catalytic pocket surrounding the zinc ion, with box dimensions of 25 × 25 × 25 Å to encompass the entire binding site. GNINA docking was performed with an exhaustiveness value of 8, generating 10 output poses per ligand, using a random seed value of 42, and a CNN rescoring function. Both the classical AutoDock Vina scoring function (affinity in kcal/mol) and GNINA’s convolutional neural network (CNN)-based scores were recorded.

For each ligand, the top-ranked pose was selected based on the CNN scores and visualized in py3Dmol (v2.5.3). Protein structures were rendered as cartoon or surface representations and ligands shown in stick format. Docking logs were parsed using custom Python scripts to extract affinity, intramolecular energy, and CNN pose and affinity scores into structured data tables.

Calculation of Ligand–Zinc Ion Distance

The minimum distance between the docked metabolite and the catalytic zinc ion in HDAC3 was calculated using a custom Python function developed with RDKit and NumPy. Docked protein–ligand complexes were parsed as RDKit molecular objects, and atomic coordinates were extracted from their conformers. The function iterated through all atoms in the protein structure to identify zinc atoms (atom symbol “Zn”) and through all atoms in the ligand to determine the closest interatomic distance. The Euclidean distance between each ligand atom and the zinc atom was computed, and the minimum value was reported as the zinc–ligand distance. All calculations were performed using the GNINA-generated docking conformations. Distance computation and visualization were performed in Jupyter Notebook using RDKit, NumPy, and py3Dmol to ensure consistency with the docking coordinate system.

Protein–Ligand Interaction Analysis

Protein–ligand interactions were predicted using BINANA 2.1 (https://durrantlab.pitt.edu/binana/), which automatically identifies potential noncovalent contacts between receptor and ligand atoms based on geometric criteria. The docked HDAC3–ligand complexes generated by GNINA were converted to PDBQT format and analyzed with BINANA using default parameters. The software classified interaction types including hydrophobic contacts, π–π stacking, cation–π interactions, hydrogen bonds, salt bridges, and metal coordination events.

Statistical Analysis

All statistical analyses were conducted in Python using NumPy, SciPy, and scikit-learn libraries. Model performance metrics, including sensitivity, specificity, ROC-AUC, PR-AUC, MCC, balanced accuracy, and F1 score, were calculated from five independent repeats for each subsampling configuration. Comparisons between models were performed using paired t-tests with a significance threshold of p < 0.05.

For experimental assays, all measurements were collected from two biological replicates with one to three technical replicates per condition. Data were analyzed in GraphPad Prism (version 10.0) using unpaired t-test against the vehicle control. Statistical significance was defined as *p < 0.05, **p < 0.01, ***p < 0.001, and ****p < 0.0001.

Supplementary Material

ao5c12875_si_001.pdf (570.3KB, pdf)

Acknowledgments

This study was supported, in part, by a Flash Grant from NC Biotech (2021-FLG-3819) to D.R., a UNC CGIBD Pilot Award (NIH NIDDK DK034987) to D.R., a Duke Cancer Institute and Duke Microbiome Center Pilot Award (NIH NCI CA014236) to D.R., the Engineering Research Center for Precision Microbiome Engineering (NSF EEC-2133504) to D.R., DK137771 to T.A., DK116868 to T.A., and K01DK135647 to E.M.E.

The Supporting Information is available free of charge at https://pubs.acs.org/doi/10.1021/acsomega.5c12875.

  • Molecular docking scores for candidate metabolites; detailed model performance comparisons for manually subsampled training sets; and comprehensive automated data composition optimization results across evaluation metrics and optimization rounds (PDF)

H.A.C. and E.M.E. Contributed equally. H.A.C. and D.R. conceptualized the study. H.A.C. developed and implemented the machine learning framework. E.M.E. performed the in vitro experiments, and J.C. performed the molecular docking simulations. T.A. and D.R. provided supervision throughout the project. H.A.C. wrote the initial manuscript draft, and D.R., E.M.E., and T.A. contributed to review and editing. All authors read and approved the final version of the manuscript.

The authors declare the following competing financial interest(s): D.R. acts as a consultant to the pharmaceutical and biotechnology industry and as a mentor for Start2.

References

  1. Park S., Kim J.. A short guide to histone deacetylases including recent progress on class II enzymes. Exp. Mol. Med. 2020;52:204–212. doi: 10.1038/s12276-020-0382-4. [DOI] [PMC free article] [PubMed] [Google Scholar]
  2. Curcio A., Rocca R., Alcaro S., Artese A.. The histone deacetylase family: Structural features and application of combined computational methods. Pharmaceuticals. 2024;17(5):620. doi: 10.3390/ph17050620. [DOI] [PMC free article] [PubMed] [Google Scholar]
  3. Seto E., Yoshida M.. Erasers of histone acetylation: The histone deacetylase enzymes. Cold Spring Harbor Perspect. Biol. 2014;6(4):a018713. doi: 10.1101/cshperspect.a018713. [DOI] [PMC free article] [PubMed] [Google Scholar]
  4. Yoon J. G., Lim S.-K., Seo H., Lee S., Cho J., Kim S. Y., Koh H. Y., Poduri A. H., Ramakumaran V., Vasudevan P., de Groot M. J., Ko J. M., Han D., Chae J.-H., Lee C.-H.. De novo missense variants in HDAC3 leading to epigenetic machinery dysfunction are associated with a variable neurodevelopmental disorder. Am. J. Hum. Genet. 2024;111(8):1588–1604. doi: 10.1016/j.ajhg.2024.06.015. [DOI] [PMC free article] [PubMed] [Google Scholar]
  5. He R., He Z., Zhang T., Liu B., Gao M., Li N., Geng Q.. HDAC3 in action: Expanding roles in inflammation and inflammatory diseases. Cell Proliferation. 2025;58(1):e13731. doi: 10.1111/cpr.13731. [DOI] [PMC free article] [PubMed] [Google Scholar]
  6. Li Y., Izhar T., Kanekiyo T.. HDAC3 as an Emerging Therapeutic Target for Alzheimer’s Disease and other Neurological Disorders. Mol. Neurobiol. 2025;62(8):9573–9585. doi: 10.1007/s12035-025-04866-w. [DOI] [PMC free article] [PubMed] [Google Scholar]
  7. Watson N., Kuppuswamy S., Ledford W. L., Sukumari-Ramesh S.. The role of HDAC3 in inflammation: mechanisms and therapeutic implications. Front. Immunol. 2024;15:1419685. doi: 10.3389/fimmu.2024.1419685. [DOI] [PMC free article] [PubMed] [Google Scholar]
  8. Kim H.-J., Bae S.-C.. Histone deacetylase inhibitors: molecular mechanisms of action and clinical trials as anti-cancer drugs. Am. J. Transl. Res. 2010;3(2):166–179. [PMC free article] [PubMed] [Google Scholar]
  9. Bondarev A. D., Attwood M. M., Jonsson J., Chubarev V. N., Tarasov V. V., Schiöth H. B.. Recent developments of HDAC inhibitors: Emerging indications and novel molecules. Br. J. Clin. Pharmacol. 2021;87(12):4577–4597. doi: 10.1111/bcp.14889. [DOI] [PubMed] [Google Scholar]
  10. Subramanian S., Bates S. E., Wright J. J., Espinoza-Delgado I., Piekarz R. L.. Clinical toxicities of histone deacetylase inhibitors. Pharmaceuticals. 2010;3(9):2751–2767. doi: 10.3390/ph3092751. [DOI] [PMC free article] [PubMed] [Google Scholar]
  11. Chung H. A., Fralish Z., Tu T., Reker D.. Profiling Biological Effects of Microbiome Metabolites via Machine Learning. ChemRxiv. 2025:15mw9. doi: 10.26434/chemrxiv-2025-15mw9. [DOI] [PMC free article] [PubMed] [Google Scholar]
  12. Spivak I., Fluhr L., Elinav E.. Local and systemic effects of microbiome-derived metabolites. EMBO Rep. 2022;23:EMBR202255664. doi: 10.15252/embr.202255664. [DOI] [PMC free article] [PubMed] [Google Scholar]
  13. Yamada T., Palm N. W.. A host-centric view of the microbiota metabolome. Immunity. 2025;58(12):2940–2956. doi: 10.1016/j.immuni.2025.11.006. [DOI] [PubMed] [Google Scholar]
  14. Miro-Blanch J., Yanes O.. Epigenetic regulation at the interplay between gut microbiota and host metabolism. Front. Genet. 2019;10:638. doi: 10.3389/fgene.2019.00638. [DOI] [PMC free article] [PubMed] [Google Scholar]
  15. Rubas N. C., Torres A., Maunakea A. K.. The gut microbiome and epigenomic reprogramming: mechanisms, interactions, and implications for human health and disease. Int. J. Mol. Sci. 2025;26(17):8658. doi: 10.3390/ijms26178658. [DOI] [PMC free article] [PubMed] [Google Scholar]
  16. Li C., Liang Y., Qiao Y.. Messengers from the gut: Gut microbiota-derived metabolites on host regulation. Front. Microbiol. 2022;13:863407. doi: 10.3389/fmicb.2022.863407. [DOI] [PMC free article] [PubMed] [Google Scholar]
  17. Krautkramer K. A., Rey F. E., Denu J. M.. Chemical signaling between gut microbiota and host chromatin: What is your gut really saying? J. Biol. Chem. 2017;292(21):8582–8593. doi: 10.1074/jbc.R116.761577. [DOI] [PMC free article] [PubMed] [Google Scholar]
  18. Waldecker M., Kautenburger T., Daumann H., Busch C., Schrenk D.. Inhibition of histone-deacetylase activity by short-chain fatty acids and some polyphenol metabolites formed in the colon. J. Nutr. Biochem. 2008;19(9):587–593. doi: 10.1016/j.jnutbio.2007.08.002. [DOI] [PubMed] [Google Scholar]
  19. Vinolo M. A. R., Rodrigues H. G., Nachbar R. T., Curi R.. Regulation of inflammation by short-chain fatty acids. Nutrients. 2011;3(10):858–876. doi: 10.3390/nu3100858. [DOI] [PMC free article] [PubMed] [Google Scholar]
  20. Bishop K. S., Xu H., Marlow G.. Epigenetic regulation of gene expression induced by butyrate in colorectal cancer: involvement of microRNA. Genet. Epigenet. 2017;9:1179237X17729900. doi: 10.1177/1179237x17729900. [DOI] [PMC free article] [PubMed] [Google Scholar]
  21. Nshanian M., Gruber J. J., Geller B. S., Chleilat F., Lancaster S. M., White S. M., Alexandrova L., Camarillo J. M., Kelleher N. L., Zhao Y., Snyder M. P.. Short-chain fatty acid metabolites propionate and butyrate are unique epigenetic regulatory elements linking diet, metabolism and gene expression. Nat. Metab. 2025;7:196–211. doi: 10.1038/s42255-024-01191-9. [DOI] [PMC free article] [PubMed] [Google Scholar]
  22. Wu S.-e., Hashimoto-Hill S., Woo V., Eshleman E. M., Whitt J., Engleman L., Karns R., Denson L. A., Haslam D. B., Alenghat T.. Microbiota-derived metabolite promotes HDAC3 activity in the gut. Nature. 2020;586(7827):108–112. doi: 10.1038/s41586-020-2604-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
  23. Mullowney M. W., Duncan K. R., Elsayed S. S., Garg N., van der Hooft J. J. J., Martin N. I., Meijer D., Terlouw B. R., Biermann F., Blin K., Durairaj J., Gorostiola González M., Helfrich E. J. N., Huber F., Leopold-Messer S., Rajan K., de Rond T., van Santen J. A., Sorokina M., Balunas M. J., Beniddir M. A., van Bergeijk D. A., Carroll L. M., Clark C. M., Clevert D.-A., Dejong C. A., Du C., Ferrinho S., Grisoni F., Hofstetter A., Jespers W., Kalinina O. V., Kautsar S. A., Kim H., Leao T. F., Masschelein J., Rees E. R., Reher R., Reker D., Schwaller P., Segler M., Skinnider M. A., Walker A. S., Willighagen E. L., Zdrazil B., Ziemert N., Goss R. J. M., Guyomard P., Volkamer A., Gerwick W. H., Kim H. U., Müller R., van Wezel G. P., van Westen G. J. P., Hirsch A. K. H., Linington R. G., Robinson S. L., Medema M. H.. Artificial intelligence for natural product drug discovery. Nat. Rev. Drug Discovery. 2023;22:895–916. doi: 10.1038/s41573-023-00774-7. [DOI] [PMC free article] [PubMed] [Google Scholar]
  24. National Center for Biotechnology Information . PubChem Bioassay Record for AID 2718, Fluorescence Cell-Free Homogeneous Primary HTS to Identify Inhibitors of Histone Deacetylase 3. Source: Broad Institute 2025.
  25. Reker, D. Cheminformatic analysis of natural product fragments. Progress in the Chemistry of Organic Natural Products; Springer International Publishing, 2019, Vol. 110, pp 143–175. 10.1007/978-3-030-14632-0_5. [DOI] [PubMed] [Google Scholar]
  26. Reker D., Perna A. M., Rodrigues T., Schneider P., Reutlinger M., Mönch B., Koeberle A., Lamers C., Gabler M., Steinmetz H., Müller R., Schubert-Zsilavecz M., Werz O., Schneider G.. Revealing the macromolecular targets of complex natural products. Nat. Chem. 2014;6:1072–1078. doi: 10.1038/nchem.2095. [DOI] [PubMed] [Google Scholar]
  27. Kim M., Hwang K.-B.. An empirical evaluation of sampling methods for the classification of imbalanced data. PLoS One. 2022;17(7):e0271260. doi: 10.1371/journal.pone.0271260. [DOI] [PMC free article] [PubMed] [Google Scholar]
  28. Wilson D. L.. Asymptotic properties of nearest neighbor rules using edited data. IEEE Trans. Syst. Man Cybern. 1972;SMC-2(3):408–421. doi: 10.1109/TSMC.1972.4309137. [DOI] [Google Scholar]
  29. Chawla N. V., Bowyer K. W., Hall L. O., Kegelmeyer W. P.. SMOTE: Synthetic Minority Over-sampling Technique. J. Artif. Intell. Res. 2002;16:321–357. doi: 10.1613/jair.953. [DOI] [Google Scholar]
  30. Dubey R., Zhou J., Wang Y., Thompson P. M., Ye J.. For the Alzheimer’s Disease Neuroimaging Initiative. Analysis of sampling techniques for imbalanced data: An n = 648 ADNI study. NeuroImage. 2014;87:220–241. doi: 10.1016/j.neuroimage.2013.10.005. [DOI] [PMC free article] [PubMed] [Google Scholar]
  31. Wongvorachan T., He S., Bulut O.. A comparison of undersampling, oversampling, and SMOTE methods for dealing with imbalanced classification in educational data mining. Information. 2023;14(1):54. doi: 10.3390/info14010054. [DOI] [Google Scholar]
  32. Maggiora G., Vogt M., Stumpfe D., Bajorath J.. Molecular similarity in medicinal chemistry. J. Med. Chem. 2014;57(8):3186–3204. doi: 10.1021/jm401411z. [DOI] [PubMed] [Google Scholar]
  33. Bajusz D., Rácz A., Héberger K.. Why is Tanimoto index an appropriate choice for fingerprint-based similarity calculations? J. Cheminf. 2015;7:20. doi: 10.1186/s13321-015-0069-3. [DOI] [PMC free article] [PubMed] [Google Scholar]
  34. Wen Y. Z.. et al. Improving molecular machine learning through adaptive subsampling with active learning. Digital Discovery. 2023;2(4):1134–1142. doi: 10.1039/D3DD00037K. [DOI] [Google Scholar]
  35. Becerra M. L., Lizarazo L. M., Rojas H. A., Prieto G. A., Martinez J. J.. Biotransformation of 5-hydroxymethylfurfural and furfural with bacteria of Bacillus genus. Biocatal. Agric. Biotechnol. 2022;39:102281. doi: 10.1016/j.bcab.2022.102281. [DOI] [Google Scholar]
  36. Alper Oral R., Dogan M., Sarioglu K., Kayacier A., Sagdic O.. Determination of HMF in some instant foods and its biodegradation by lactic acid bacteria in medium and food. Ann. Chromatogr. Sep. Tech. 2015;1(1):1–4. doi: 10.36876/acst.1004. [DOI] [Google Scholar]
  37. Hong H. A., Khaneja R., Tam N. M. K., Cazzato A., Tan S., Urdaci M., Brisson A., Gasbarrini A., Barnes I., Cutting S. M.. Bacillus subtilis isolated from the human gastrointestinal tract. Res. Microbiol. 2009;160(2):134–143. doi: 10.1016/j.resmic.2008.11.002. [DOI] [PubMed] [Google Scholar]
  38. Husøy T., Haugen M., Murkovic M., Jöbstl D., Stølen L. H., Bjellaas T., Rønningborg C., Glatt H., Alexander J.. Dietary exposure to 5-hydroxymethylfurfural from Norwegian food and correlations with urine metabolites of short-term exposure. Food Chem. Toxicol. 2008;46(12):3697–3702. doi: 10.1016/j.fct.2008.09.048. [DOI] [PubMed] [Google Scholar]
  39. Kimura Y., Tani S., Hayashi A., Ohtani K., Fujioka S., Kawano T., Shimada A.. Nematicidal activity of 5-hydroxymethyl-2-furoic acid against plant-parasitic nematodes. Z. Naturforsch., C: J. Biosci. 2007;62(3–4):234–238. doi: 10.1515/znc-2007-3-413. [DOI] [PubMed] [Google Scholar]
  40. Zhang B.. et al. Characteristics of serum metabolites and gut microbiota in diabetic kidney disease. Front. Pharmacol. 2022;13:872988. doi: 10.3389/fphar.2022.872988. [DOI] [PMC free article] [PubMed] [Google Scholar]
  41. Gao S., Sun R., Singh R., Yu So S., Chan C. T. Y., Savidge T., Hu M.. The role of gut microbial β-glucuronidase in drug disposition and development. Drug Discovery Today. 2022;27(10):103316. doi: 10.1016/j.drudis.2022.07.001. [DOI] [PMC free article] [PubMed] [Google Scholar]
  42. Hu S., Ding Q., Zhang W., Kang M., Ma J., Zhao L.. Gut microbial β-glucuronidase: A vital regulator in female estrogen metabolism. Gut Microbes. 2023;15(1):2236749. doi: 10.1080/19490976.2023.2236749. [DOI] [PMC free article] [PubMed] [Google Scholar]
  43. Salvi P. S., Cowles R. A.. Butyrate and the intestinal epithelium: Modulation of proliferation and inflammation in homeostasis and disease. Cells. 2021;10(7):1775. doi: 10.3390/cells10071775. [DOI] [PMC free article] [PubMed] [Google Scholar]
  44. Davie J. R.. Inhibition of histone deacetylase activity by butyrate. J. Nutr. 2003;133(7):2485S–2493S. doi: 10.1093/jn/133.7.2485S. [DOI] [PubMed] [Google Scholar]
  45. Parada-Venegas D., De la Fuente López M., Dubois-Camacho K., Landskron G., Blokzijl T., Molina H., Casanova M.-C., Cui Y., Liu M., Da Costa De Pina A. M., Simian D., González M.-J., Weersma R. K., Quera R., Dijkstra G., Faber K. N., Hermoso M. A.. Butyrate suppresses mucosal inflammation in inflammatory bowel disease primarily through HDAC3 inhibition in monocytes and macrophages. FEBS J. 2025;292:6134–6157. doi: 10.1111/febs.70289. [DOI] [PMC free article] [PubMed] [Google Scholar]
  46. Schulthess J., Pandey S., Capitani M., Rue-Albrecht K. C., Arnold I., Franchini F., Chomka A., Ilott N. E., Johnston D. G. W., Pires E., McCullagh J., Sansom S. N., Arancibia-Cárcamo C. V., Uhlig H. H., Powrie F.. The short-chain fatty acid butyrate imprints an antimicrobial program in macrophages. Immunity. 2019;50(2):432–445. doi: 10.1016/j.immuni.2018.12.018. [DOI] [PMC free article] [PubMed] [Google Scholar]
  47. McNutt A. T., Francoeur P., Aggarwal R., Masuda T., Meli R., Ragoza M., Sunseri J., Koes D. R.. GNINA 1.0: molecular docking with deep learning. J. Cheminf. 2021;13:43. doi: 10.1186/s13321-021-00522-2. [DOI] [PMC free article] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

ao5c12875_si_001.pdf (570.3KB, pdf)

Articles from ACS Omega are provided here courtesy of American Chemical Society

RESOURCES