Abstract
Inflammatory bowel disease (IBD) is characterized by a refractory, relapsing inflammatory state driven by a multifaceted and poorly understood interplay between host mucosal metabolic disturbances and immune microenvironment dysregulation. To uncover robust non-invasive diagnostic indicators, this study implemented an integrated analytical framework combining advanced machine learning feature selection, penalized logistic regression modeling, and leave-one-dataset-out cross-validation to construct a novel diagnostic signature derived from metabolic cell death-related genes (MCDRGs). Systematic multiple-testing correction and rigorous data harmonization were applied across all independent discovery and validation cohorts to control for potential batch effects. Through this approach, we successfully identified a core three-gene candidate panel comprising indoleamine 2,3-dioxygenase 1 (IDO1), lipocalin 2 (LCN2), and solute carrier family 6 member 14 (SLC6A14). Methodologically, we demonstrated that the initial near-perfect apparent discrimination within the discovery cohort was mathematically attributable to quasi-complete data separation rather than systemic model overfitting. This identified signature exhibited consistent cross-cohort validation and correlated tightly with the coordinated infiltration and functional states of multiple mucosal immune cell subsets. Furthermore, independent clinical assays substantiated the synchronized elevation of these markers during active clinical phases, while orthogonal single-cell RNA-sequencing confirmed their prominent, cell-type-specific enrichment within the myeloid, epithelial, and stromal compartments of inflamed intestinal mucosa. Collectively, the identified MCDRGs, IDO1, LCN2, and SLC6A14, link metabolic dysregulation, immune infiltration, and regulated cell death, offering insights into IBD pathophysiology and providing a transcriptomic candidate signature with exploratory diagnostic potential for IBD stratification.
Supplementary Information
The online version contains supplementary material available at https://doi.org/10.1038/s41598-026-57908-1.
Keywords: Inflammatory Bowel Disease, Metabolic Cell Death, Machine Learning, Diagnostic Biomarkers, Immune Infiltration, Single-Cell RNA Sequencing
Subject terms: Biomarkers, Computational biology and bioinformatics, Diseases, Gastroenterology, Immunology
Introduction
Inflammatory bowel disease (IBD), which includes Crohn’s disease (CD) and ulcerative colitis (UC), is a chronic, relapsing inflammatory disorder of the gastrointestinal tract that has emerged as a global health burden1. The etiology of IBD is multifactorial and involves a complex interplay between genetic susceptibility, environmental factors, gut microbiota dysbiosis, and aberrant immune responses2. Persistent mucosal inflammation in patients with IBD leads to irreversible intestinal damage, which significantly increases the risk of colorectal cancer3. Despite advances in biological therapies, a substantial proportion of patients fail to achieve long-term remission, primarily because of the heterogeneous nature of the disease and lack of precise early diagnostic biomarkers4,5. Therefore, identifying novel molecular signatures that reflect the underlying pathophysiological mechanisms of IBD is imperative to improve diagnostic accuracy and personalize therapeutic strategies.
Beyond conventional regulated cell death, the emerging paradigm of metabolic cell death, encompassing ferroptosis, cuproptosis, and the recently identified disulfidptosis, has been recognized as a new driver of several diseases6–9. Notably, in IBD, these forms of cell death play a particularly significant role and are characterized by an intimate interplay between metabolic dysregulation and cell death signaling pathways. For instance, the inactivation of glutathione peroxidase 4 (GPX4) drives ferroptosis in IBD by impairing lipid peroxide detoxification, leading to iron-dependent oxidative damage, intestinal epithelial cell death, and exacerbation of mucosal inflammation10. Another study found that copper depletion ameliorates intestinal barrier damage in mice with dextran sulfate sodium-induced experimental colitis by inhibiting cuproptosis11. Recent evidence indicates that metabolic cell death-related genes (MCDRGs) do not function in isolation but are intricately linked to the remodeling of the intestinal immune microenvironment12,13. Although bioinformatic landscapes have identified individual MCDRGs as potential players in mucosal inflammation, their integrated diagnostic utility and synergistic impact on disease progression remain largely uncharacterized. Consequently, there is a need to decode these complex metabolic signatures to develop multifaceted candidate biomarkers for diagnosis and therapeutic monitoring of IBD.
In this era of precision medicine, the integration of high-throughput multi-omics data with advanced computational algorithms has revolutionized biomarker discovery14,15. Traditional statistical methods often struggle with high dimensionality and the noise inherent in genomic data. In contrast, machine learning (ML) frameworks, such as the least absolute shrinkage and selection operator (LASSO), Boruta random forest, and support vector machine (SVM)-recursive feature elimination (SVM-RFE), offer robust capabilities for feature selection and predictive modeling16,17. These tools allow the identification of candidate hub genes that are not only statistically significant but also biologically relevant to disease progression, providing a foundation for translating complex bioinformatic outputs into potential clinical decision-making tools.
In this study, we integrated multiple transcriptomic datasets and used advanced ML algorithms, combined with penalized regression modeling and LODO-CV, to identify a candidate MCDRG-related signature including IDO1, LCN2, and SLC6A14 for diagnostic stratification of patients with IBD. By leveraging bioinformatic frameworks, we characterized the immunological landscape associated with this signature and revealed correlations between these genes and the functional states of several immune cells. Systematic BH-FDR correction was applied throughout to ensure statistical transparency. Furthermore, the clinical relevance of the identified biomarkers was validated in an independent clinical cohort, and single-cell RNA-sequencing analysis was additionally performed to validate cell-type-specific expression of the three genes in IBD-inflamed intestinal tissue.
Materials and methods
Public data acquisition
IBD expression profile data were retrieved from the Gene Expression Omnibus (GEO) database via the “GEOquery” R package version 2.72.0. The discovery set comprised the dataset GSE75214 from the GPL6244 platform, which included 133 IBD samples and 22 non-IBD controls. For external validation, the dataset GSE87466 from the GPL13158 platform included 87 IBD samples and 21 controls, and the dataset GSE47908 from the GPL570 platform included 45 IBD samples and 15 controls. Metabolic cell death encompasses multiple regulated cell death modalities, including ferroptosis, cuproptosis, disulfidptosis, lysosomal zinc-induced cell death, and alkaliptosis, as described by Mao et al.6. Ferroptosis-related genes were identified following the strategy described by Liu et al.18 and downloaded from the FerrDb database, yielding 415 genes. Data collected from published studies included 36 cuproptosis-related genes, 23 disulfidptosis-related genes, 1 lysosomal zinc death-related gene, and 7 alkaliptosis-related genes19–21. After merging all gene sets and removing duplicates, 478 unique MCDRGs were retained for subsequent analysis.
Differential gene expression analysis
The “limma” R package was used to perform differential gene expression analysis to identify differentially expressed genes (DEGs) within IBD and control samples. The p-values were adjusted using the Benjamini-Hochberg (BH) method. The screening criteria included a Benjamini-Hochberg false discovery rate (BH-FDR) < 0.05 and |log2FC| > 1. The “ggplot2” R package was used to draw volcano plots for visual display of the results, labeling the names of the top 20 up- and down-regulated genes.
Weighted gene co-expression network analysis
Weighted Gene Co-expression Network Analysis (WGCNA) was performed on IBD to explore gene expression pattern networks. The “goodSamplesGenes” function of the R package WGCNA was used to cluster all samples in the discovery set, and to check and remove any outliers. An adjacency matrix was constructed to calculate the Pearson correlation coefficient between all gene pairs in the selected samples, with beta = 14 (scale-free R² = 0.85) as the soft threshold to ensure a scale-free network. Functional modules in WGCNA were identified by calculating the topological overlap measure using the adjacency matrix. Gene modules and screening module eigengenes were established using a dynamic tree-building method. To obtain highly correlated gene modules, the correlation coefficient matrix between module eigengenes and traits was calculated, and modules with an absolute correlation > 0.3 and p-value < 0.01 were selected as key module genes. BH-FDR correction was applied to the module-trait correlations.
Identification and functional enrichment analysis of candidate MCDRGs
The “Venn” R package was used to draw a Venn diagram showing the overlap between DEGs, key WGCNA module genes, and MCDRGs. The “clusterProfiler” R package was used to perform gene ontology (GO) and Kyoto Encyclopedia of Genes and Genomes (KEGG) functional enrichment analyses on the candidate MCDRGs with a screening threshold of BH-FDR < 0.05. The top 10 enrichments for each ontology of significant GO and the top 20 pathways of KEGG were shown in bubble charts.
ML screening for key MCDRGs
Three ML algorithms, LASSO, Boruta, and SVM-RFE, were independently applied to the discovery set, in which the expression matrix of candidate MCDRGs was defined as the predictor variable, and disease status as the response variable. All data preprocessing and parameter tuning were strictly integrated within the cross-validation framework. For LASSO analysis, gene expression data were first standardized, and a logistic regression model with L1 regularization was constructed using the “glmnet” R package. The optimal penalty parameter (lambda) was determined via 10-fold cross-validation based on the minimum mean cross-validated error. For Boruta analysis, the “Boruta” R package was used to perform feature selection based on a random forest classifier. Shadow features were generated by permuting original gene expression values. Genes exhibiting significantly higher importance than the maximum shadow importance were classified as confirmed. For SVM-RFE analysis, a linear kernel SVM was implemented using the “e1071” R package and the optimal feature subset was determined by minimizing the 10-fold cross-validation error through recursive feature elimination. Finally, genes identified by all three algorithms were defined as key MCDRGs for subsequent validation and downstream analysis.
Validation of key MCDRGs
The independent GSE87466 and GSE47908 datasets were used as external validation sets to examine inter-group expression patterns of the identified hub genes. Statistical significance of expression differences was assessed using an unpaired t-test, with BH-FDR correction applied within each validation cohort. The diagnostic performance of each gene was quantified via ROC curve analysis using the “pROC” R package. Genes were prioritized for inclusion in the final diagnostic model based on a dual-threshold strategy requiring both a significant inter-group difference (BH-FDR < 0.05) across all datasets and an area under the curve (AUC) exceeding 0.85.
Diagnostic model construction and evaluation
To address concerns about the high apparent AUC observed in the discovery cohort, we applied penalized regression to construct a robust diagnostic model. Firth penalized logistic regression was used as the primary modeling method, with five additional methods (unpenalized logistic regression, ridge, lasso, elastic net, and Bayesian logistic regression) fitted in parallel for comparative evaluation. Apparent AUC, bootstrap-optimism-corrected AUC (1,000 resamples), repeated 10-fold cross-validation AUC, and LOOCV AUC were computed for each method.
To provide a rigorous unbiased estimate of generalization performance, leave-one-dataset-out cross-validation (LODO-CV) was performed: in each of three iterations, the model was trained on two cohorts and the held-out cohort served as the test set, yielding three out-of-sample AUC estimates. A permutation test (n = 1,000 permutations, random label shuffling) was conducted to confirm non-random discriminatory ability. Calibration of the LODO-CV held-out predictions was assessed by binned calibration analysis (15 equal-count bins with Wilson 95% CI), Brier score, mean absolute calibration error (Eavg), and calibration slope. A nomogram was constructed using the Firth penalized logistic regression model to facilitate clinical score estimation. The “rms” package was used for nomogram visualization.
Spearman’s rank correlation analysis
Based on all samples within the discovery set and the in-house clinical collection cohort, Spearman’s rank correlation analysis was performed between key MCDRGs using the “stats” R package. BH-FDR correction was applied to all pairwise correlations within each analysis family. Results were displayed as correlation heatmaps.
Gene set enrichment analysis
The “msigdbr” R package was used to obtain KEGG pathway-related gene sets. All genes were ranked in descending order according to their correlation coefficients with each key gene to generate ranked gene lists for downstream analyses. Gene set enrichment analysis (GSEA) was conducted using the “clusterProfiler” package to identify significantly associated pathways, with significance thresholds set at a normalized enrichment score > 1 and BH-FDR < 0.05. The “enrichplot” package was used to visualize the top five most significantly enriched positively and negatively associated signaling pathways.
Clinical validation in human peripheral blood
Peripheral blood samples were obtained from patients with active IBD (n = 16) and age-matched healthy controls (n = 30) following approval by the Ethics Committee of the Affiliated Hospital of North Sichuan Medical College (Approval No. 2024ER187-2), and informed consent was obtained from all participants.
For leukocyte RNA extraction, whole blood collected in EDTA tubes was treated with red blood cell lysis buffer, washed with PBS, pelleted by centrifugation, and processed for RNA isolation. Total RNA was extracted and reverse transcribed using TRIzol (Thermo Fisher Scientific) and PrimeScript RT reagent kits (Takara). The mRNA expression was quantified by qRT-PCR using SYBR Premix Ex Taq II kit (RR820A, Takara Inc.). Relative expression levels were normalized to ACTB and calculated using the 2−ΔΔCt method. Primer sequences: IDO1 forward, GATGTCCGTAAGGTCTTGCCA; reverse, TGCAGTCTCCATCACGAAATG; LCN2 forward, GTGAGCACCAACTACAACCAGC; reverse, GTTCCGAAGTCAGCTCCTTGGT; SLC6A14 forward, TGCACCTGCTACCAGTCAAG; reverse, GTCCATGGTTCACTCCCTCG; ACTB forward, TGACGTGGACATCCGCAAAG; reverse, CTGGAAGGTGGACAGCGAGG.
Immune cell infiltration analysis
The CIBERSORT algorithm implemented in the “IOBR” R package was applied to estimate the relative proportions of 22 immune cell types in each sample. Only samples with CIBERSORT output p-value < 0.05 were retained for subsequent analysis. Differences in immune cell infiltration between IBD and control samples were assessed using the Wilcoxon rank-sum test, with BH-FDR correction applied within the immune infiltration analysis family. Correlation heatmaps were constructed to evaluate relationships among immune cell subsets and key MCDRGs, with Spearman’s rank correlations BH-FDR corrected. Single-sample GSEA (ssGSEA) was performed to quantify the relative enrichment of 26 predefined immune cell-related gene signatures. Subsequent Spearman correlations between ssGSEA scores and key gene expression levels were BH-FDR corrected within the ssGSEA analysis family.
Single-cell RNA sequencing analysis
To provide orthogonal cell-type-level validation, publicly available single-cell RNA-sequencing data from IBD intestinal tissue (GSE214695) were analyzed. The dataset comprised 18 samples from CD (n = 6), UC (n = 6), and healthy control (HC; n = 6) donors. Raw 10X Genomics matrices were processed using the “Seurat” R package (v5). Quality control thresholds were set at 200-5,000 features per cell and mitochondrial fraction < 20%. After filtering, 44,941 high-quality cells were retained. Data were normalized using the “NormalizeData” function, variable features were identified with “FindVariableFeatures” (nfeatures = 2,000), and data were scaled. Dimensionality reduction was performed using PCA (30 principal components), and batch effects across samples were corrected using Harmony. Clustering was performed using FindNeighbors and FindClusters (resolution = 1.43), yielding 35 clusters. Cell type annotation was performed using a three-source strategy integrating (1) per-cluster SingleR annotation with the HumanPrimaryCellAtlasData reference, (2) marker score analysis using tissue-specific marker gene sets, and (3) top marker gene inspection. Final annotations were assigned to seven major cell types: T/NK cells, B cells, Plasma cells, Myeloid cells, Epithelial cells, Fibroblasts, and Endothelial cells. Mean normalized expression of IDO1, LCN2, and SLC6A14 was calculated across the seven annotated major cell types and disease groups (Supplementary Table 22).
Statistical analyses and multiple-testing correction
Statistical analyses and visualizations were performed using R software (v4.4.1). Continuous data are presented as mean ± standard deviation (SD) or standard error of the mean (SEM). Intergroup differences were evaluated using the unpaired t-test or Wilcoxon rank-sum test based on the normality of the data distribution. To control for multiple testing, BH-FDR correction was applied within predefined analysis families, including: WGCNA module-trait correlations, differential expression across three cohorts (9 comparisons), pairwise Spearman correlations among the three key genes, clinical qRT-PCR validation comparisons, CIBERSORT group comparisons (20 cell types), CIBERSORT gene-immune correlations, ssGSEA group comparisons (26 signatures), and ssGSEA gene-immune correlations. A permutation test (n = 1,000) was performed to confirm non-random discrimination of the diagnostic model. All statistical tests were two-tailed, and BH-FDR < 0.05 was considered statistically significant unless otherwise specified. Complete software versions, parameter settings, and threshold rationale are provided in Supplementary Tables 23–25.
Results
Screening of metabolic cell death-related DEGs
Accumulating evidence suggests that the interplay between aberrant immune responses and metabolic dysregulation drives the progression of mucosal inflammation in IBD22,23. Using the GSE75214 cohort as the discovery set, we identified a robust pool of DEGs characterizing the transcriptomic shift from healthy mucosa to IBD lesions. Principal component analysis (PCA) demonstrated distinct transcriptomic clustering between IBD and control samples, reflecting high inter-group heterogeneity (Fig. 1A). Preprocessing quality control for all three transcriptomic cohorts confirmed adequate normalization and the absence of systematic outliers (Figure S1; Supplementary Tables 1–3). Differential gene expression analysis revealed 751 DEGs in IBD samples compared with controls, consisting of 475 up-regulated and 276 down-regulated genes, with the top 10 most significantly altered genes in each category highlighted to pinpoint the core transcriptomic shifts (Fig. 1B-C).
Fig. 1.

Identification of DEGs and transcriptomic landscape of IBD in the discovery set. (A) PCA plot demonstrating the distinct transcriptomic clustering and high inter-group heterogeneity between IBD (orange dots) and control (blue dots) samples in the discovery set (GSE75214). (B-C) Visualization of DEGs via volcano plot (B) and hierarchical clustering heatmap with dendrogram (Ward’s linkage, Euclidean distance) (C), with the top 10 significantly up-regulated and down-regulated genes highlighted.
Identification and functional enrichment analysis of candidate MCDRGs
To mine the module genes related to IBD phenotypes, WGCNA was performed on the discovery set. Prior to network construction, sample clustering analysis demonstrated good overall consistency among the samples with no obvious outliers detected (Figure S2A). Scale-free topology analysis indicated that the selected soft-thresholding power (beta = 14) was sufficient to approximate a scale-free network structure, as evidenced by high scale independence and acceptable mean connectivity (Figure S2B-C). Hierarchical clustering combined with dynamic tree cutting identified multiple distinct gene modules with coherent expression patterns (Figure S2D), whereas network topology analysis confirmed robust intramodular connectivity within each module (Figure S2E). As shown in Fig. 2A, seven modules exhibited significant correlations with IBD after BH-FDR correction (Supplementary Tables 4–5) and the most strongly correlated modules were selected as disease-relevant modules. A total of 4,377 key module genes within these IBD-associated modules were extracted for downstream analysis.
Fig. 2.

Identification and functional enrichment of candidate MCDRGs in IBD. (A) Heatmap of module-trait relationships illustrating the correlation coefficients between WGCNA co-expression modules and clinical phenotypes, with BH-FDR-adjusted significance indicated by asterisks in the discovery set. (B) Venn diagram displaying the intersection of 751 DEGs, 4,377 key module genes and 478 MCDRGs, yielding 25 candidate MCDRGs. (C) GO enrichment analysis of candidate MCDRGs categorized by Biological Process (BP), Cellular Component (CC), and Molecular Function (MF). (D) KEGG pathway enrichment analysis for the overlapping candidate MCDRGs.
To restrict our focus to metabolic cell death pathways directly perturbed in disease pathology, we performed an intersection analysis. A Venn diagram was drawn to identify the intersecting DEGs between IBD and control samples, key module genes, and MCDRGs, and a total of 25 genes were screened (Fig. 2B) and defined as candidate MCDRGs. GO and KEGG enrichment analysis were performed on the 25 candidate MCDRGs. The GO results revealed that these MCDRGs were involved in several cellular processes, including positive regulation of the inflammatory response, oxidoreductase complex, and long-chain fatty acid-CoA ligase activity (Fig. 2C). Notably, the KEGG pathway analysis highlighted a significant convergence on the IL-17 signaling pathway, ferroptosis, HIF-1 signaling, and fatty acid metabolism (Fig. 2D). Collectively, these findings underscore the functional role of these candidate genes in bridging metabolic dysregulation with core inflammatory cascades.
Screening of key MCDRGs via integrated ML algorithms
To precisely identify the key MCDRGs in IBD, three distinct ML algorithms—LASSO regression, Boruta, and SVM-RFE—were used to screen the 25 previously identified candidate MCDRGs. First, a LASSO regression model was constructed and the optimal lambda value was determined to be 0.005852 through 10-fold cross-validation, representing the point of minimum error (Fig. 3A-B), resulting in the selection of 5 key MCDRGs. The Boruta algorithm was implemented to evaluate gene importance by comparing real feature variables with shadow variables, where genes represented by green boxplots exhibited significantly higher importance scores and were classified as confirmed (Fig. 3C-D). Additionally, SVM-RFE analysis identified the optimal feature combination by minimizing the 10-fold cross-validation error (Fig. 3E). Finally, the intersection of the genes identified by all three algorithms was determined using a Venn diagram, yielding five overlapping genes—IDO1, LCN2, MIR21, RRM2, and SLC6A14—as key MCDRGs (Fig. 3F; Supplementary Tables 6–8).
Fig. 3.

Screening of key MCDRGs in IBD by three ML algorithms. (A-B) LASSO logistic regression analysis with 10-fold cross-validation, showing the distribution of coefficients (A) and the partial likelihood deviance plot for selecting the optimal penalty parameter lambda (B). (C) Feature importance trajectory plot of the Boruta algorithm across multiple iterations, where green, yellow, and blue lines represent confirmed, tentative, and rejected features, respectively. (D) Importance ranking of candidate MCDRGs compared against shadow variables as determined by the Boruta classifier. (E) SVM-RFE algorithm illustrating the relationship between the 10-fold cross-validation error rate and the number of feature variables. (F) Venn diagram showing the overlap of genes identified by all three algorithms.
Multi-cohort validation and robust evaluation of key MCDRGs
To further refine the core diagnostic signature, we evaluated the ML-based MCDRGs across the discovery and two independent validation cohorts. Among the initial candidates, IDO1, LCN2, RRM2, and SLC6A14 consistently demonstrated significant expression differences with uniform directional shifts across all cohorts after BH-FDR correction, whereas MIR21 showed poor consistency and was excluded (Fig. 4A-C). Notably, IDO1, LCN2, and SLC6A14 exhibited superior diagnostic accuracy with AUCs > 0.85 in both the discovery and validation cohorts, while RRM2 was excluded owing to its lower diagnostic performance (Fig. 4D-F). All three genes showed statistically significant differential expression across three independent cohorts (9/9 BH-FDR-corrected comparisons; maximum FDR = 1.12 × 10⁻⁵), confirming their statistical robustness. Subtype-stratified analysis in the discovery cohort further confirmed that all three genes were significantly upregulated in both CD and UC relative to controls, with no significant difference between CD and UC (Figure S3 and Supplementary Tables 9, 14–18), supporting the rationale for combined IBD analysis in the primary model.
Fig. 4.

Expression validation and diagnostic efficacy evaluation of the key MCDRGs across multiple cohorts. (A-C) Relative expression levels of the key MCDRGs in the discovery set GSE75214 (A) and two independent validation sets GSE87466 (B) and GSE47908 (C) between IBD and control groups, with BH-FDR correction applied within each cohort. (D-F) ROC curves assessing the diagnostic performance of the key MCDRGs in the discovery set (D) and the respective validation sets (E-F). Error bars represent mean ± SD. ***FDR < 0.001, **FDR < 0.01; ns, not significant.
To address concerns regarding the near-perfect apparent AUC observed in the discovery cohort, we performed systematic evaluation using six statistical methods and multiple performance frameworks. The apparent AUC of 1.000 in the unpenalized logistic regression model was attributable to quasi-complete separation of the three-gene expression profile in GSE75214, as evidenced by highly divergent regression coefficients (IDO1: 169.27; LCN2: 60.00; SLC6A14: 187.63), a well-documented statistical phenomenon in linearly separable data that does not indicate overfitting or data leakage. Applying Firth penalized logistic regression and four additional penalized or Bayesian methods yielded stable, finite coefficients with consistent directionality across all methods (Figure S4D-E). Bootstrap-optimism-corrected AUC and repeated cross-validation AUC were consistently high across all six methods (Figure S4D; Supplementary Tables 10–13).
Importantly, LODO-CV provided an unbiased estimate of generalization performance across independent cohorts: the mean LODO-CV AUC was 0.975 (range 0.936–0.998), with held-out AUCs of 0.998 (GSE75214), 0.990 (GSE87466), and 0.936 (GSE47908) (Figure S4A). A permutation test (n = 1,000) confirmed non-random discrimination (p < 0.001). Calibration analysis of the LODO-CV held-out predictions demonstrated acceptable overall agreement between predicted and observed probabilities (Brier score = 0.044; mean absolute calibration error Eavg = 0.033; calibration slope = 0.995; n = 323; Figure S4B). Based on the Firth penalized regression, a clinical nomogram was constructed to facilitate individualized risk estimation (Figure S4C). Precision-recall analysis further confirmed consistently high precision across all three cohorts (Figure S4F). In summary, the consistent differential expression and robust cross-cohort discrimination of IDO1, LCN2, and SLC6A14 across multiple independent cohorts support their candidacy as metabolic cell death-related biomarkers for IBD diagnostic stratification.
Correlation analysis and GSEA of key MCDRGs
To further elucidate the potential synergism and biological significance of the identified key MCDRGs, we conducted an integrated analysis of their expression correlations and functional enrichment profiles. Spearman’s rank correlation analysis in the discovery set revealed that the expression levels of IDO1, LCN2, and SLC6A14 were significantly and positively correlated after BH-FDR correction (IDO1-LCN2: rho = 0.456, FDR = 2.55 × 10⁻⁹; IDO1-SLC6A14: rho = 0.578, FDR = 5.08 × 10⁻¹⁵; LCN2-SLC6A14: rho = 0.845, FDR = 5.13 × 10⁻⁴³; Fig. 5A and Supplementary Table 19), suggesting a coordinated regulatory role in IBD pathogenesis. Single-gene GSEA revealed a high degree of functional convergence among these three genes, with predominant enrichment in pathways associated with immune-inflammatory responses and protein metabolic regulation. Notably, all three genes were significantly associated with “Cytokine-cytokine receptor interaction” and “Proteasome” pathways, with consistent enrichment also observed in “Leishmania infection,” “ECM-receptor interaction,” and “Protein export” across the gene sets (Fig. 5B-D).
Fig. 5.

Correlation analysis and GSEA of key MCDRGs. (A) Correlation heatmap evaluating the BH-FDR-corrected inter-gene expression relationships among IDO1, LCN2, and SLC6A14 in the discovery set. (B-D) Single-gene GSEA identifying enriched KEGG signaling pathways associated with each hub gene (B: IDO1, C: LCN2, D: SLC6A14). ***FDR < 0.001 (by Spearman’s rank correlation analysis with BH-FDR correction).
Clinical validation and correlation analysis of key MCDRGs in human peripheral blood
To validate the ML-identified signature at the clinical level, peripheral blood samples from patients with active IBD (n = 16) and healthy controls (n = 30) were analyzed. The qRT-PCR analysis revealed that IDO1, LCN2 and SLC6A14 mRNA expression levels were significantly upregulated in patients with active IBD compared with healthy individuals (FDR = 2.02 × 10⁻¹²; Fig. 6A-C). Spearman’s rank correlation analysis performed on all peripheral blood samples with BH-FDR correction showed strong and significant positive correlations among IDO1, LCN2 and SLC6A14 mRNA expression (all FDR < 5.75 × 10⁻⁵; Fig. 6D and Supplementary Table 20). These clinical findings confirm the biological consistency of the metabolic cell death-derived candidate signature and support the coordinated alteration of these three markers during the active phase of IBD.
Fig. 6.

Clinical validation and correlation analysis of key MCDRGs in human peripheral blood. (A-C) Relative mRNA expression of IDO1 (A), LCN2 (B) and SLC6A14 (C) in peripheral blood samples from patients with active IBD (n = 16) and healthy controls (n = 30) measured by qRT-PCR. (D) Correlation heatmap evaluating mutual relationships among IDO1, LCN2 and SLC6A14 mRNA expression in the clinical cohort, with BH-FDR correction applied. Error bars represent mean ± SEM. ***FDR < 0.001 (by unpaired t-test or Spearman’s rank correlation analysis with BH-FDR correction).
Characterization of the immune infiltration landscape and its correlation with key MCDRGs
To characterize the immune microenvironment in IBD, the CIBERSORT algorithm was used to quantify the relative proportions of 22 infiltrating immune cell types within the discovery set, and significant disparities in immune infiltration were observed between IBD and control samples after BH-FDR correction. Specifically, 14 immune cell types exhibited distinct inter-group variations, including eight types such as M0 macrophages and activated mast cells that were significantly enriched in IBD samples, whereas the remaining six types, including M2 macrophages and resting mast cells, were predominantly enriched in the control group (Fig. 7A and Supplementary Table 21). Spearman’s rank correlation analysis with BH-FDR correction revealed that the three key MCDRGs exhibited consistent and significant positive correlations with resting CD4 + memory T cells, resting NK cells, and M1 macrophages, whereas significant negative correlations were observed with CD8 + T cells, regulatory T cells, and activated NK cells (Fig. 7B). Notably, IDO1 expression demonstrated a strong positive correlation with M1 macrophage abundance and a significant negative correlation with activated NK cell abundance (Fig. 7C-D).
Fig. 7.

Characterization of the immune landscape and its correlation with key MCDRGs. (A) Boxplot comparing the relative proportions of 22 immune cell types between IBD and control samples from the discovery set using the CIBERSORT algorithm, with BH-FDR correction applied. (B) Correlation heatmap illustrating the BH-FDR-corrected relationships between the three key MCDRGs and the abundance of various infiltrating immune cells. (C-D) Correlation scatter plots displaying the specific association between IDO1 expression and infiltrating levels of M1 macrophages (C) and activated NK cells (D). Error bars represent mean ± SD. *FDR < 0.05, **FDR < 0.01, ***FDR < 0.001, ****FDR < 0.0001; ns, not significant (by Wilcoxon rank-sum test or Spearman’s rank correlation analysis with BH-FDR correction). NK, natural killer.
Additionally, the ssGSEA results showed statistically significant differences in the infiltration abundance of 26 immune cell subsets between IBD and control groups after BH-FDR correction (Figure S5A-B). Correlation analysis showed that key MCDRGs were significantly and positively correlated with most immune cells after FDR correction, among which IDO1 had the highest positive correlation with activated dendritic cells (DCs) and effector memory CD8 + T cells, whereas SLC6A14 had the highest negative correlation with CD56dim NK cells (Figure S5C). Taken together, these results indicate that the identified key MCDRGs may play pivotal roles in modulating the inflammatory immune response in IBD by regulating specific cell recruitment or activation.
Single-cell RNA sequencing validation provides orthogonal, cell-type-resolved evidence for key MCDRG expression in IBD intestinal tissue
The immune cell infiltration analysis described above was based on computational deconvolution of bulk transcriptomic data (CIBERSORT and ssGSEA), which cannot resolve cell-type-specific gene expression at the individual-cell level. To provide orthogonal, cell-type-resolved validation of IDO1, LCN2, and SLC6A14 expression in the IBD inflammatory microenvironment—independent of bulk deconvolution assumptions—we analyzed a publicly available IBD single-cell RNA-sequencing dataset (GSE214695). This dataset comprises 18 intestinal mucosal biopsy samples from CD (n = 6), UC (n = 6), and healthy control (HC; n = 6) donors, representing an independent single-cell cohort distinct from the three transcriptomic cohorts used in the primary analysis.
After stringent quality control (200–5,000 features per cell, mitochondrial fraction < 20%), 44,941 high-quality cells were retained. Harmony-based batch correction across 18 donors effectively removed sample-level variation while preserving biologically meaningful cell-type structure, as confirmed by the intermingling of cells from different donors and disease groups in the corrected UMAP space (Fig. 8A). Unsupervised clustering identified 35 distinct clusters (Fig. 8A), which were annotated into seven major cell types using a three-source consensus strategy: (1) SingleR automated annotation using the HumanPrimaryCellAtlasData reference, (2) tissue-specific marker score analysis covering immune, epithelial, and stromal lineages, and (3) top differential marker gene inspection. The annotated atlas comprised T/NK cells (31.1%), Plasma cells (28.8%), Epithelial cells (16.5%), Myeloid cells (8.6%), B cells (7.7%), Fibroblasts (5.9%), and Endothelial cells (0.9%), consistent with the known cellular composition of IBD-inflamed intestinal mucosa (Fig. 8B). Cell-type composition analysis revealed distinct proportional differences across CD, UC, and HC groups, with T/NK and Plasma cells collectively comprising over 60% of cells across all groups (Fig. 8C; Supplementary Table 22).
Fig. 8.

Single-cell RNA sequencing validation of IDO1, LCN2, and SLC6A14 in IBD intestinal tissue. (A) Cluster UMAP split by disease group (HC, CD, UC) showing 44,941 cells from 18 donors (6 HC, 6 CD, 6 UC; GSE214695) after Harmony batch correction and unsupervised clustering (35 clusters), demonstrating comparable cluster distributions across disease groups. (B) Celltype UMAP showing the same cells colored by major cell type annotation: T/NK cells (31.1%), Plasma cells (28.8%), Epithelial cells (16.5%), Myeloid cells (8.6%), B cells (7.7%), Fibroblasts (5.9%), and Endothelial cells (0.9%). (C) Stacked bar chart showing mean proportional composition of the seven major cell types across HC, CD, and UC groups. (D) FeaturePlot showing normalized expression of IDO1, LCN2, and SLC6A14 overlaid on the UMAP embedding. (E) Violin plots with individual cell data points showing IDO1, LCN2, and SLC6A14 expression stratified by disease group (HC, CD, UC). (F) Dotplot showing cell-type proportions per sample across HC, CD, and UC groups with Wilcoxon pairwise significance testing (ns, not significant; *, p < 0.05; **, p < 0.01). UMAP, uniform manifold approximation and projection; CD, Crohn’s disease; UC, ulcerative colitis; HC, healthy control.
Single-cell-level expression analysis revealed that IDO1, LCN2, and SLC6A14 were each detectable in distinct cellular compartments (Fig. 8D). IDO1 expression was most prominent in Myeloid cells, consistent with its established role as an IFN-γ-inducible enzyme in activated myeloid-lineage cells, consistent with macrophage/DC-related inflammatory signals observed in the bulk CIBERSORT analysis24,25. LCN2 showed concentrated expression in Epithelial and Myeloid cells, supporting LCN2 as a mucosal inflammatory marker with epithelial and myeloid-associated expression patterns in IBD, in line with its association with inflammatory immune-cell infiltration observed in bulk immune profiling. SLC6A14 expression was enriched in Epithelial cells, consistent with its role as an epithelial amino acid transporter. Violin plots stratified by disease group further demonstrated that all three genes showed higher expression in IBD (CD and UC) compared to HC across the relevant cell types (Fig. 8E). The cell-type-specific expression patterns of all three genes are directionally consistent with the computational immune infiltration correlations derived from bulk transcriptomic deconvolution, while offering the additional resolution of cell-type-of-origin information not accessible by bulk methods alone.
To further characterize disease-associated cellular remodeling, we examined the proportional abundance of each major cell type across disease groups (Fig. 8F). Most cell types, including T/NK, B, Plasma, Myeloid, Fibroblast, and Endothelial cells, showed broadly comparable proportions between CD and UC, consistent with the shared IBD inflammatory transcriptomic signature identified in the bulk analysis. However, Epithelial cell proportion was significantly lower in UC compared to both HC (Wilcoxon test, p = 0.005) and CD (p = 0.013). Plasma cell proportion also differed significantly between UC and HC (p = 0.045), and between UC and CD (p = 0.031). Myeloid cell proportion was significantly elevated in UC versus HC (p = 0.045). These findings are consistent with the bulk subtype analysis in which UC-specific patterns in colonic mucosal gene expression were observed, and with the established biology of UC as a predominantly colonic mucosal disease with pronounced epithelial disruption. The selective reduction in Epithelial cell proportion in UC may reflect the more extensive mucosal erosion and epithelial barrier disruption characteristic of UC compared with CD. While the primary aim of this study was not to dissect CD versus UC differences at the single-cell level, these exploratory observations suggest that cell-type-specific composition analysis in larger single-cell cohorts may provide additional insights into the divergent mucosal biology of CD and UC, representing a promising direction for future investigation.
Discussion
In this study, we systematically integrated transcriptomic data with multiple ML algorithms to identify a metabolic cell death-related candidate diagnostic signature for IBD. Three key MCDRGs, IDO1, LCN2, and SLC6A14, demonstrated consistent differential expression and cross-cohort discriminatory performance across multiple independent datasets. To address the near-perfect apparent AUC observed in the discovery cohort, we systematically evaluated model performance using six statistical methods and multiple validation frameworks. The apparent AUC of 1.000 was attributable to quasi-complete separation, a statistically well-characterized phenomenon in linearly separable data, rather than overfitting or data leakage. Applying Firth penalized logistic regression and LODO-CV provided a more conservative and unbiased estimate of generalization performance, with a mean LODO-CV AUC of 0.975 (range 0.936–0.998) across three independent cohorts. Taken together, these results support the diagnostic potential of the identified candidate signature, while acknowledging that prospective clinical validation is necessary to establish its utility as a diagnostic tool.
IBD is characterized by pronounced heterogeneity across multiple dimensions, including disease subtype, inflammatory activity, anatomical location, and therapeutic response, posing significant challenges for the discovery and validation of reproducible transcriptomic biomarkers26,27. Previous studies have shown that pronounced biological heterogeneity in IBD frequently results in inconsistent or cohort-specific differentially expressed gene signatures, thereby limiting the reproducibility and clinical applicability of conventional statistical approaches28,29. The cross-cohort stability of our MCDRG-based signature suggests that it captured core molecular features of IBD pathogenesis. ML provides a disease-adapted analytical framework by integrating signals across heterogeneous patient subsets and prioritizing genes based on their collective discriminatory ability30. Importantly, integrating multiple feature selection algorithms helps accommodate the complexity and heterogeneity of transcriptomic data, as different methods capture complementary aspects of feature importance31–33.
IDO1 encodes a key enzyme involved in tryptophan catabolism, which plays a critical immunoregulatory role in mucosal environments. Clinical and experimental studies have consistently shown that IDO1 expression is markedly upregulated in the inflamed intestinal mucosa in IBD, and its activity correlates with disease severity and immune cell infiltration34,35. Recent mechanistic evidence has demonstrated that inflammation-induced IDO1 overexpression promotes kynurenine accumulation and subsequent activation of the aryl hydrocarbon receptor, which in turn is associated with lipid peroxidation and ferroptosis under certain conditions, suggesting a potential molecular link between IDO1 activity, redox imbalance, and regulated metabolic cell death36. Whether similar mechanisms operate in the inflamed gut warrants further mechanistic investigation.
LCN2 is highly expressed in the inflamed colonic mucosa of patients with UC and has been identified as a ferroptosis-related gene associated with disease progression through modulation of lipid peroxidation, iron accumulation, and GPX4 levels in colitis models37,38. Recent evidence indicates that Brusatol effectively suppresses LCN2-mediated ferroptosis to alleviate colonic injury by upregulating the ubiquitin ligase UBE3A39. In addition to ferroptosis, LCN2 has been implicated in pyroptosis through NLRP3 inflammasome activation in UC, suggesting that it may associate with multiple regulated cell death pathways and immune responses40.
SLC6A14, a Na+/Cl−-dependent amino acid transporter, has been shown to be upregulated in colonic tissues from patients with UC and mice with experimental colitis, where it is associated with epithelial injury through modulation of C/EBPbeta-PAK6 signaling41–43. Given that SLC6A14-mediated substrate uptake, such as glutamine and methionine, directly modulates intracellular antioxidant pools, it may occupy a strategic position in cellular susceptibility to metabolic stress. Further mechanistic studies are warranted to explore whether SLC6A14 serves as a broader regulator of metabolic cell death under nutrient or oxidative perturbation conditions.
Notably, our immune cell infiltration analysis revealed that IDO1 expression was positively correlated with M1 macrophages and DCs, whereas SLC6A14 exhibited the strongest negative correlation with CD56dim NK cells. Growing evidence indicates that IDO1 is associated with pro-inflammatory macrophage states in different contexts44,45, and IDO1 is upregulated by IFN-gamma to catalyze tryptophan-kynurenine metabolism in DCs, thereby modulating downstream immune homeostasis24,25. NK cells in patients with IBD exhibit signs of metabolic exhaustion, characterized by impaired mitochondrial function46. Given that the cytotoxic CD56dim NK subset relies on amino acid availability for its effector functions47,48, the inverse association between SLC6A14 and CD56dim NK cells may reflect competition for amino acid resources within the inflammatory microenvironment49. These associations, combined with the single-cell-level expression patterns in IBD mucosa, suggest that IDO1 and SLC6A14 may participate in immune-metabolic cross-talk relevant to IBD pathophysiology, though causal relationships require direct experimental validation.
An exploratory analysis of cell-type proportions in the single-cell cohort revealed exploratory differences between CD and UC that complement the bulk subtype analysis. While T/NK, B, Fibroblast, and Endothelial cells showed relatively stable proportions across HC, CD, and UC (T/NK: 30.0%, 32.9%, and 29.7%, respectively), the three most abundant non-lymphoid populations showed disease-associated variation. In particular, Epithelial cells showed a progressive reduction from HC (29.2%) through CD (17.8%) to UC (3.6%), with UC significantly lower than both HC (p = 0.005) and CD (p = 0.013). This reduction in epithelial cells in UC is consistent with the established pathobiology of UC as a colonic mucosal disease and aligns with the bulk expression finding that SLC6A14—an epithelial amino acid transporter enriched in epithelial cells—showed consistent differential expression in the UC-enriched GSE87466 and GSE47908 validation cohorts43,50,51. Concurrently, Plasma cells were increased in UC (43.1%) compared with CD (20.6%) and HC (23.2%), with UC significantly different from both comparators (p < 0.05). The increase in Plasma cells in UC may reflect mucosal humoral immune responses and is broadly consistent with the elevated immune-related signals observed in the bulk CIBERSORT analysis52. In contrast, Myeloid cells were higher in CD (15.6%) relative to HC (3.5%, p = 0.045) and UC (7.3%), suggesting disease-associated variation in the myeloid compartment and broadly aligning with the positive association between IDO1—which is most highly expressed in Myeloid cells—and macrophage-related inflammatory infiltration observed in bulk immune profiling53,54. These exploratory cell-type composition patterns are broadly consistent with the bulk subtype analysis, which showed subtype-related expression trends for the three signature genes without establishing subtype-specific mechanisms. Together, these exploratory observations suggest that cell-type-specific composition analysis in larger, well-annotated single-cell cohorts—with deeper examination of IDO1 expression in Myeloid subclusters in CD, Plasma cell IgA/IgG isotype distribution in UC, and SLC6A14-expressing epithelial progenitor dynamics—could provide additional insights into the mucosal biology of CD and UC, representing a promising direction for future investigation.
This study has several limitations. First, the analyses were primarily based on retrospective transcriptomic datasets. Although multiple external cohorts and peripheral blood samples were used for validation, prospective and multicenter studies are required to confirm clinical utility. Second, the mechanistic links between the identified MCDRGs and specific metabolic cell death modalities were inferred from bioinformatic and correlative analyses, and further perturbation-based experimental validation is necessary to establish causal relationships. Third, immune infiltration was assessed using bulk transcriptomic deconvolution methods, which may not fully capture cell-type-specific or spatial heterogeneities; the single-cell validation partially addresses this by providing cell-type-resolved expression evidence. Fourth, the clinical cohort was relatively small, limiting statistical power for subgroup analyses. Fifth, the three transcriptomic cohorts predominantly represent colonic IBD with variable treatment exposure, and the diagnostic performance of the signature in ileal CD, treatment-naive patients, or pediatric populations has not been evaluated. These recognized limitations are committed to addressing them in the comprehensive and functional validation studies in the future.
Conclusion
In summary, we identified a metabolic cell death-related candidate diagnostic signature for IBD using an integrated ML framework. The identified MCDRGs, IDO1, LCN2, and SLC6A14, collectively associate with metabolic dysregulation, immune infiltration, and regulated cell death pathways, offering insights into IBD pathophysiology. Cross-cohort validation using LODO-CV, combined with penalized regression and permutation testing, demonstrated robust generalization performance with a mean AUC of 0.975. Orthogonal single-cell validation confirmed cell-type-specific expression of all three genes in IBD-inflamed intestinal tissue. While these findings identify IDO1, LCN2, and SLC6A14 as a transcriptomic candidate signature with exploratory diagnostic potential for IBD stratification, prospective clinical studies and functional mechanistic validation are needed to establish their utility as diagnostic and therapeutic targets.
Supplementary Information
Below is the link to the electronic supplementary material.
Acknowledgements
The authors thank the NCBI GEO (https://www.ncbi.nlm.nih.gov/geo/) and FerrDb database (http://www.zhounan.org/ferrdb/current/) for providing valuable data resources. We also acknowledge the R programming language and its open-source community for enabling the bioinformatics analyses.
Author contributions
F.X. conceptualized the study, developed the methodology (including advanced machine learning optimization and statistical framework reconstruction), performed the formal analysis and investigation, curated the data, wrote the original draft, and executed the comprehensive text rewriting during revision. She also handled the visualization, supervision, and project administration. Z.H. contributed to validation, formal analysis, investigation, resources, data curation, visualization, and manuscript review and editing. Q.L. validated the findings, provided resources, and acquired funding. Y.C. contributed to the clinical investigation by collecting human peripheral blood samples and performing the biochemical assay and molecular verifications. F.X. and Z.H. contributed equally to this work as co-first authors. All authors reviewed and approved the final manuscript.
Funding
This work was supported by grants from the Cooperation Project on Scientific Research between Nanchong City and North Sichuan Medical College (No. 19SXHZ0344).
Data availability
The transcriptomic data used in this study were obtained from the publicly available GEO database under accession numbers GSE75214, GSE87466, GSE47908, and GSE214695. The datasets generated and analyzed during the current study, including processed expression matrices, model performance tables, and supplementary tables, are available from the corresponding author upon reasonable request. R code supporting the findings is available upon request.
Declarations
Competing interests
The authors declare no competing interests.
Ethics Statement
This study was approved by the Ethics Committee of the Affiliated Hospital of North Sichuan Medical College (Approval No. 2024ER187-2). Informed consent was obtained from all participants. Analysis of publicly available transcriptomic datasets (GSE75214, GSE87466, GSE47908, GSE214695) was performed in accordance with the terms of the original deposited studies.
Footnotes
Publisher’s note
Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
Fengmin Xu and Zhenting He contributed equally to this work.
References
- 1.Ng, S. C. et al. Worldwide incidence and prevalence of inflammatory bowel disease in the 21st century: a systematic review of population-based studies. Lancet (London England). 390, 2769–2778. 10.1016/s0140-6736(17)32448-0 (2017). [DOI] [PubMed] [Google Scholar]
- 2.Graham, D. B. & Xavier, R. J. Pathway paradigms revealed from the genetics of inflammatory bowel disease. Nature578, 527–539. 10.1038/s41586-020-2025-2 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 3.Greten, F. R. & Grivennikov, S. I. Inflammation and Cancer: Triggers, Mechanisms, and Consequences. Immunity51, 27–41. 10.1016/j.immuni.2019.06.025 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 4.Zeng, Z., Jiang, M., Li, X., Yuan, J. & Zhang, H. Precision medicine in inflammatory bowel disease. Precis Clin. Med.6, pbad033. 10.1093/pcmedi/pbad033 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 5.Li, T. et al. Identification of key genes as diagnostic biomarkers for IBD using bioinformatics and machine learning. J. translational Med.23 10.1186/s12967-025-06531-1 (2025). [DOI] [PMC free article] [PubMed]
- 6.Mao, C., Wang, M., Zhuang, L. & Gan, B. Metabolic cell death in cancer: ferroptosis, cuproptosis, disulfidptosis, and beyond. Protein cell.15, 642–660. 10.1093/procel/pwae003 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 7.Artusi, I., Rubin, M., Cravin, G. & Cozza, G. Ferroptosis in Human Diseases: Fundamental Roles and Emerging Therapeutic Perspectives. Antioxid. (Basel Switzerland). 14 10.3390/antiox14121411 (2025). [DOI] [PMC free article] [PubMed]
- 8.Yang, W., Xiao, C., Zheng, J., Song, J. & Li, X. G. Copper homeostasis and cuproptosis represent emerging targets for therapeutic intervention in inflammatory diseases. Pharmacol. Res.221, 107988. 10.1016/j.phrs.2025.107988 (2025). [DOI] [PubMed] [Google Scholar]
- 9.Wan, Y. et al. The Mechanism and Regulation of Disulfidptosis and Its Role in Disease. Biomedicines14 10.3390/biomedicines14010228 (2026). [DOI] [PMC free article] [PubMed]
- 10.Ocansey, D. K. W., Yuan, J., Wei, Z., Mao, F. & Zhang, Z. Role of ferroptosis in the pathogenesis and as a therapeutic target of inflammatory bowel disease (Review). Int. J. Mol. Med.51 10.3892/ijmm.2023.5256 (2023). [DOI] [PMC free article] [PubMed]
- 11.Ma, J. et al. Penicillamine ameliorates intestinal barrier damage in dextran sulfate sodium-induced experimental colitis mice by inhibiting cuproptosis. Front. Immunol.16 10.3389/fimmu.2025.1580963 (2025). [DOI] [PMC free article] [PubMed]
- 12.Tang, H., Li, P. & Guo, X. Ferroptosis-Mediated Immune Microenvironment and Therapeutic Response in Inflammatory Bowel Disease. DNA Cell Biol.42, 720–734. 10.1089/dna.2023.0260 (2023). [DOI] [PubMed] [Google Scholar]
- 13.Pu, Y., Meng, X. & Zou, Z. Identification and immunological characterization of cuproptosis-related molecular clusters in ulcerative colitis. BMC Gastroenterol.23, 221. 10.1186/s12876-023-02831-2 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 14.Judge, E., Gramatikoff, K., Milovic, L., Minchev, A. & Karabaliev, M. Precision Medicine and Multi-Omics Integration: Transforming Drug Discovery Through FAIR-Enabled Systems. EuroBiotech J.10, 1–6. 10.2478/ebtj-2026-0001 (2026). [DOI] [Google Scholar]
- 15.Jiang, Z., Zhang, H., Gao, Y. & Sun, Y. Multi-omics strategies for biomarker discovery and application in personalized oncology. Mol. Biomed.6, 115. 10.1186/s43556-025-00340-0 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 16.Bian, R., Xu, X. & Li, W. Uncovering the molecular mechanisms between heart failure and end-stage renal disease via a bioinformatics study. Front. Genet.13, 1037520. 10.3389/fgene.2022.1037520 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 17.Dai, Z. et al. A multi-omics pipeline integrating machine learning and spatial-cellular analysis identifies SASH1 as a prognostic biomarker and therapeutic target in head and neck squamous cell carcinoma. Int. J. Surg. (London England). 111, 9178–9195. 10.1097/js9.0000000000003647 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 18.Liu, X. et al. Generalized machine learning based on multi-omics data to profile the effect of ferroptosis pathway on prognosis and immunotherapy response in patients with bladder cancer. Environ. Toxicol.39, 680–694. 10.1002/tox.23949 (2024). [DOI] [PubMed] [Google Scholar]
- 19.Chen, Y. et al. Identification of a prognostic cuproptosis-related signature in hepatocellular carcinoma. Biol. Direct. 18 10.1186/s13062-023-00358-w (2023). [DOI] [PMC free article] [PubMed]
- 20.Ni, L., Yang, H., Wu, X., Zhou, K. & Wang, S. The expression and prognostic value of disulfidptosis progress in lung adenocarcinoma. Aging15, 7741–7759. 10.18632/aging.204938 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 21.Li, J. et al. Association Between Diverse Cell Death Patterns Related Gene Signature and Prognosis, Drug Sensitivity, and Immune Microenvironment in Glioblastoma. J. Mol. neuroscience: MN. 74, 10. 10.1007/s12031-023-02181-4 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 22.Zheng, M., Tian, Q., Shen, J. & Li, S. Dysregulated immunometabolism in gut inflammation. Acta Biochim. Biophys. Sin.58, 169–182. 10.3724/abbs.2025192 (2026). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 23.Verdugo-Meza, A., Ye, J., Dadlani, H., Ghosh, S. & Gibson, D. L. Connecting the Dots Between Inflammatory Bowel Disease and Metabolic Syndrome: A Focus on Gut-Derived Metabolites. Nutrients12 10.3390/nu12051434 (2020). [DOI] [PMC free article] [PubMed]
- 24.Ketelhuth, D. F. J. The immunometabolic role of indoleamine 2,3-dioxygenase in atherosclerotic cardiovascular disease: immune homeostatic mechanisms in the artery wall. Cardiovascular. Res.115, 1408–1415. 10.1093/cvr/cvz067 (2019). [DOI] [PubMed] [Google Scholar]
- 25.Gargaro, M. et al. Indoleamine 2,3-dioxygenase 1 activation in mature cDC1 promotes tolerogenic education of inflammatory cDC2 via metabolic communication. Immunity55, 1032–1050e1014. 10.1016/j.immuni.2022.05.013 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 26.Tien, N. T. N. et al. An exploratory multi-omics study reveals distinct molecular signatures of ulcerative colitis and Crohn’s disease and their correlation with disease activity. J. Pharm. Biomed. Anal.255, 116652. 10.1016/j.jpba.2024.116652 (2025). [DOI] [PubMed] [Google Scholar]
- 27.Rubin, S. J. S. et al. Mass cytometry reveals systemic and local immune signatures that distinguish inflammatory bowel diseases. Nat. Commun.10, 2686. 10.1038/s41467-019-10387-7 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 28.Yang, C. et al. Identification of cuproptosis hub genes contributing to the immune microenvironment in ulcerative colitis using bioinformatic analysis and experimental verification. Front. Immunol.14, 1113385. 10.3389/fimmu.2023.1113385 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 29.Xie, X. P., Xie, Y. F., Liu, Y. T. & Wang, H. Q. Adaptively capturing the heterogeneity of expression for cancer biomarker identification. BMC Bioinform.19, 401. 10.1186/s12859-018-2437-2 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 30.He, J. et al. The practical implementation of artificial intelligence technologies in medicine. Nat. Med.25, 30–36. 10.1038/s41591-018-0307-0 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 31.Fleuren, L. M. et al. Machine learning for the prediction of sepsis: a systematic review and meta-analysis of diagnostic test accuracy. Intensive Care Med.46, 383–400. 10.1007/s00134-019-05872-y (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 32.He, T. et al. Novel Ensemble Feature Selection Approach and Application in Repertoire Sequencing Data. Front. Genet.13, 821832. 10.3389/fgene.2022.821832 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 33.Wang, X. & Zhang, L. Integrative machine learning identifies robust inflammation-related diagnostic biomarkers and stratifies immune-heterogeneous subtypes in Kawasaki disease. Pediatr. Rheumatol. Online J.23 10.1186/s12969-025-01114-2 (2025). [DOI] [PMC free article] [PubMed]
- 34.Proietti, E. et al. Modulation of Indoleamine 2,3-Dioxygenase 1 During Inflammatory Bowel Disease Activity in Humans and Mice. Int. J. tryptophan research: IJTR. 16, 11786469231153109. 10.1177/11786469231153109 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 35.Shon, W. J., Lee, Y. K., Shin, J. H., Choi, E. Y. & Shin, D. M. Severity of DSS-induced colitis is reduced in Ido1-deficient mice with down-regulation of TLR-MyD88-NF-kB transcriptional networks. Sci. Rep.5, 17305. 10.1038/srep17305 (2015). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 36.Cheng, Z. et al. Ferroptosis mediated by the IDO1/Kyn/AhR pathway triggers acute thymic involution in sepsis. Cell Death Dis.16 10.1038/s41419-025-07882-9 (2025). [DOI] [PMC free article] [PubMed]
- 37.Zollner, A. et al. Faecal Biomarkers in Inflammatory Bowel Diseases: Calprotectin Versus Lipocalin-2-a Comparative Study. J. Crohn’s colitis. 15, 43–54. 10.1093/ecco-jcc/jjaa124 (2021). [DOI] [PubMed] [Google Scholar]
- 38.Deng, L. et al. Identification of Lipocalin 2 as a Potential Ferroptosis-related Gene in Ulcerative Colitis. Inflamm. Bowel Dis.29, 1446–1457. 10.1093/ibd/izad050 (2023). [DOI] [PubMed] [Google Scholar]
- 39.Guan, X. et al. Brusatol Exerts Therapeutic Effects in Ulcerative Colitis by Regulating the UBE3A-LCN2-Mediated Ferroptosis and Inflammation Axis. Chem. Biol. Drug Des.106, e70201. 10.1111/cbdd.70201 (2025). [DOI] [PubMed] [Google Scholar]
- 40.Yang, Y. et al. Lipocalin-2-mediated intestinal epithelial cells pyroptosis via NF-κB/NLRP3/GSDMD signaling axis adversely affects inflammation in colitis. Biochim. et Biophys. acta Mol. basis disease. 1870, 167279. 10.1016/j.bbadis.2024.167279 (2024). [DOI] [PubMed] [Google Scholar]
- 41.Papalazarou, V. et al. Phenotypic profiling of solute carriers characterizes serine transport in cancer. Nat. metabolism. 5, 2148–2168. 10.1038/s42255-023-00936-2 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 42.Ren, X. J., Zhang, M. L., Shi, Z. H. & Zhu, P. P. SLC6A14 as a Key Diagnostic Biomarker for Ulcerative Colitis: An Integrative Bioinformatics and Machine Learning Approach. Biochem. Genet.64, 197–212. 10.1007/s10528-025-11027-0 (2026). [DOI] [PubMed] [Google Scholar]
- 43.Chen, Y. et al. SLC6A14 facilitates epithelial cell ferroptosis via the C/EBPβ-PAK6 axis in ulcerative colitis. Cell. Mol. Life Sci.79, 563. 10.1007/s00018-022-04594-7 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 44.Gao, Z. et al. IDO1 induced macrophage M1 polarization via ER stress-associated GRP78-XBP1 pathway to promote ulcerative colitis progression. Front. Med.12, 1524952. 10.3389/fmed.2025.1524952 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 45.Kuang, S. et al. Indoleamine 2,3-dioxygenase 1 mediated macrophage oxidative phosphorylation impairment drives pro-inflammatory M1 polarization aggravates acetaminophen-induced acute liver injury. Ecotoxicol. Environ. Saf.310, 119774. 10.1016/j.ecoenv.2026.119774 (2026). [DOI] [PubMed] [Google Scholar]
- 46.Zaiatz Bittencourt, V., Jones, F., Tosetto, M., Doherty, G. A. & Ryan, E. J. Dysregulation of Metabolic Pathways in Circulating Natural Killer Cells Isolated from Inflammatory Bowel Disease Patients. J. Crohn’s colitis. 15, 1316–1325. 10.1093/ecco-jcc/jjab014 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 47.De Federicis, D. et al. Nutrient transporter pattern in CD56(dim) NK cells: CD16 (FcγRIIIA)-dependent modulation and association with memory NK cell functional profile. Front. Immunol.15, 1477776. 10.3389/fimmu.2024.1477776 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 48.Presnell, S. R. et al. Differential Fuel Requirements of Human NK Cells and Human CD8 T Cells: Glutamine Regulates Glucose Uptake in Strongly Activated CD8 T Cells. ImmunoHorizons4, 231–244. 10.4049/immunohorizons.2000020 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 49.Domagala, J. et al. The Tumor Microenvironment-A Metabolic Obstacle to NK Cells’ Activity. Cancers12 10.3390/cancers12123542 (2020). [DOI] [PMC free article] [PubMed]
- 50.Sikder, M. O. F. et al. SLC6A14, a Na+/Cl–coupled amino acid transporter, functions as a tumor promoter in colon and is a target for Wnt signaling. Biochem. J.477, 1409–1425. 10.1042/bcj20200099 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 51.Ahmadi, S. et al. SLC6A14, an amino acid transporter, modifies the primary CF defect in fluid secretion. eLife7 10.7554/eLife.37963 (2018). [DOI] [PMC free article] [PubMed]
- 52.Scheid, J. F. et al. Remodeling of colon plasma cell repertoire within ulcerative colitis patients. J. Exp. Med.220 10.1084/jem.20220538 (2023). [DOI] [PMC free article] [PubMed]
- 53.Suau, R., Pardina, E., Domènech, E., Lorén, V. & Manyé, J. The Complex Relationship Between Microbiota, Immune Response and Creeping Fat in Crohn’s Disease. J. Crohn’s colitis. 16, 472–489. 10.1093/ecco-jcc/jjab159 (2022). [DOI] [PubMed] [Google Scholar]
- 54.Sewell, G. W., Marks, D. J. & Segal, A. W. The immunopathogenesis of Crohn’s disease: a three-stage model. Curr. Opin. Immunol.21, 506–513. 10.1016/j.coi.2009.06.003 (2009). [DOI] [PMC free article] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Data Availability Statement
The transcriptomic data used in this study were obtained from the publicly available GEO database under accession numbers GSE75214, GSE87466, GSE47908, and GSE214695. The datasets generated and analyzed during the current study, including processed expression matrices, model performance tables, and supplementary tables, are available from the corresponding author upon reasonable request. R code supporting the findings is available upon request.
