Skip to main content
BMC Gastroenterology logoLink to BMC Gastroenterology
. 2026 Jun 15;26:517. doi: 10.1186/s12876-026-05008-9

Shared molecular signatures linking gastroesophageal reflux disease and major depressive disorder revealed by integrated machine learning

Zirong Wang 1,#, Yongjia Chai 1,#, Sida Pi 2,#, Dilihumaer Tuerxun 1, Yixuan Zhu 1, Zihan Liu 1, Huayang Zhang 1, Jingyu Li 1, Chen Wang 1, Zihao Li 1, Ying Kang 1, Yanan Li 1, Na Zhao 1, Yan Tian 1, Kai Xiong 1,✉, Peng Zhang 1,✉
PMCID: PMC13495282  PMID: 42298439

Abstract

Background

Gastroesophageal reflux disease (GERD) and major depressive disorder (MDD) frequently co-occur, yet their shared molecular mechanisms remain poorly understood. This study aimed to identify a robust gene signature linking GERD and MDD and to evaluate its clinical relevance.

Methods

Transcriptomic data from five GEO datasets were analyzed, including three GERD datasets (52 GERD samples and 23 controls) and two MDD datasets (138 MDD samples and 76 controls). The GERD datasets were integrated after batch correction for differential expression analysis, whereas the MDD datasets were analyzed separately for WGCNA, model development, and external validation. Candidate genes were identified from the intersection of GERD differentially expressed genes and MDD-associated co-expression module genes. A total of 71 machine learning models were systematically evaluated in the MDD training cohort using stratified 10-fold cross-validation, with model selection based on area under the curve (AUC). The clinical relevance of the prioritized genes was further examined in an independent cohort of 70 GERD patients, in whom gene expression was measured and depressive symptom severity was stratified into three PHQ-9-defined groups: non-depressed (n = 39), mild-to-moderate depression (n = 19), and major depression (n = 12).

Results

A three-gene signature (SAMD14, PROC, and NRG1) was identified. Among 71 candidate models, ridge regression achieved the best performance, with a cross-validated AUC of 0.709. After refitting, the model yielded an AUC of 0.896 (95% CI, 0.851–0.941) in the training cohort and 0.875 (95% CI, 0.712–1.000) in the external test cohort. Additional performance metrics included a sensitivity of 0.805, specificity of 0.844, accuracy of 0.818, and F1 score of 0.855 in the training set. In the clinical cohort, expression levels of SAMD14, PROC, and NRG1 differed significantly across depression severity groups (all P < 0.001). DeMeester scores also differed among groups (P < 0.001), with higher values observed in patients with more severe depressive symptoms. Correlation analysis further demonstrated significant associations between gene expression and reflux-related parameters.

Conclusions

We identified a three-gene signature derived from molecular features shared between GERD and MDD. The signature showed discriminative performance for MDD classification and was associated with reflux-related clinical parameters in GERD patients. These findings support a hypothesis-generating molecular framework for investigating GERD-MDD overlap, while further prospective and mechanistic validation remains required.

Supplementary Information

The online version contains supplementary material available at 10.1186/s12876-026-05008-9.

Keywords: Gastroesophageal reflux disease, Major depressive disorder, Gut–brain axis, Machine learning, Transcriptomics, Molecular biomarkers

Highlights

An integrative machine learning framework was used to investigate molecular overlap between GERD and MDD. A three-gene signature (SAMD14, PROC, NRG1) derived from shared GERD-MDD features was identified. The gene signature was associated with immune infiltration patterns and reflux-related clinical indices, supporting its potential value for hypothesis generation and future validation.

Supplementary Information

The online version contains supplementary material available at 10.1186/s12876-026-05008-9.

Introduction

Gastroesophageal reflux disease (GERD) and major depressive disorder (MDD) have emerged as major global health challenges, affecting millions of individuals worldwide. Both conditions are characterized by complex pathophysiological mechanisms in which biopsychosocial factors play a critical role [1]. Clinically, GERD presents with a broad spectrum of manifestations. Typical symptoms include heartburn, chest pain, and acid regurgitation, whereas atypical manifestations encompass chronic cough with sputum production, chest tightness, dyspnea, hoarseness, globus sensation, nausea, abdominal pain, and other dyspeptic symptoms [2]. Notably, these mild or atypical symptoms are often underestimated by patients, leading to prolonged physical discomfort and psychological distress, which may gradually precipitate psychiatric disorders.

Accumulating evidence suggests that GERD is associated with an increased risk of subsequent depression, anxiety, and sleep disorders, all of which substantially impair quality of life. Therefore, clinicians should pay particular attention to psychiatric comorbidities in patients with GERD [3]. Mendelian randomization studies have further provided genetic evidence supporting a potential causal relationship between GERD and psychiatric disorders, indicating that these conditions may share common biological underpinnings. A study published in 2023 demonstrated a significant association between GERD and an increased risk of MDD [4]. MDD is one of the most prevalent psychiatric disorders globally and is characterized by persistent low mood and loss of interest or pleasure [5].

According to the World Health Organization Report Depression and Other Common Mental Disorders: Global Health Estimates, approximately 322 million people worldwide suffer from depression, accounting for 4.4% of the global population. In addition, anxiety disorders affect more than 260 million individuals (3.6% of the global population). The prevalence of these common mental disorders continues to rise, particularly in low- and middle-income countries, where comorbidity between depression and anxiety is highly prevalent [6]. Previous studies have extensively explored the association between GERD and various comorbid conditions, including Sjögren’s syndrome [7], smoking behavior [8], and chronic obstructive pulmonary disease [9]. However, investigations into the intrinsic biological link between GERD and MDD remain relatively limited. Notably, the association between GERD and depressive symptoms remains controversial. For instance, a prospective study involving 225 patients with GERD reported no significant increase in depressive symptoms [10].

In that study, participants underwent 24-hour ambulatory pH-impedance monitoring, and anxiety and depression were assessed using the Hospital Anxiety and Depression Scale (HADS). The null finding may be attributed to several factors: first, the sample size was relatively small (n = 225), which may limit the statistical power to detect subtle differences in depressive symptoms; second, the study focused on the “effects of anxiety and depression on GERD”, with a primary endpoint of evaluating the impact of psychiatric symptoms on GERD severity rather than the incidence or progression of depressive symptoms in GERD patients. In contrast, large-scale population-based cohort studies [3] and Mendelian randomization studies [4] have confirmed a positive association between GERD and MDD risk, suggesting that the inconsistency across studies may stem from differences in study design, sample size, and outcome measures. These discrepant findings underscore the need to clarify the biological relationship between GERD and MDD. Moreover, MDD has been identified as a mediating factor in the association between GERD and migraine [11].

Collectively, these conflicting and complementary findings have directed research attention toward the “brain–gut axis.” In recent years, the brain–gut axis theory has been increasingly recognized as a key mechanism underlying the interaction between gastrointestinal and psychiatric disorders. As a bidirectional communication network involving the central nervous system (CNS), enteric nervous system (ENS), and the gut microbiota, the brain–gut axis regulates gastrointestinal function and emotional responses through neural, endocrine, and immune pathways [12]. More directly relevant to our research, a recent cross-sectional study on young adults demonstrated that depression is strongly associated with GERD and proposed that psychological factors may alter pain perception and exacerbate GERD symptoms through the brain–gut axis, forming a self-perpetuating cycle [13]. Taken together, these observations support the brain–gut axis as a plausible framework for understanding the overlap between GERD and depression. Against this background, we sought to identify shared molecular features of GERD and MDD, to construct a GERD-informed transcriptomic classifier of MDD, and to examine the clinical relevance of the resulting gene signature in an independent GERD cohort

Methods

Data acquisition and preprocessing

Transcriptomic datasets related to GERD and MDD were retrieved from the Gene Expression Omnibus (GEO) database (https://www.ncbi.nlm.nih.gov/geo/). To ensure a clear analytical framework, GERD and MDD datasets were assigned distinct roles in this study: GERD datasets were used for differential expression analysis and feature discovery, whereas MDD datasets were used for network construction and machine learning model development. Three GERD-related datasets (GSE148381, GSE190027, and GSE226303; n = 75, including 52 patients and 23 controls) were integrated for differential expression analysis. To minimize non-biological variability across platforms, batch effect correction was performed using the ComBat function from the sva package, and its effectiveness was assessed by principal component analysis (PCA). Two MDD-related datasets (GSE98793 and GSE52790; n = 214, including 138 patients and 76 controls) were used for downstream analyses(Table 1). GSE98793 was designated as the training cohort for model construction and internal cross-validation, while GSE52790 served as an independent external validation cohort. Because GSE98793 and GSE52790 were not merged for downstream analysis, cross-dataset batch correction was not applied; instead, each dataset was preprocessed and normalized separately.

Table 1.

GEO datasets included in this study and their analytical roles

Dataset ID Disease Samples (Case/Control) Platform / Technology Role in Study
GSE148381 GERD 6/7(13 total) Illumina NovaSeq 6000 (Homo sapiens) GERD DEG Training
GSE190027 GERD 8/8(16 total) Illumina NovaSeq 6000 (Homo sapiens) GERD DEG Training
GSE226303 GERD 38/8(46 total) Illumina NextSeq 500 (Homo sapiens) GERD DEG Training
GSE98793 MDD 128/ 64 (192 total) Affymetrix Human Genome U133 Plus 2.0 Array MDD Training / WGCNA
GSE52790 MDD 10 / 12 (22 total) Affymetrix Human hGlue_3_0_v1 Array MDD External Validation

Differential expression analysis in GERD

Differentially expressed genes (DEGs) were identified in the batch-corrected merged GERD cohort using the limma package. Gene-wise linear models were fitted with disease status (GERD vs. control) as the grouping variable, and differential expression statistics were obtained using the empirical Bayes–moderated t-test implemented in eBayes. Multiple testing correction was performed using the Benjamini–Hochberg method, and genes with |log2 fold change| > 1 and adjusted P < 0.05 were considered statistically significant.

Identification of MDD-related gene signatures

Unlike the GERD datasets, which were used to identify robust differential signals after cross-dataset integration, the MDD datasets were analyzed using a network-based approach to capture phenotype-associated co-expression structure. Weighted gene co-expression network analysis (WGCNA) was performed to identify molecular features associated with MDD. Because transcriptomic differences in the MDD cohort were relatively subtle, a network-based strategy was adopted to capture coordinated gene expression patterns associated with the MDD phenotype. Initially, hierarchical clustering was conducted to detect and remove outlier samples. The soft-thresholding power (β) was selected using the pickSoftThreshold function to achieve a scale-free topology fit index close to 0.9. An adjacency matrix was then constructed and transformed into a topological overlap matrix (TOM), followed by hierarchical clustering based on TOM dissimilarity. Gene modules were identified using the dynamic tree cut algorithm (cutreeDynamic), with a minimum module size of 30 genes. Highly similar modules were subsequently merged. To identify modules most relevant to MDD, module eigengenes (MEs), defined as the first principal component of each module, were calculated across all samples and correlated on a sample-by-sample basis with MDD status encoded as a binary trait (MDD = 1, Control = 0) within the WGCNA framework. Modules showing the strongest positive associations with MDD were selected for further analysis. Genes within these modules were defined as MDD-related candidate genes.

Identification of GERD–MDD comorbidity-related genes

To identify biologically meaningful genes associated with GERD–MDD comorbidity, MDD-related WGCNA module genes were intersected with GERD DEGs. The resulting shared genes were then further intersected with genes present in both the MDD training dataset (GSE98793) and the validation dataset (GSE52790), and the final overlapping gene set was used as the candidate feature set for subsequent machine learning analyses.

Integrated machine learning modeling and cross-validation

Using the expression profiles of the selected candidate genes, classification models were constructed in the training MDD dataset (GSE98793) to distinguish MDD patients from healthy controls. Importantly, GERD datasets were not used for model training or validation but contributed exclusively through the upstream feature-selection step, in which candidate genes were restricted to shared GERD-MDD molecular features. The caret R package was used to train and compare multiple machine learning algorithms, including Random Forest, Support Vector Machine, Elastic Net, and XGBoost. Within the training cohort, 10-fold cross-validation was applied for hyperparameter tuning, with the objective of maximizing the area under the receiver operating characteristic curve (AUC). The optimized models were subsequently evaluated in an independent external dataset (GSE52790) to assess generalizability. Importantly, GERD datasets were not involved in model training or validation but contributed to feature selection through prior integration with MDD-associated gene modules. This design aimed to enhance biological relevance of the candidate features while minimizing potential information leakage between datasets. Detailed model specifications and associated parameter labels/settings used in model comparison are provided in Supplementary Table S1.

CIBERSORT analysis of feature genes

To provide immunological context for the feature genes identified by the optimal model, CIBERSORT analysis was performed to estimate the relative proportions of 22 immune cell types from bulk transcriptomic data. Correlations between the model-derived feature genes and immune cell fractions were then assessed to explore whether these genes were linked to shared immune-related mechanisms potentially relevant to GERD–MDD comorbidity.

Identification of candidate therapeutic agents

To explore potential drug–gene relationships for the final prioritized key genes identified from the optimal model, Enrichr was used as an over-representation-based enrichment query platform, focusing on the Drug Signatures Database (DSigDB). The analysis was performed using the final key genes selected for downstream biological interpretation, rather than the full set of shared genes identified in the earlier screening steps.

Clinical sample collection and qRT-PCR validation

A STROBE-compliant flow diagram illustrating recruitment into the clinical validation cohort is provided in Fig. 1. Between 2024 and 2025, a total of 70 adult patients with GERD were enrolled in this study. Depressive symptom severity was assessed using the Patient Health Questionnaire-9 (PHQ-9) [14]. Patients were then stratified into three predefined PHQ-9 severity categories: non-depressed group (0–4 points), mild-to-moderate depression symptoms group (5–14 points), and major depression group (15–27 points). qRT-PCR-based gene expression comparisons and clinical subgroup analyses were conducted across these three strata. GERD inclusion criteria required fulfillment of at least one of the following: a: typical reflux symptoms (heartburn, regurgitation, and/or acid reflux) responsive to proton pump inhibitor (PPI) therapy; b: endoscopic evidence of reflux esophagitis (Los Angeles grade B or higher), Barrett’s esophagus (histologically confirmed), or reflux-related esophageal stricture; c: atypical upper gastrointestinal symptoms with negative endoscopic findings (no erosions or LA grade A) and abnormal acid exposure confirmed by 24-hour esophageal pH monitoring. Exclusion criteria included: a: other major upper gastrointestinal diseases mimicking or coexisting with GERD (e.g., peptic ulcer disease); b: severe comorbidities such as advanced cardiac disease, active malignancy, or other conditions that could interfere with treatment or outcomes; c: pregnancy or lactation; and d: recent surgery or treatments potentially confounding study results. For each participant, 5 mL of peripheral blood was collected. Additional reflux-related clinical parameters were collected when available, including the DeMeester score and the total number of reflux episodes obtained from 24-hour esophageal pH monitoring, for subsequent correlation and subgroup analyses. Red blood cells were lysed using erythrocyte lysis buffer (Solarbio, China). After erythrocyte lysis, the remaining nucleated cells were collected by centrifugation at 450 × g for 10 min at room temperature, and the resulting cell pellets were lysed with TRIzol reagent (Takara Bio, China) for total RNA extraction. Total RNA was extracted and purified using chloroform, isopropanol, and ethanol, followed by RNA quantification and normalization to a final concentration of 1000 ng/µL. RNA was reverse-transcribed into cDNA, and quantitative real-time PCR (qRT-PCR) was performed using a Bio-Rad system (USA) in triplicate. Relative gene expression levels were calculated using the 2^−ΔΔCt method, with GAPDH as the housekeeping gene. Primer sequences were as follows:

  • SAMD14: Forward 5′-CCGTGGATGAAGTTTTTGACCT-3′; Reverse 5′-GATCCAGGCAGAAAGAGCCC-3′

  • PROC: Forward 5′-GGACGGCGAACTTGCAGTAT-3′; Reverse 5′-GACCAGAAGGCCAGTGTGTC-3′

  • NRG1: Forward 5′-GGAGCATATGTGTCTTCAGCTAC -3′; Reverse 5′- TGCTTGTAGAAGCTGGCCATT -3′

  • GAPDH: Forward 5′-GAAGGTGAAGGTCGGAGTC-3′; Reverse 5′-GGGTGGAATCATATTGGAAC-3′.

Fig. 1.

Fig. 1

STROBE flow diagram of patient recruitment for the clinical validation cohort. After exclusions due to missing depression assessment, missing gene expression data, or other criteria, 70 patients were included and categorized into no depression (n = 39), mild-to-moderate depression (n = 19), and major depression (n = 12)

Statistical analysis

All statistical analyses were performed using R software (R version 4.5.1; The R Foundation for Statistical Computing, 2025). For differential expression analysis in the GERD datasets, gene-wise linear modeling and empirical Bayes moderation were performed using the limma framework, and multiple testing correction was applied using the Benjamini–Hochberg method. For module–trait association analysis in WGCNA, Pearson correlation was used to evaluate the relationship between module eigengenes and MDD status encoded as a binary trait. For downstream association analyses involving gene expression, immune cell fractions, or clinical variables, Spearman correlation was used as a nonparametric measure of association. For two-group comparisons of continuous variables, including qRT-PCR validation analyses, either Student’s t-test or the Wilcoxon rank-sum test was applied as appropriate according to data distribution. Unless otherwise specified, a two-sided P value < 0.05 was considered statistically significant. For comparisons across the three PHQ-9-defined clinical groups, continuous variables were analyzed using one-way ANOVA or the Kruskal–Wallis test, as appropriate. Categorical variables were compared using the chi-square test or Fisher’s exact test, as appropriate. For correlation analyses in the clinical cohort, Spearman correlation coefficients were calculated, and the corresponding P values were reported as indicated.

Results

Research flowchart

Figure 2 summarizes the overall analytical workflow of the study. Briefly, the GERD datasets were integrated for differential expression analysis after preprocessing and batch-effect correction, whereas the MDD datasets were preprocessed separately and used for WGCNA, model training, and external validation without cross-dataset merging. Intersection of GERD DEGs with MDD-associated module genes yielded 78 shared genes, which were further restricted to 70 genes present in both the MDD training and validation cohorts for downstream machine learning analyses.

Fig. 2.

Fig. 2

Overview of the study workflow for identifying and validating shared molecular signatures between GERD and MDD

Multi-cohort Integration, Batch Effect Correction, and Visualization of GERD Datasets

Batch effect correction was performed on gene expression matrices from different platforms using the ComBat algorithm. Principal component analysis (PCA) showed that batch-driven clustering of samples was markedly attenuated after correction, whereas disease-specific clustering became more prominent. These results indicate that batch effects were effectively mitigated, and the integrated GERD dataset was suitable for subsequent downstream analyses (Fig. 3A, B).

Fig. 3.

Fig. 3

Integration of GERD transcriptomic datasets and identification of MDD-associated co-expression modules. A Principal component analysis (PCA) of GERD samples before batch correction, showing separation primarily driven by dataset origin. B PCA of GERD samples after ComBat correction, demonstrating reduction of batch-associated variation and improved integration across datasets. C Volcano plot of differential gene expression analysis in GERD samples, showing significantly upregulated and downregulated genes based on adjusted P-value and log2 fold-change thresholds. D Hierarchical clustering dendrogram of genes based on topological overlap, with module colors assigned using the dynamic tree cut method. E Module–trait correlation matrix showing associations between module eigengenes and MDD status in the MDD training cohort; color intensity represents Pearson correlation coefficients

Differential expression analysis in the GERD cohort and construction of MDD-related WGCNA modules

In the merged GERD cohort, a total of 2,172 differentially expressed genes (DEGs) were identified, including 1,729 upregulated and 443 downregulated genes (Fig. 3C). In contrast, differential expression analysis in the primary MDD cohort revealed relatively subtle global transcriptional differences. Under stringent thresholds (|log2FC| > 1, adjusted P < 0.05), no statistically significant DEGs were detected, suggesting substantial molecular heterogeneity in MDD. Therefore, a network-based analytical strategy was adopted to capture coordinated gene expression patterns associated with MDD. Weighted gene co-expression network analysis (WGCNA) was performed on 20,814 genes across 192 samples. Based on scale-free topology criteria, a soft-thresholding power of β = 7 was selected. Multiple co-expression modules were subsequently identified. Module–trait correlation analysis revealed that four modules—black, salmon, lightgreen, and cyan—were significantly positively correlated with the MDD phenotype (R = 0.18, 0.19, 0.17, and 0.16; P = 0.02, 0.01, 0.03, and 0.04, respectively), whereas no significant associations were observed for the remaining modules. A total of 726 genes contained within these four modules were therefore considered putative MDD-associated functional genes and were carried forward for comorbidity-focused analyses. The module–trait correlation matrix was visualized to summarize the strength and direction of associations, rather than to represent gene-level clustering (Fig. 3D, E).

3.4Identification of GERD–MDD comorbidity genes and machine learning candidate features

Intersection of GERD DEGs with MDD-associated WGCNA module genes yielded 78 shared genes, representing a potential molecular transcriptional overlap between GERD and MDD. These genes may reflect biological processes involved in gut–brain axis-mediated pathophysiology. Further intersection with genes present in both the MDD training dataset (GSE98793) and the independent validation dataset (GSE52790) resulted in a final set of 70 candidate genes, which were used as input features for subsequent machine learning modeling (Fig. 4).

Fig. 4.

Fig. 4

Identification and filtering of shared genes for GERD-MDD analysis. A Venn diagram showing the overlap between GERD differentially expressed genes (DEGs) and MDD-associated genes derived from WGCNA modules, yielding 78 shared genes. B Venn diagram showing the further restriction of these 78 shared genes to genes present in both the MDD training cohort (GSE98793) and the independent validation cohort (GSE52790), resulting in a final set of 70 candidate genes used for downstream machine learning analysis

Model selection and identification of the optimal classifier

A total of 71 machine-learning models, including single and combined algorithms, were evaluated in the training cohort using stratified 10-fold cross-validation. Model selection was based primarily on cross-validated AUC in the training dataset, whereas Fig. 5A additionally displays the AUC in the independent test dataset and the arithmetic mean of the two AUC values as descriptive comparison metrics. Among all candidate models, ridge regression (RR) achieved the highest cross-validated AUC (0.709) and was therefore selected as the optimal model for subsequent analyses.

Fig. 5.

Fig. 5

Performance comparison of candidate machine learning models and coefficient distribution of the optimal ridge regression model. A Model-comparison matrix summarizing the performance of 71 machine learning models. Model development involved stratified 10-fold cross-validation within the training dataset, and AUC values are shown for the training dataset and the independent test dataset. B Distribution of coefficients derived from the optimal ridge regression model. Because ridge regression retains all predictors with non-zero coefficients, coefficient magnitude reflects relative contribution within the fitted model rather than formal feature selection. Dashed lines are shown as visual reference markers and do not indicate a predefined feature-selection cutoff

Importantly, the AUC of 0.709 reflects the out-of-fold performance during cross-validation, providing an internal estimate of generalization ability, and was used strictly for model selection. To further characterize the structure of the selected model, we examined the distribution of gene coefficients derived from the RR model (Fig. 5B). Because ridge regression retains all predictors with non-zero coefficients, coefficient magnitude should be interpreted as reflecting relative contribution within the fitted model rather than formal feature selection. The dashed lines are shown as visual reference markers and do not represent a formal feature-selection cutoff. Unlike sparse models that perform hard feature selection, ridge regression retains all input features with non-zero coefficients. The coefficients were broadly distributed around zero, with both positive and negative contributions, indicating that the model integrates information from multiple genes rather than relying on a small subset of dominant predictors. This global distribution provides a basis for ranking genes according to their relative importance in subsequent analyses.

3.6Performance evaluation and cross-cohort generalizability of the optimal model

The predictive performance of the optimal RR model was evaluated in both the training cohort and the independent external test cohort. During model selection, cross-validation yielded an AUC of 0.709 for the training data (Fig. 5A), reflecting the internal out-of-fold performance. After refitting the model on the full training dataset, the apparent AUC increased to 0.896 (95% CI, 0.851–0.941), while the external test cohort achieved an AUC of 0.875 (95% CI, 0.712–1.000) (Fig. 6A). In addition to AUC, classification performance was summarized using sensitivity, specificity, accuracy, error rate, positive predictive value (PPV), negative predictive value (NPV), and F1 score (Table S2). In the training cohort, the model achieved a sensitivity of 0.805, specificity of 0.844, accuracy of 0.818, PPV of 0.912, NPV of 0.684, and F1 score of 0.855. In both cohorts, predicted scores were significantly higher in MDD samples than in controls, indicating effective rank-based class separation (Fig. 6B). However, the score distributions also revealed an important limitation. Compared with the training cohort, predicted probabilities in the external test cohort were shifted toward uniformly higher values and showed a narrower dynamic range; this distributional shift is further illustrated by the density plots in Figure S1. When the training-derived cutoff was directly transferred to the external cohort, all test samples were classified as positive, yielding a sensitivity of 1.000 but a specificity of 0.000, with an accuracy of 0.455, PPV of 0.455, and F1 score of 0.625 (Table S2, Fig. 6C).

Fig. 6.

Fig. 6

Performance evaluation and top gene coefficients of the optimal RR model. A Receiver operating characteristic (ROC) curves of the optimal RR model in the training and independent test cohorts. B Distribution of predicted probabilities stratified by class in the training and test cohorts. In both datasets, predicted scores were significantly higher in MDD samples than in controls, demonstrating effective class separation. C Classification performance metrics of the RR model, including AUC, accuracy, sensitivity, specificity, precision, and F1 score, evaluated in the training and test cohorts. D Top-ranked gene coefficients derived from the optimal RR model. Genes are ordered by the absolute magnitude of their coefficients, reflecting their relative contribution to model prediction. Positive and negative coefficients indicate the direction of association with the predicted outcome

Under this condition, some threshold-dependent metrics, particularly NPV and the diagnostic odds ratio (DOR), were not stably estimable in the external test cohort. Taken together, these findings indicate that the RR model retained discriminative ranking performance across cohorts, but its absolute probability scale was not stable, resulting in limited transferability of a fixed classification threshold. In addition to performance evaluation, we further examined the contribution of individual genes to model prediction based on the coefficient structure of the RR model. Genes were ranked according to the absolute magnitude of their coefficients, and the top-ranked features are visualized in Fig. 6D. This analysis provides a more interpretable view of the model by highlighting genes with the strongest positive and negative effects on prediction, complementing the global coefficient distribution shown in Fig. 5B.

CIBERSORT analysis based on key genes

To further define the immune context of the three prioritized genes, we examined the correlations between SAMD14, PROC, and NRG1 expression and the relative abundance of CIBERSORT-estimated immune cell populations in the MDD training cohort. Spearman correlation analysis followed by Benjamini–Hochberg correction identified 12 significant gene–immune cell associations (adjusted p < 0.05; Fig. 7; Table S3-S4). Among these, all three genes showed significant negative correlations with T cells gamma delta, including SAMD14 (r = -0.295, adjusted p = 0.0021), PROC (r = -0.264, adjusted p = 0.0035), and NRG1 (r = -0.205, adjusted p = 0.0357). SAMD14 was additionally negatively correlated with Eosinophils (r = -0.275, adjusted p = 0.0025) and T cells CD4 memory resting (r = -0.275, adjusted p = 0.0025), while showing a positive correlation with T cells CD8 (r = 0.247, adjusted p = 0.0062). PROC showed a positive correlation with Macrophages M0 (r = 0.256, adjusted p = 0.0044) and a weaker positive correlation with T cells CD8 (r = 0.189, adjusted p = 0.0474) but was negatively correlated with T cells CD4 memory resting (r = -0.214, adjusted p = 0.0273). NRG1 was positively correlated with Monocytes (r = 0.201, adjusted p = 0.0378) and Macrophages M0 (r = 0.196, adjusted p = 0.0394), and negatively correlated with Eosinophils (r = -0.196, adjusted p = 0.0394). Overall, these results indicate that the three-gene signature is associated with selective immune-cell infiltration patterns, particularly involving gamma delta T cells, macrophage-related compartments, and selected lymphoid and myeloid subsets, thereby supporting an immune-related component in the shared molecular architecture of GERD–MDD comorbidity.

Fig. 7.

Fig. 7

Correlation of the three prioritized genes with CIBERSORT-estimated immune cell fractions in the MDD training cohort. Spearman correlation analysis between SAMD14, PROC, and NRG1 expression and the relative abundance of immune cell populations estimated by CIBERSORT in the MDD training cohort. Dot size represents the absolute value of the Spearman correlation coefficient (|r|). Dot color indicates the direction of correlation and the significance level based on Benjamini–Hochberg-adjusted p-values, with red denoting positive correlations and blue denoting negative correlations

Candidate drug identification based on key genes

Drug-signature enrichment analysis identified several compounds whose annotated gene sets overlapped with the three prioritized genes. PROC appeared in enriched signatures including Clozapine (P = 3.83 × 10⁻⁴), Norgestimate (P = 5.99 × 10⁻³), and Warfarin (P = 8.72 × 10⁻³). NRG1 overlapped with the largest number of enriched signatures, including Clozapine, 5-Fluorouracil, Glibenclamide, Danazol, Terbufos, and Paracetamol (all P < 0.01). Because this analysis was exploratory and based on a small gene set, these results should be interpreted as hypothesis-generating rather than as direct therapeutic recommendations (Fig. 8).

Fig. 8.

Fig. 8

The circos plot illustrates the relationships between significantly enriched drug signatures and overlapping genes. Each sector represents either a drug term or a gene. Links indicate the presence of shared genes between drug-associated gene sets and the final prioritized genes submitted to Enrichr. Colors are used to distinguish different categories and do not represent quantitative values

Clinical Characteristics of the GERD clinical cohort stratified by depressive symptom severity

The clinical cohort comprised 70 patients with GERD, who were stratified into three PHQ-9-defined groups: non-depressed groups (n = 39), mild-to-moderate depression groups (n = 19), and major depression groups (n = 12). Baseline characteristics are summarized in Table 2. No significant between-group differences were observed for age, body mass index, symptom duration, pulmonary function parameters (FEV1 and FEV1/FVC), hemoglobin level, or categorical variables including sex, smoking status, alcohol consumption, chest pain, heartburn, reflux type, hiatal hernia, and cardia function (all P > 0.05). In contrast, DeMeester scores differed significantly across the three groups (P < 0.001). Using the non-depressed group as the clinical reference, median DeMeester values were higher in the mild-to-moderate depression group and highest in the major depression group, indicating an increasing reflux burden with worsening depressive status. Similarly, the total number of reflux episodes over 24 hours also differed among groups (P = 0.04), with higher values observed in the more major depression strata. The prevalence of esophagitis also differed significantly across groups (P = 0.03). The proportion of patients recorded as having esophagitis was 48.7% (19/39) in the no-depressive-symptoms group, 73.7% (14/19) in the mild-to-moderate group, and 83.3% (10/12) in the major group; however, this pattern should be interpreted cautiously because the proportion of unknown endoscopic status differed across groups (Table 2).

Table 2.

Clinical characteristics of GERD patients stratified by depression severity

Variable non-depressed mild-to-moderate depression major depression P
N = 39 N = 19 N = 12
Age 62.00 (57.00, 69.00) 69.00 (58.00, 71.00) 66.00 (62.50, 71.00) 0.15
BMI 25.70 (23.14, 27.27) 25.00 (23.40, 28.40) 25.94 (20.60, 27.82) 0.7
Symptom Duration 36.00 (12.00, 108.00) 60.00 (24.00, 120.00) 30.00 (5.00, 120.00) 0.6
DeMeester 4.80 (1.50, 19.90) 72.60 (67.90, 170.20) 253.25 (209.95, 264.60) < 0.001
Total Reflux 30.00 (22.00, 53.00) 40.00 (28.00, 71.00) 70.00 (34.50, 90.50) 0.04
FEV1 2.36 (1.79, 2.80) 2.28 (1.88, 2.71) 2.32 (1.61, 2.51) 0.6
FEV1/FVC 79.19 (76.24, 85.10) 78.59 (74.76, 81.60) 78.59 (71.04, 86.12) 0.7
Hemoglobin 130.00 (118.00, 142.00) 127.00 (108.00, 135.00) 126.50 (116.50, 136.00) 0.4
Gender 0.7
 Man 17 (43.6%) 6 (31.6%) 5 (41.7%)
 Woman 22 (56.4%) 13 (68.4%) 7 (58.3%)
Smoking 0.6
 No 32 (82.1%) 17 (89.5%) 11 (91.7%)
 Yes 7 (17.9%) 2 (10.5%) 1 (8.3%)
Drinking Alcohol > 0.9
 No 34 (87.2%) 17 (89.5%) 11 (91.7%)
 Yes 5 (12.8%) 2 (10.5%) 1 (8.3%)
Chest Pain 0.5
 No 4 (10.3%) 2 (10.5%) 0 (0.0%)
 Yes 35 (89.7%) 17 (89.5%) 12 (100.0%)
Heartburn 0.2
 No 6 (15.4%) 2 (10.5%) 4 (33.3%)
 Yes 33 (84.6%) 17 (89.5%) 8 (66.7%)
Reflux Types 0.6
 Acid Reflux 33 (84.6%) 15 (78.9%) 12 (100.0%)
 Alkaline Reflux 2 (5.1%) 1 (5.3%) 0 (0.0%)
 Mixed Reflux 4 (10.3%) 3 (15.8%) 0 (0.0%)
Hiatal Hernia 0.5
 No 8 (20.5%) 1 (5.3%) 2 (16.7%)
 Yes 28 (71.8%) 17 (89.5%) 10 (83.3%)
 Unknown 3 (7.7%) 1 (5.3%) 0 (0.0%)
Cardia Function 0.3
 Normal 7 (17.9%) 1 (5.3%) 3 (25.0%)
 Relaxation 23 (59.0%) 16 (84.2%) 8 (66.7%)
 Unknown 9 (23.1%) 2 (10.5%) 1 (8.3%)
Esophageal Inflammation 0.03
 No 5 (12.8%) 4 (21.1%) 1 (8.3%)
 Yes 19 (48.7%) 14 (73.7%) 10 (83.3%)
 Unknown 15 (38.5%) 1 (5.3%) 1 (8.3%)

Expression of SAMD14, PROC, and NRG1 in relation to depression severity

To assess the clinical relevance of the identified gene signature, we examined the expression patterns of SAMD14, PROC, and NRG1 across the clinically defined depressive-symptom strata rather than treating the binary RR model as a formal multi-class classifier. The expression levels of SAMD14, PROC, and NRG1 differed significantly among the non-depressed, mild-to-moderate depressive symptoms, and major depressive symptoms groups (all P < 0.001) (Fig. 9A-C). The relative expression levels of all three genes were positively correlated with GERD-related parameters. Specifically, Spearman correlations with the DeMeester score were 0.68 for SAMD14, 0.70 for PROC, and 0.72 for NRG1, while correlations with total reflux episodes were 0.31, 0.21, and 0.33, respectively (Fig. 9D). Strong positive correlations were also observed among the three genes themselves. In contrast, correlations with other clinical indicators such as BMI and pulmonary function parameters (e.g., FEV1) were weak.

Fig. 9.

Fig. 9

Key gene expression and clinical trait correlations in GERD. A-C Relative expression levels of SAMD14, PROC, and NRG1 across non-depressed, mild-to-moderate depression, and major depression groups. D Module-trait correlation table showing relationships between gene expression levels and clinical parameters. SAMD14, PROC, and NRG1 exhibited strong positive correlations with each other and with GERD-related indices (including DeMeester score and total reflux episodes), but weak correlations with BMI and pulmonary function parameters

Discussion

Collectively, this study integrated multi-cohort transcriptomic data from GERD and MDD to explore molecular features shared across the two conditions and identified a three-gene signature comprising SAMD14, PROC, and NRG1. The findings provide a data-driven starting point for examining GERD-MDD overlap within the broader framework of the gut-brain axis, a bidirectional communication network linking the central nervous system and the gastrointestinal tract [15]. However, broad gut-brain axis concepts need to be translated into more specific and testable molecular hypotheses. Previous research has often been confined to single-disease contexts or hypothesis-driven analyses based on known pathways, which may overlook shared molecular patterns across digestive and psychiatric disorders [16, 17].

In contrast to some prior studies that reported no significant increase in depressive symptoms among GERD patients, our transcriptome-based analysis suggests that a subset of molecular features may be relevant to both GERD biology and MDD classification. This finding should be interpreted cautiously, because the present study does not prove that these genes define a true comorbid molecular subtype. Nevertheless, the three prioritized genes provide biologically plausible candidates for future investigation. NRG1 has been implicated in psychiatric disorders, including mood disorders and schizophrenia, and is involved in neuroplasticity, glial function, and neuro-immune crosstalk [18]. SAMD14 has been associated with cellular signaling and survival-related processes [19], which may be relevant to tissue response or inflammatory remodeling. PROC, although less studied in psychiatric or gastrointestinal contexts, is linked to endothelial function and coagulation-inflammation biology [20]. We therefore propose that these genes may participate in biological processes involving immune regulation, tissue remodeling, endothelial or barrier-related function, and neuroimmune signaling; however, this interpretation remains hypothesis-generating and requires functional validation [21]. Recent studies have also highlighted inflammatory cytokine modulation in recurrent GERD and inflammation–miRNA-related convergence across gut health and depressive disorders [22, 23].

The immune-cell correlation analysis provided additional context for the three prioritized genes. Associations with gamma delta T cells, macrophage-related compartments, and selected lymphoid and myeloid subsets suggest that immune-related processes may contribute to the molecular overlap observed between GERD and MDD. The discriminatory performance of the model in an external MDD dataset, together with the directionally consistent associations between gene expression and MDD status, supports the potential relevance of these shared molecular features. Importantly, however, the current model should be interpreted within its appropriate scope. Although the input features were constrained to genes shared between GERD and MDD, model training and validation were performed only within MDD cohorts. Therefore, the present framework identifies GERD-informed molecular features relevant to MDD classification, rather than constituting a formal predictive model of GERD-MDD comorbidity. This distinction is important when interpreting the translational scope of the model.

The clinical cohort provided further context for the relationship between depressive symptom burden and reflux-related severity. DeMeester scores and total 24-hour reflux episodes differed across PHQ-9-defined groups, with higher values observed in patients with more severe depressive symptoms. The prevalence of esophagitis also differed among groups. These findings are consistent with epidemiological evidence linking GERD and depressive symptoms [3]. However, because the clinical analysis was cross-sectional, it cannot establish whether depressive symptoms exacerbate reflux, whether more severe reflux contributes to depressive symptoms, or whether both are influenced by shared biological or behavioral factors.

This study has several limitations. First, all transcriptomic discovery data were derived from public GEO datasets, and inherent heterogeneity and potential confounding remain unavoidable despite preprocessing and external validation. Second, the available GERD and MDD datasets differed substantially in biological source and technical platform: the GERD datasets were generated from esophageal tissue RNA-seq, whereas the MDD datasets were based on peripheral blood microarray profiles. For this reason, we did not perform a direct meta-analysis combining GERD and MDD expression matrices for model fitting or feature ranking, because such an approach could be strongly confounded by tissue-specific transcriptional programs and platform-related differences. Instead, we adopted a conservative intersection-based strategy to identify genes supported by both disease contexts. Future studies using harmonized datasets from comparable biological materials and platforms will be needed to test whether the present three-gene signature remains stable under formal meta-analysis-based approaches. Third, the machine learning framework was developed and validated only in MDD cohorts. Although candidate genes were selected from the molecular overlap between GERD and MDD, the model itself was not trained to distinguish patients with confirmed GERD-MDD comorbidity from appropriate comparator groups. Therefore, the current framework should be interpreted as a GERD-informed MDD classification model based on shared molecular features, rather than as a formal predictive model of GERD-MDD comorbidity. In addition, the observed shift in predicted probability distributions between the training and external test cohorts suggests that, although rank-based discrimination was preserved, the fixed classification threshold showed limited transferability across datasets. Fourth, while peripheral blood transcriptomics is advantageous for non-invasive biomarker development, it cannot directly capture molecular events within the central nervous system or esophageal tissue. Fifth, the clinical validation cohort was single-center and relatively small, and depressive symptom severity was assessed using the PHQ-9 rather than a structured psychiatric diagnostic interview. Missing endoscopic information in a subset of patients also limits interpretation of the esophagitis analysis. Future work will require prospective clinical cohorts with confirmed dual-diagnosis labels, harmonized tissue or blood transcriptomic profiling, and finer phenotyping, including depression subtypes, symptom dimensions, reflux monitoring parameters, and treatment response.

Finally, the present study remains primarily computational and associative. Functional validation of the prioritized genes in experimental systems, such as esophageal epithelial models, immune-cell models, organoid systems, or animal models, will be necessary to move from computational association toward mechanistic understanding.

Conclusion

In summary, this study explored molecular features shared between GERD and MDD by combining multi-cohort transcriptomic analysis, network modeling, and machine learning. We identified a three-gene signature (SAMD14, PROC, and NRG1) derived from the overlap between GERD-related differential expression and MDD-associated co-expression modules. This signature showed discriminative performance for MDD classification and was associated with reflux-related clinical indices in an independent GERD cohort. These findings provide a hypothesis-generating molecular entry point for future investigation of GERD-MDD overlap, but they should not yet be interpreted as a clinically validated diagnostic tool or as evidence of causality. Prospective studies with confirmed comorbid phenotypes and mechanistic validation are needed before translational application can be considered.

Supplementary Information

12876_2026_5008_MOESM1_ESM.xlsx (14.8KB, xlsx)

Supplementary Material 1: Table S1. Model specifications and parameter labels for the 71 candidate machine-learning models.

12876_2026_5008_MOESM2_ESM.xlsx (8.6KB, xlsx)

Supplementary Material 2: Table S2. Threshold-independent and threshold-dependent performance metrics of the optimal ridge regression model in the training and independent external test cohorts.

12876_2026_5008_MOESM3_ESM.xlsx (10.3KB, xlsx)

Supplementary Material 3: Table S3. Significant correlations between prioritized gene expression and CIBERSORT-estimated immune cell fractions in the MDD training cohort.

12876_2026_5008_MOESM4_ESM.xlsx (14KB, xlsx)

Supplementary Material 4: Table S4. Full Spearman correlation results between prioritized gene expression and immune cell fractions in the MDD training cohort, including Benjamini–Hochberg-adjusted P values.

12876_2026_5008_MOESM5_ESM.png (65KB, png)

Supplementary Material 5: Figure S1. Density distributions of model-predicted probabilities in the training and independent test cohorts. Relative to the training cohort, predicted probabilities in the external test cohort are shifted toward higher values and are more concentrated within the upper probability range, consistent with the limited transferability of a fixed training-derived classification cutoff across cohorts.

Supplementary Material 7. (157.3KB, png)
Supplementary Material 8. (128.7KB, png)
Supplementary Material 9. (221.8KB, png)

Acknowledgements

Not applicable.

Authors' contributions

PZ and KX were responsible for the overall conception, supervision, and coordination of the study. ZRW, YJC, and SDP performed the main data analysis and contributed to interpretation of the results. DT, YXZ, and ZHLi contributed to the study design and clinical data organization. ZRW, YJC, SDP, DT, YXZ, ZHLiu, HYZ, JYL, CW, ZHLi, YK, YNL, NZ, and YT contributed to manuscript drafting and revision. All authors read and approved the final manuscript.

Funding

None.

Data availability

The datasets generated and analyzed during the current study are not publicly available due to patient confidentiality and ethical restrictions but are available from the corresponding author on reasonable request.

Declarations

Ethics approval and consent to participate

Publicly available, de-identified GEO datasets were used for the transcriptomic discovery phase and therefore did not require additional institutional ethics approval. The clinical validation component of this study was approved by the Medical Ethics Committee of Tianjin Medical University General Hospital (IRB2025-KY-686). The study was conducted in accordance with the Declaration of Helsinki and Good Clinical Practice guidelines. Written informed consent to participate was obtained from all participants prior to enrollment and blood collection.

Consent for publication

Written informed consent for the procedure and management was obtained from all individual participants included in the study.

Competing interests

The authors declare no competing interests.

Footnotes

Publisher’s note

Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.

Zirong Wang, Yongjia Chai and Sida Pi contributed equally to this work as co-first authors.

Contributor Information

Kai Xiong, Email: xiongkai1229@tmu.edu.cn.

Peng Zhang, Email: pengzhang01@tmu.edu.cn.

References

  • 1.Clarrett DM, Hachem C. Gastroesophageal Reflux Disease (GERD). Mo Med. 2018;115(3):214–8. [PMC free article] [PubMed] [Google Scholar]
  • 2.Gyawali CP, Yadlapati R, Fass R, et al. Updates to the modern diagnosis of GERD: Lyon consensus 2.0. Gut. 2024;73(2):361–71. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 3.You Z-H, Perng C-L, Hu L-Y, et al. Risk of psychiatric disorders following gastroesophageal reflux disease: a nationwide population-based cohort study. Eur J Intern Med. 2015;26(7):534–9. [DOI] [PubMed] [Google Scholar]
  • 4.Zheng X, Zhou X Tong L, et al. Mendelian randomization study of gastroesophageal reflux disease and major depression. PLoS ONE. 2023;18(9):e0291086. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 5.Schramm E, Klein DN, Elsaesser M, et al. Review of dysthymia and persistent depressive disorder: history, correlates, and clinical implications. Lancet Psychiatry. 2020;7(9):801–12. [DOI] [PubMed] [Google Scholar]
  • 6.Friedrich MJ. Depression Is the Leading Cause of Disability Around the World. JAMA. 2017;317(15):1517. [DOI] [PubMed] [Google Scholar]
  • 7.Liu J, Li J, Yuan G, et al. Relationship between Sjogren’s syndrome and gastroesophageal reflux: A bidirectional Mendelian randomization study. Sci Rep. 2024;14(1):15400. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 8.Yan Z, Xu Y, Li K, et al. Genetic correlation between smoking behavior and gastroesophageal reflux disease: insights from integrative multi-omics data. BMC Genomics. 2024;25(1):642. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 9.Liu B, Chen M, You J, et al. The Causal Relationship Between Gastroesophageal Reflux Disease and Chronic Obstructive Pulmonary Disease: A Bidirectional Two-Sample Mendelian Randomization Study. Int J Chron Obstruct Pulmon Dis. 2024;19:87–95. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 10.Kessing BF, Bredenoord AJ, Saleh CMG, et al. Effects of anxiety and depression in patients with gastroesophageal reflux disease. Clin Gastroenterol Hepatol. 2015;13(6):1089–95.e1. [DOI] [PubMed]
  • 11.Shen Z, Bian Y, Huang Y, et al. Migraine and gastroesophageal reflux disease: Disentangling the complex connection with depression as a mediator. PLoS ONE. 2024;19(7):e0304370. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12.Mayer EA, Nance K. The Gut-Brain Axis. Annu Rev Med. 2021;73:439–53. [DOI] [PubMed] [Google Scholar]
  • 13.Asim M, Khan HS, Khan MH, et al. Association between gastroesophageal reflux disease severity and depression and anxiety in young adults: a cross-sectional study. Cureus. 2025;17(9):e92993. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 14.Kroenke K, et al. The PHQ-9: validity of a brief depression severity measure. J Gen Intern Med. 2001;16(9):606–13. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15.Yan M, Man S, Sun B, et al. Gut liver brain axis in diseases: the implications for therapeutic interventions. Signal Transduct Target Ther. 2023;8(1):443. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16.Yan M, Chen M, Jiang Q, et al. A network analysis of gene co-expression in post-mortem brain tissues identifying novel genes and biological pathways underlying major depression. Front Psychiatry. 2025;16:1556983. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 17.Zhou Y, Shi W. Identification of Immune-Associated Genes in Diagnosing Aortic Valve Calcification With Metabolic Syndrome by Integrated Bioinformatics Analysis and Machine Learning. Front Immunol. 2022;13:937886. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 18.Schosser A, Cohen-Woods S, Gaysina D, et al. NRG1 gene in recurrent major depression: no association in a large-scale case-control association study. Am J Med Genet B Neuropsychiatr Genet. 2010;153B(1):141–7. [DOI] [PubMed] [Google Scholar]
  • 19.Ray S, Chee L, Matson DR, et al. Sterile α-motif domain requirement for cellular signaling and survival. J Biol Chem. 2020;295(20):7113–25. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 20.Fu J, Ling J, Li C-F, et al. Nardilysin-regulated scission mechanism activates polo-like kinase 3 to suppress the development of pancreatic cancer. Nat Commun. 2024;15(1):3149. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 21.DuU S, Zhang L, Chen Y, et al. Exploring the Mechanisms of Gastroesophageal Reflux Disease Based on the Brain-Gut Axis Theory. Int J Gen Med. 2025;18:6833–45. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22.Wen T-K, Shiue S-J, Huang Y-J, et al. Scalp and auricular acupuncture attenuate recurrent gastroesophageal reflux disease and related inflammatory cytokines. J Tradit Complement Med. 2024;15(7):794–802. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 23.Miquel-Rio L, Jericó-Escolar J, Yanes-Castilla C, et al. A molecular convergence in the triad of parkinson’s disease, depressive disorder and gut health is revealed by the inflammation-miRNA axis. J Neuroinflammation. 2025;22(1):260. [DOI] [PMC free article] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

12876_2026_5008_MOESM1_ESM.xlsx (14.8KB, xlsx)

Supplementary Material 1: Table S1. Model specifications and parameter labels for the 71 candidate machine-learning models.

12876_2026_5008_MOESM2_ESM.xlsx (8.6KB, xlsx)

Supplementary Material 2: Table S2. Threshold-independent and threshold-dependent performance metrics of the optimal ridge regression model in the training and independent external test cohorts.

12876_2026_5008_MOESM3_ESM.xlsx (10.3KB, xlsx)

Supplementary Material 3: Table S3. Significant correlations between prioritized gene expression and CIBERSORT-estimated immune cell fractions in the MDD training cohort.

12876_2026_5008_MOESM4_ESM.xlsx (14KB, xlsx)

Supplementary Material 4: Table S4. Full Spearman correlation results between prioritized gene expression and immune cell fractions in the MDD training cohort, including Benjamini–Hochberg-adjusted P values.

12876_2026_5008_MOESM5_ESM.png (65KB, png)

Supplementary Material 5: Figure S1. Density distributions of model-predicted probabilities in the training and independent test cohorts. Relative to the training cohort, predicted probabilities in the external test cohort are shifted toward higher values and are more concentrated within the upper probability range, consistent with the limited transferability of a fixed training-derived classification cutoff across cohorts.

Supplementary Material 7. (157.3KB, png)
Supplementary Material 8. (128.7KB, png)
Supplementary Material 9. (221.8KB, png)

Data Availability Statement

The datasets generated and analyzed during the current study are not publicly available due to patient confidentiality and ethical restrictions but are available from the corresponding author on reasonable request.


Articles from BMC Gastroenterology are provided here courtesy of BMC

RESOURCES