Abstract
It is well known that bipolar disorder (BD) and epilepsy (EP) are common neurological diseases. The objective of this study was to screen for potential biomarkers applicable to the diagnosis of EP and BD. The gene expression profiles from both the BD and EP datasets were sourced from the Gene Expression Omnibus database. To pinpoint the core shared genes, we conducted differential expression analysis as well as weighted gene co-expression network analysis. Additionally, we leveraged protein–protein interaction, Gene Ontology, and Kyoto Encyclopedia of Genes and Genomes pathway enrichment to uncover the pathogenic genes of BD and EP, as well as their underlying mechanisms. Using least absolute shrinkage and selection operator regression, support vector machine–recursive feature elimination, and random forest, hub genes were determined via rigorous examination. Subsequently, predictive nomograms and receiver operating characteristic curves were crafted to forecast BD and EP. A single-gene set enrichment analysis was executed meticulously on every diagnostic gene, aiming to identify shared signaling pathways. To round things off, the cell-type identification by estimating relative subsets of RNA transcripts algorithm analysis explored immune cell infiltration within BD and EP samples. After analyzing the intersection of weighted gene co-expression network analysis significant module genes and the differentially expressed genes, we pinpointed 113 genes of interest. Our protein–protein interaction analysis revealed 3 pivotal modules, each harboring 14 genes, which are considered pivotal for diagnosing BD and EP. The machine learning models consistently highlighted 2 genes – Regulators of G-protein signaling 4 and gamma-aminobutyric acid type A receptor subunit alpha1 – as universal diagnostic biomarkers. Furthermore, the immune infiltration analysis disclosed that activated M2 macrophages and mast cells are integral players in the onset of BD and EP.
Keywords: biomarker, bipolar disorder, epilepsy, machine learning, transcriptomic analysis
1. Introduction
Bipolar disorder (BD), often simply referred to as BD, is a multifaceted set of serious long-term conditions. This encompasses bipolar I and bipolar II disorders,[1] impacting roughly 2% of the world inhabitants.[2] It marked by alternating periods of intense mania and severe depression, leading to substantial health issues and an increased risk of an early demise.[3] BD patients exhibit an approximate yearly suicide incidence of 0.9%, contrasting with 0.014% in the general populace.[4] Epilepsy (EP) involves spontaneous, recurring seizures. Seizures affect 10% of the global population,[5] and most focal seizures originate in the temporal region.[6] Some evidence suggests that inflammation in the brain may be a major contributor to acquired seizures.[7]
Although the clinical manifestations of the 2 diseases are different, more and more clinical evidence suggests that there are significant comorbidities between the 2 diseases.[8] Epidemiological findings indicate a 2- to 3-fold increased EP occurrence among BD patients, with BD prevalence in EP cases reaching 6% to 12%.[9] Some antidrugs (such as sodium valproate) are used to treat BD,[10] suggesting that the 2 diseases may have a common pathophysiological mechanism. Current research is limited to identify BD and EP risk genes in single diseases, but cross-disease analysis is only seen in theoretical discussions.
This study set out to identify promising biomarkers and the biological pathways they influence in the development of BD and EP. To this end, we pooled transcriptomic data on BD and EP from the Gene Expression Omnibus database. We then leveraged the “limma” package (The Walter and Eliza Hall Institute of Medical Research [WEHI], Parkville, Victoria, Australia) and weighted gene co-expression network analysis (WGCNA) to pinpoint genes with significantly altered expression levels and key functional modules specific to each disease. We then looked at the overlap and applied 3 machine learning techniques – least absolute shrinkage and selection operator (LASSO) regression, support vector machine–recursive feature elimination (SVM-RFE) algorithm, and random forest (RF) – to these genes. We identified diagnostic genes shared by 2 diseases, regulators of G-protein signaling 4 (RGS4) and gamma-aminobutyric acid type A receptor subunit alpha1 (GABRA1), which was verified through external datasets. Furthermore, gene set enrichment analysis (GSEA) was conducted on key genes to identify common pathways implicated in both BD and EP. What is more, we dug into how immune cells contribute to the diseases common origin with an immune infiltration analysis. To sum it up, this research delves into the common molecular roots of BD and EP, and it puts the spotlight on RGS4 and GABRA1 as promising diagnostic indicators for these conditions.
2. Methods
2.1. Materials
We obtained all transcriptomic datasets of BD and EP from the Gene Expression Omnibus database (https://www.ncbi.nlm.nih.gov/geo/), and thus no ethical approval was required for the use of these data.[11] For BD, we downloaded the gene expression omnibus series (GSE) 5389 dataset based on gene expression omnibus platform (GPL) 96, which included brain tissue from the orbitofrontal cortexes of 10 BD samples and 11 control samples. To evaluate diagnostic performance, the GSE5388 dataset from GPL96 was acquired, which included 30 BD samples and 31 control samples from the prefrontal cortexes. For EP, the GSE28674 dataset based on GPL6480 was downloaded, which included brain tissue from the hippocampus of 6 temporal lobe epilepsy samples and 12 control samples. To evaluate diagnostic accuracy, the GSE134697 dataset from GPL16791 was also acquired, which included hippocampal and neocortex samples from 17 temporal lobe epilepsy patients.
2.2. Identification and visualization of differentially expressed genes
To pinpoint the genes showing differential expression in the GSE5389 and GSE28674 datasets, we turned to the “limma” package in R.[12] For GSE5389, we homed in on genes with an adjusted P < .05 and |log FC| ≥ 0. The selection criteria for GSE28674 were slightly more stringent, requiring an adjusted P < .05 and |log FC| ≥ 1. To really drive home the differences in gene expression between BD and EP, we whipped up some waterfall plots using the ever-reliable “ggplot2” package (Posit Software, PBC, Boston) in R.
2.3. WGCNA
WGCNA is a systematic biology that describes correlation patterns between genes in microarray samples.[13] It can be used to find gene clusters (modules) that are related to the degree. We employed the R package “WGCNA” (University of California, Los Angeles [UCLA], Los Angeles) for gene co-expression network assembly and utilized the “Hclust” function in R for outlier sample identification through hierarchical clustering. To pinpoint the optimal soft-thresholding power, we leveraged the “pickSoftThreshold” function within the WGCNA package. Employing a power function (β), the expression profile similarity matrix became an adjacency network, then a topological overlap map. Module eigengene was then used to summarize the expression patterns of each module and calculate and streamline the association between module eigengene and clinical characteristics for further analysis.
2.4. The protein–protein interaction (PPI) network construction
By plotting the Venn diagram, we conducted a combined analysis of the differentially expressed genes (DEGs) and WGCNA screening genes. These so-called overlapping genes, which are essentially the building blocks of shared genetic information, were meticulously deposited in the Search Tool for the Retrieval of Interacting Genes/Proteins database (https://string-db.org/). We then went ahead and crafted a comprehensive PPI network for these fundamental genes.[14] To get a clearer picture of this intricate network, we utilized the Cytoscape software, version 3.10.3 (The Cytoscape Consortium, San Diego). With Cytoscape molecular complex detection plugin, we effectively sifted through the PPI network inferential module. The inference module uses default settings, degree truncation = 2, node score truncation = 0.2, K-core = 2, and maximum depth = 100.
2.5. Functional enrichment analysis
To figure out what these key genes actually *do* and which biological pathways they are involved in, we ran Gene Ontology (GO) and Kyoto Encyclopedia of Genes and Genomes (KEGG) pathway enrichment analyses using the “clusterProfiler” package (Southern Medical University, Guangzhou, China). We were essentially trying to pinpoint their common biological roles and the pathways they influence, zeroing in on those key genes we’d previously identified in the PPI network. Findings below P value < .05 indicated statistical significance.
2.6. Machine learning analysis
For the discovery of putative diagnostic genes linked to BD and EP, the following 3 machine learning techniques were applied: LASSO regression, SVM-RFE algorithm, and RF.[15] LASSO regression algorithm reduces estimates of square deviations and improves the accuracy of predictions.[16] SVM-RFE algorithm is a powerful tool for analyzing biomedical data and can achieve precision in predicting the intensity between factors.[17] In contrast, RF is a flexible prediction algorithm that can provide reliable or better predictions.[18] Then, we constructed Venn diagrams to determine the diagnostic genes.
2.7. Constructing nomograms and evaluating predictive models for diagnostic markers
To develop a diagnostic tool, we leveraged the “rms” package (Vanderbilt University Medical Center, Nashville) to whip up a nomogram[19] based on our identified diagnostic genes. This nomogram hinges on “points,” which reflect each gene specific metric. Ultimately, the “total points,” representing the grand sum of these individual scores, serves as the critical yardstick for diagnosing BD and EP. Moreover, the prognostic potential of the candidate genes and nomograms was appraised through receiver operating characteristic (ROC) curve analysis.[20] Area under the curve (AUC) of 0.7 is considered to have strong diagnostic accuracy. Ultimately, nomogram prediction accuracy in BD and concurrent EP was assessed via calibration curves.
2.8. GSEA
Having pinpointed the diagnostic gene, we then ran a single GSEA[21] on each diagnosis within both datasets, leveraging the “clusterProfiler” R package. GSEA, as you know, is a way of checking if certain gene sets are overrepresented. By using GSEA, we were able to compare the signaling pathways of these diagnostic genes, contrasting disease groups with the healthy control subjects to see what stood out. The molecular signature database was used in GSEA, and the study used signature genes.[22] Pathway analysis highlighted the top 5 activated and repressed signaling cascades for each diagnostic gene within both disease cohorts.
2.9. Immune cell infiltration analysis
To assess the immune landscape of both BD and EP, we employed the cell-type identification by estimating relative subsets of RNA transcripts (CIBERSORT) R package (Stanford University, Stanford)[23] to deconvolve the gene expression profiles, hence determining the infiltration levels of 22 diverse immune cell categories.[24] CIBERSORT output, representing immune cell proportions, was normalized to sum to 1 for each sample, allowing for a fair comparison across different immune cell populations and datasets. We then leveraged the R packages “corplot,” “viplot,” and “ggplot2” to bring these findings to life through visual representations. Ultimately, we investigated the association between our identified diagnostic genes and the density of immune cells using Spearman rank correlation test to see how they stacked up against each other.
3. Results
3.1. Identification of DEGs
In the BD dataset, GSE5389, a grand total of 2551 DEGs were pinpointed, with 1481 being upregulated and a notable 1080 being downregulated. Similarly, the EP dataset, GSE28674, unearthed 980 DEGs, where 750 were upregulated and 230 went the other way. To visualize the expression patterns of these genes across the 2 conditions, we employed waterfall plots, as illustrated in Figure 1A and B.
Figure 1.
Identification of in BD and EP. (A) The waterfall plot revealing DEGs in the BD dataset. (B) The waterfall plot revealing DEGs in the EP dataset. BD = bipolar disorder, DEGs = differentially expressed genes, EP = epilepsy.
3.2. Establishment of a weighted gene co-expression network and subsequent delineation of pivotal modules
To delve deeper into the crucial genes implicated in both conditions, WGCNA was employed to pinpoint interconnected gene modules within the 2 cohorts. In order to establish a scale-free network, the scale-free fitting index and average connectivity were computed. A soft-thresholding power of 8 was chosen for the BD group, based on the average connectivity across varying individual sizes. By using adjacency functions, an adjacency matrix was formed (Fig. 2A). Hierarchical clusters were constructed using the topological overlap map dissimilarity measure (Fig. 2C). From the co-expression network built using the BD dataset, 15 distinct modules were identified and selected brown, pink, and red as key modules, among which| r| ≥ 0.5 and P < .05. For EP, a soft-thresholding power of 9 was chosen. We identified the pink module as the key module, among which |correlation| ≥ 0.5 and P < .05 (Fig. 2B and D). In addition, we further intersected the DEGs and genes from the key modules of WGCNA in BD and EP and obtained a total of 113 core shared genes, which were further analyzed (Fig. 2E).
Figure 2.
Weighted gene co-expression network analysis (WGCNA) of BD and EP. (A) Cluster dendrogram of BD highly connected genes in key modules. (B) Relationships between modules and traits in BD. Correlations and P values are included in each cell. (C) Cluster dendrogram of EP highly connected genes in key modules. (D) Relationships between modules and traits in EP. Correlations and P values are included in each cell. (E) Venn diagram shows that 113 genes overlap in the DEGs and the key modules of WGCNA from BD and EP. BD = bipolar disorder, DEGs = differentially expressed genes, EP = epilepsy.
3.3. PPI network construction and key gene identification
Following this, we uploaded 113 genes to the Search Tool for the Retrieval of Interacting Genes/Proteins database and constructed a PPI) network. Figure 3A displays a network of 54 genes (nodes) with 124 interaction edges. To pinpoint key modules within the PPI network, we employed the molecular complex detection plugin in Cytoscape (Fig. 3B), using the following parameters: a degree cutoff of 2, a node score cutoff of 0.2, a K-core of 2, and a maximum depth of 100. Ultimately, we homed in on 14 key genes within the identified module for subsequent analysis.
Figure 3.
PPI network and functional enrichment between BD and EP. (A) The PPI network of 113 genes. (B) Key modules of the PPI network. (C and D) Key genes were represented by bar plots displaying GO and KEGG enrichment. BD = bipolar disorder, EP = epilepsy, GO = Gene Ontology, KEGG = Kyoto Encyclopedia of Genes and Genomes, PPI = protein–protein interaction.
3.4. Functional enrichment
To pinpoint shared regulatory pathways, we ran GO and KEGG enrichment analyses on those 14 key genes. Looking at the GO annotations, we broke things down into 3 categories: cellular component, biological process, and molecular function. This let us really dig into the functional enrichment of these key genes and see what makes them tick. For cellular component, the key genes were chiefly associated with synaptic membrane and gamma-aminobutyric acid A receptor (GABAAR) complex, etc. The biological process , the key genes might be related to inhibitory synapse assembly and GABA signaling pathway, etc. Lastly, for molecular function, the key genes were primarily interested in GABAAR activity, etc. KEGG pathway enrichment indicated primary involvement of key genes in neuroactive ligand-receptor communication and GABAergic synapses (Fig. 3C and D).
3.5. Identification of candidate hub genes through multiple machine-learning methods
To pinpoint key genes for BD and EP diagnosis, LASSO regression, SVM-RFE, and RF algorithms were employed. Initially, using the LASSO regression algorithm, we identified 4 genes in GSE5389 at the most appropriate λ = 0.08946765 (Fig. 4A and B). In GSE28674, we successfully identified 6 genes at the most appropriate λ = 0.00926975 (Fig. 5A and B).
Figure 4.
Three machine learning methods identify hub genes associated with bipolar disorder. (A and B) LASSO regression was used to screen and identify 4 key genes. (C and D) Five crosstalk genes were selected by using the SVM-RFE algorithm for bipolar disorder. (E and F) Random forest ranks all key genes to determine their importance in the model. LASSO = least absolute shrinkage and selection operator, SVM-RFE = support vector machine–recursive feature elimination.
Figure 5.
Three machine learning methods identify hub genes associated with epilepsy. (A and B) LASSO regression was used to screen and identify 6 key genes. (C and D) Five crosstalk genes were selected by using the SVM-RFE algorithm for bipolar disorder. (E and F) Random forest ranks all key genes to determine their importance in the model. LASSO = least absolute shrinkage and selection operator, SVM-RFE = support vector machine–recursive feature elimination.
Following the application of the SVM-RFE algorithm, we meticulously tested all the genes using a 5-fold cross-validation method. After analyzing the average rankings, we pinpointed the 9 most crucial genes associated with GSE5389, as depicted in Figure 4C and D, and similarly, we identified the top 9 pivotal genes linked to GSE28674, as shown in Figure 5C and D. We also implemented the RF algorithm to rank 14 genes based on their variable importance and extracted genes with importance > 0.7. In GSE5389 we obtained 5 genes (Fig. 4E and F) and in GSE28674 we obtained 6 genes (Fig. 5E and F). Finally, we constructed a Venn diagram to identify overlapping genes from 3 learning methods, resulting in 6 hub genes significantly associated with BD (Fig. 6A) and 8 hub genes significantly associated with EP (Fig. 6B).
Figure 6.
Selection and validation of the 2 shared diagnostic genes. (A) The Venn diagram showed 6 hub genes in BD by intersecting the results of 3 algorithms. (B) The Venn diagram showed 8 hub genes in EP by intersecting the results of 3 algorithms. (C) The Venn plot showed the 2 shared diagnostic genes. BD = bipolar disorder, EP = epilepsy, LASSO = least absolute shrinkage and selection operator, SVM-RFE = support vector machine–recursive feature elimination.
3.6. Diagnostic value and validation of diagnostic biomarkers
In order to understand the relationship between BD and EP more accurately, we intersected the hub genes of BD and EP and obtained 2 shared diagnostic genes, RGS4 and GABRA1 (Fig. 6C). To get a clearer picture in terms of diagnosis and forecasting, we came up with a nomogram that utilizes the information from those 2 diagnostic genes, thanks to some solid logistic regression analysis (Fig. 7A and D). Furthermore, we also mapped out the ROC curves, to get a better sense of the outcomes. Right on cue, both diagnostic genes delivered AUC values north of 0.75. Furthermore, the nomogram AUC values mirrored those of each individual diagnostic gene, suggesting it could be a real game-changer in diagnosing BD and EP (Fig. 7B and E). Figure 7C and F illustrate that the diagnostic model predictive accuracy closely mirrors that of an ideal model.
Figure 7.
Development of the diagnostic nomogram model. (A) The nomogram was constructed for BD based on the diagnostic biomarkers. (B) The ROC curve for the diagnostic performance of each candidate biomarker including GABRA1, RGS4, and the nomogram model constructed for BD. (C) The calibration curve of nomogram model prediction in BD. The dash line is marked as “ideal,” which represents the standard curve, and is on behalf of the perfect prediction of the ideal model. The dotted line is marked as “apparent,” which indicates the uncalibrated prediction curve, while the solid line is marked as “bias-corrected” and represents the calibrated prediction curve. (D) The nomogram was constructed for EP based on the diagnostic biomarkers. (E) The ROC curve for the diagnostic performance of each candidate biomarker including GABRA1, RGS4, and the nomogram model constructed for EP. (F) The calibration curve of nomogram model prediction in EP. BD = bipolar disorder, EP = epilepsy, GABRA1 = gamma-aminobutyric acid type A receptor subunit alpha1, RGS4 = regulators of G-protein signaling 4, ROC = receiver operating characteristic.
Furthermore, to bolster the credibility of the previously detailed bioinformatics study, we devised a nomogram using data from 2 diagnostic markers within the validation dataset (Fig. 8A and D). Within the BD validation dataset, GSE5388, based on the ROC curve, it can be observed that the maximum AUC of the nomogram between the control group and BD is comparable to that of each marker (Fig. 8B and C). Similarly, in the EP validation dataset GSE134697, the nomograms showed AUC values for each diagnostic gene, bolstering the hypothesis that the nomogram possesses substantial diagnostic utility in EP (Fig. 8E and F). The results show that these results confirm the capabilities of RGS4 and GABRA1 as key diagnostic genes for BD and EP.
Figure 8.
Validation of the expression patterns of 2 hub genes in BD and EP samples. (A) The nomogram was constructed for BD validation dataset based on the diagnostic biomarkers. (B) The ROC curve for the diagnostic performance of each candidate biomarker including RGS4, GABRA1, and the nomogram model constructed for BD validation dataset. (C). The calibration curve of nomogram model prediction in BD validation dataset. (D). The nomogram was constructed for EP validation dataset based on the diagnostic biomarkers. (E). The ROC curve for the diagnostic performance of each candidate biomarker including RGS4, and GABRA1, and the nomogram model constructed for EP validation dataset. (F). The calibration curve of nomogram model prediction in EP validation dataset. BD = bipolar disorder, EP = epilepsy, GABRA1 = gamma-aminobutyric acid type A receptor subunit alpha1, RGS4 = regulators of G-protein signaling 4, ROC = receiver operating characteristic.
3.7. Single-gene GSEA analysis of diagnostic genes
To delve deeper into the roles of these 2 diagnostic genes, we embarked on a single-gene GSEA across both the BD and EP datasets. Utilizing the “GSEA” software package (Broad Institute, Inc., Cambridge), we mapped out 5 key pathways – showing both up-regulation and down-regulation (Fig. 9A–D). Across the board in both disease cohorts, these genes were integral to the cation channel complex, GABAergic synapses, and neuron-to-neuron synapse.
Figure 9.
GSEA for the single diagnostic gene. (A) GSEA analysis for GABRA1 in BD group. (B) GSEA analysis for GABRA1 in EP group. (C) GSEA analysis for RGS4 in BD group. (D) GSEA analysis for RGS4 in EP group. BD = bipolar disorder, EP = epilepsy, GABRA1 = gamma-aminobutyric acid type A receptor subunit alpha1, GSEA = gene set enrichment analysis, RGS4 = regulators of G-protein signaling 4.
3.8. Immune cell infiltration and its correlation with candidate biomarkers
In this study, we delved into the immunological intricacies of both BD and EP by deploying the CIBERSORT algorithm to identify distinct immune cell traits. We scrutinized the interplay between immunoregulation, diagnostic markers, and the immune cell influx, with visualizations in Figure 10A and D providing insightful comparisons. What emerged was a striking contrast to the control group, where the BD samples showcased elevated levels of naive B cells, follicular helper T cells, M2 macrophages, monocytes, activated mast cells, dendritic cells, and T cells regulatory, while simultaneously displaying reduced cluster of differentiation 8 T cell counts (Fig. 10B).
Figure 10.
Immune cell infiltration analysis in BD and EP. (A) A stacked histogram displaying the immune cell proportions between BD and control groups. (B) Bar charts show the relative proportions of different immune cell types in BD and control groups. (C) The correlation map representing the association of the differentially infiltrated immune cells with 2 hub genes in BD. (D) A stacked histogram displaying the immune cell proportions between EP and control groups. (E) Bar charts show the relative proportions of different immune cell types in EP and control groups. (F) The correlation map representing the association of the differentially infiltrated immune cells with 2 hub genes in EP. BD = bipolar disorder, EP = epilepsy. * P < .05, ** P < .01, *** P < .001.
In comparison to the control group, the EP datasets showcased a higher prevalence of quiescent M2 macrophages and dendritic cells, while exhibiting lower numbers of activated mast cells, natural killer (NK) cells, Plasma cells, and resting cluster of differentiation 4 memory cells (Fig. 10E). Our research also revealed that both the BD and EP datasets presented elevated counts of M2 macrophages and activated mast cells. Furthermore, our analysis of immune cell correlations with potential biomarkers indicated a positive association between activated NK cells and RGS4, along with neutrophils, and a negative correlation between activated NK cells and RGS4, B cells naive and GABRA1 (Fig. 10C and F).
4. Discussion
Advancements in microarray and sequencing methods have occurred lately,[25] it has become possible to obtain large-scale gene expression data, and machine learning algorithms[26] have significant advantages in complex data analysis and pattern recognition. By leveraging comprehensive bioinformatics analyses in conjunction with machine learning techniques, we were able to identify genes that contribute to both Behçet disease and EP. This work lays the groundwork for more targeted treatments and a deeper understanding of the underlying disease mechanisms in both conditions.
This study employs diverse bioinformatics tools to identify novel biomarkers and pathways associated with the development of BD and EP. We speculate that synapse assembly, GABAergic synapse,[27] and monoatomic ion channel activity might be potential mechanisms of BD and EP. A diagnostic Nomogram model has also been developed using 2 diagnostic genes, RGS4 and GABRA1, to predict the risk of BD and EP. Our findings suggest that these diagnostic genes are excellent predictors for BD and EP, as demonstrated by ROC curve analysis. Furthermore, external validation in our patient group confirmed that a diagnostic nomogram, relying on RGS4 and GABRA1 levels, effectively differentiated between BD and EP; it really hit the mark. Digging deeper, immune infiltration analysis highlighted the significant roles of M2 macrophages and activated mast cells in the development of both BD and EP; they seem to be key players in the disease process.
The results of this study suggest that nuclear cross-dialogue genes in BD and EP are related to the cation channel complex, GABAergic synapse and neuron-to-neuron synapse pathways. The cation channel complex pathway may lead to the abnormal membrane potential of neurons through influence, triggering hyperactivity of the isolated channel, leading to abnormal synchronous discharge of neurons, causing EP.[28] In addition, the positive and dissociative channel complex also affects the mitochondrial potential, resulting in changes in the excitability of the limbic system loop, causing emotional fluctuations, in bipolar affective disorder.[29] The GABAergic synapse pathway may cause oversynchronization of local neural networks by affecting the functional defects of GABAergic interneurons, causing seizures,[30] and GABAergic synapses may cause an imbalance of GABAergic regulation in the prefrontal lobe and amygdala, causing disorder of emotion regulation and triggering BD.[31] The neuron-to-neuron synapse pathway may lead to changes in the efficiency of contact transmission, leading to the formation of abnormal circuits (e.g., hippocampal sclerosis), resulting in EP,[32] similarly, abnormalities in neuron-to-neuron synapses may lead to impaired synaptic plasticity, dysfunction of the prefrontal-striatal circuit leading to BD.[33] Three pathways may interact with each other, firstly, the mutation of the cation channel complex causes changes in the basic excitability of neurons, and the GABAergic synapse pathway attempts to balance overexcitement, but fails due to its flaws (such as GABRA1 mutation), which causes abnormal amplification of synaptic connections and disorder of local circuits,[34] ultimately cause emotional symptoms RGS4 of BD and abnormal discharges of EP.
RGS4 was regulatory molecules that act as GTPase activating proteins for G alpha subunits of heterotrimeric G proteins.[35] It pivotal in rendering the Gi alpha, Go alpha, and Gq alpha subtypes of G protein subunits inactive. And elevated levels of RGS4 are reportedly linked to a variety of human nervous system diseases, such as schizophrenia, Parkinson disease, and addiction.[36] In addition, in our study, it was identified as a new marker of BD and EP. RGS4 is located in brain regions such as the ventral tegmental area, nucleus accumbens, amygdala, prefrontal cortex, and hippocampus. These brain areas play a role in anxiety, depression, and drug reward.[37] RGS4 has been widely reported to perform its functions through several important biological processes, such as emotional regulation, addiction, and the transmission and maintenance of chronic pain.[38] This study also demonstrated that RGS4 is involved in many immune cells. In this regard, our research suggests that RGS4 may provide a potential diagnostic indicator for BD and EP.
In addition, this study identified GABRA1 as a potential contributor to the diagnosis of BD and EP. GABA serves as the primary inhibitory neurotransmitter within the mammalian brain, influencing GABAARs that function as chloride channels regulated by ligands. The GABRA1 gene plays a crucial role in the growth and development of the brain,[39] as well as in neurodevelopmental processes and related disorders. It is responsible for producing the α1 subunit, a key component found in high quantities and expressed throughout development. This subunit is integral to the GABAARs, which are pivotal for inhibitory functions in the brain.[40] Some studies suggest that when the metabolism of membrane lipids like sphingomyelin and phospholipids goes awry, it can mess with the structure of synaptic membranes and the activity of neurotransmitter receptors, specifically GABRA1, potentially triggering inflammation.[41] Furthermore, there is a growing body of research indicating that inflammation is a major player in how BD and EP develop.[42] So, to cut a long story short, GABRA1 is shaping up to be a pretty good biomarker for both BD and EP.
Our study has several strengths. First, we used transcriptomic analysis to understand the connection between BD and EP. To pinpoint possible diagnostic genes that bipolar disorder and epilepsy might share, we used LASSO, SVM-RFE, and RF algorithm. The predictive power gets a real boost when you validate against external datasets. Now, this study isn’t without its limitations, mind you. Since our results hinge on distinct patient groups, we couldn’t confirm them in patients dealing with multiple health issues at once. Looking ahead, we’ll need to construct a really solid model of both bipolar disorder and epilepsy to see if there is a real connection between them. Plus, we didn’t factor in things like patient age, sex, meds, or co-existing conditions in our sample, and that could throw a wrench into how dependable our current findings are.
5. Conclusion
In wrapping up, this research has broken new ground by pinpointing pivotal genes shared in brain tissue across those suffering from both bipolar disorder and epilepsy. RGS4 and GABRA1 stand out as the key communicators bridging the gap between BD and EP. These revelations offer fresh perspectives on the common disease processes that link these disorders and point toward exciting prospects for therapies aimed at effectively treating both BD and EP.
Acknowledgments
This study received no financial support. The authors declare no conflicts of interest. We thank the researchers who provided data from their original studies, and we sincerely appreciate Dr Ying Ji for his expert guidance in statistical methodology design and rigorous data analysis.
Author contributions
Conceptualization: Yixuan Zhang.
Data curation: Yixuan Zhang.
Formal analysis: Yixuan Zhang.
Investigation: Yixuan Zhang.
Methodology: Yixuan Zhang.
Validation: Yixuan Zhang.
Visualization: Yixuan Zhang, Youzhi Xu.
Writing - original draft: Yixuan Zhang.
Writing – review & editing: Yixuan Zhang, Chunyue Huo.
Abbreviations:
- AUC
- area under the curve
- BD
- bipolar disorder
- CIBERSORT
- cell-type identification by estimating relative subsets of RNA transcripts
- DEGs
- differentially expressed genes
- EP
- epilepsy
- GABAARs
- gamma-aminobutyric acid A receptors,
- GABRA1
- gamma-aminobutyric acid type A receptor subunit alpha1
- GO
- Gene Ontology
- GPL
- gene expression omnibus platform
- GSE
- gene expression omnibus series
- GSEA
- gene set enrichment analysis
- KEGG
- Kyoto Encyclopedia of Genes and Genomes
- LASSO
- least absolute shrinkage and selection operator
- NK
- natural killer
- PPI
- protein–protein interaction
- RF
- random forest
- RGS4
- regulators of G-protein signaling 4
- ROC
- receiver operating characteristic
- SVM-RFE
- support vector machine–recursive feature elimination
- WGCNA
- weighted gene co-expression network analysis.
The authors have no funding and conflicts of interest to disclose.
The datasets generated during and/or analyzed during the current study are publicly available.
How to cite this article: Zhang Y, Xu Y, Huo C. Transcriptomic analysis and machine learning have identified shared diagnostic genes and a possible mechanism linking bipolar disorder and epilepsy. Medicine 2026;105:10(e47980).
Contributor Information
Yixuan Zhang, Email: 3539582852@qq.com.
Youzhi Xu, Email: 2235255922@qq.com.
References
- [1].McIntyre RS, Berk M, Brietzke E, et al. Bipolar disorders. Lancet. 2020;396:1841–56. [DOI] [PubMed] [Google Scholar]
- [2].Goes FS. Diagnosis and management of bipolar disorders. BMJ. 2023;381:e073591. [DOI] [PubMed] [Google Scholar]
- [3].Lane NM, Smith DJ. Bipolar disorder: diagnosis, treatment and future directions. J R Coll Physicians Edinb. 2023;53:192–6. [DOI] [PubMed] [Google Scholar]
- [4].Nierenberg AA, Agustini B, Köhler-Forsberg O, et al. Diagnosis and treatment of bipolar disorder: a review. JAMA. 2023;330:1370–80. [DOI] [PubMed] [Google Scholar]
- [5].Falco-Walter J. Epilepsy-definition, classification, pathophysiology, andepidemiology. Semin Neurol. 2020;40:617–23. [DOI] [PubMed] [Google Scholar]
- [6].Vinti V, Dell’Isola GB, Tascini G, et al. Temporal lobe epilepsy and psychiatric comorbidity. Front Neurol. 2021;12:775781. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [7].Grassi P. Re: Cazana et al. A comparison of long-term results after Baerveldt 250 implantation in advanced uveitic vs. other forms of glaucoma. Graefes Arch Clin Exp Ophthalmol. 2018;261:2089–90. [DOI] [PubMed] [Google Scholar]
- [8].Li J, Ledoux-Hutchinson L, Toffa DH. Prevalence of bipolar symptoms or disorder in epilepsy: a systematic review and meta-analysis. Neurology. 2022;98:e1913–22. [DOI] [PubMed] [Google Scholar]
- [9].Lu E, Pyatka N, Burant CJ, Sajatovic M. Systematic literature review of psychiatric comorbidities in adults with epilepsy. J Clin Neurol. 2021;17:176–86. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [10].Ranjith S, Joshi A. Measures to mitigate sodium valproate use in pregnant women with epilepsy. Cureus. 2022;14:e30144. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [11].Clough E, Barrett T, Wilhite SE, et al. NCBI GEO: archive for gene expression and epigenomics data sets: 23-year update. Nucleic Acids Res. 2024;52:D138–44. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [12].Liu S, Wang Z, Zhu R, Wang F, Cheng Y, Liu Y. Three differential expression analysis methods for RNA sequencing: limma, EdgeR, DESeq2. J Vis Exp. 2021:175. [DOI] [PubMed] [Google Scholar]
- [13].Langfelder P, Horvath S. WGCNA: an R package for weighted correlation network analysis. BMC Bioinf. 2008;9:559. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [14].Martino E, Chiarugi S, Margheriti F, Garau G. Mapping, Structure and Modulation of PPI. Front Chem. 2021;9:718405. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [15].DiStefano J, 3rd, Hannah-Shmouni F, Clément F. Editorial: mechanistic, machine learning and hybrid models of the “other” endocrine regulatory systems in health & disease. Front Endocrinol (Lausanne). 2023;14:1104746. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [16].Tibshirani R. The lasso method for variable selection in the cox model. Stat Med. 1997;16:385–95. [DOI] [PubMed] [Google Scholar]
- [17].Sanz H, Valim C, Vegas E, Oller JM, Reverter F. SVM-RFE: selection and visualization of the most relevant features through non-linear kernels. BMC Bioinf. 2018;19:432. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [18].Hu J, Szymczak S. A review on longitudinal data analysis with random forest. Brief Bioinform. 2023;24:bbad002. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [19].Mulligan BP, Carniellov TN. A procedure for predicting, illustrating, communicating, and optimizing patient-centered outcomes of epilepsy surgery using nomograms and Bayes’ theorem. Epilepsy Beha. 2023;140:109088. [DOI] [PubMed] [Google Scholar]
- [20].Martínez Pérez JA, Pérez Martin PS. ROC curve. Semergen. 2023;49:101821. [DOI] [PubMed] [Google Scholar]
- [21].Franchini M, Pellecchia S, Viscido G, Gambardella G. Single-cell gene set enrichment analysis and transfer learning for functional annotation of scRNA-seq data. NAR Genom Bioinform. 2023;5:lqad024. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [22].Liberzon A, Birger C, Thorvaldsdóttir H, Ghandi M, Mesirov JP, Tamayo P. The molecular signatures database (MSigDB) hallmark gene set collection. Cell Syst. 2015;1:417–25. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [23].Zhao P, Zhen H, Zhao H, Huang Y, Cao B. Identification of hub genes and potential molecular mechanisms related to radiotherapy sensitivity in rectal cancer based on multiple datasets. J Transl Med. 2023;21:176. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [24].Gu J, Qian K, Wu G, Xu H. Development of an inflammation-related gene-based diagnostic risk model and immune infiltration analysis in bipolar disorder [published online ahead of print 2025]. Curr Med Chem. [DOI] [PubMed] [Google Scholar]
- [25].Schaudy E, Lietard J, Somoza MM. Enzymatic synthesis of high-density RNA microarrays. Curr Protoc. 2023;3:e667. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [26].Sarker IH. Machine learning: algorithms, real-world applications and research directions. SN Comput Sci. 2021;2:160. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [27].Brunskine C, Passlick S, Henneberger C. Structural heterogeneity of the GABAergic tripartite synapse. Cells. 2022;11:3150. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [28].Vetro A, Pelorosso C, Balestrini S, et al. Stretch-activated ion channel TMEM63B associates with developmental and epileptic encephalopathies and progressive neurodegeneration. Am J Hum Genet. 2023;110:1356–76. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [29].Clifton NE, Collado-Torres L, Burke EE, et al. Developmental profile of psychiatric risk associated with voltage-gated cation channel activity. Biol Psychiatry. 2021;90:399–408. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [30].Li J, Long S, Zhang Y, et al. Molecular mechanisms and diagnostic model of glioma-related epilepsy. NPJ Precis Oncol. 2024;8:223. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [31].Heider J, Stahl A, Sperlich D, et al. Defined co-cultures of glutamatergic and GABAergic neurons with a mutation in DISC1 reveal aberrant phenotypes in GABAergic neurons. BMC Neurosci. 2024;25:12. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [32].Purnell BS, Alves M, Boison D. Astrocyte-neuron circuits in epilepsy. Neurobiol Dis. 2023;179:106058. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [33].Vanderplow AM, Eagle AL, Kermath BA, Bjornson KJ, Robison AJ, Cahill ME. Akt-mTOR hypoactivity in bipolar disorder gives rise to cognitive impairments associated with altered neuronal structure and function. Neuron. 2021;109:1479–96.e6. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [34].Koh W, Kwak H, Cheong E, Lee CJ. GABA tone regulation and its cognitive functions in the brain. Nat Rev Neurosci. 2023;24:523–39. [DOI] [PubMed] [Google Scholar]
- [35].Del Calvo G, Baggio Lopez T, Lymperopoulos A. The therapeutic potential of targeting cardiac RGS4. Ther Adv Cardiovasc Dis. 2023;17:17539447231199350. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [36].Guda MR, Velpula KK, Asuthkar S, Cain CP, Tsung AJ. Targeting RGS4 ablates glioblastoma proliferation. Int J Mol Sci . 2020;21:3300. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [37].D’Souza MS, Seeley SL, Emerson N, et al. Attenuation of nicotine-induced rewarding and antidepressant-like effects in male and female mice lacking regulator of G-protein signaling 2. Pharmacol Biochem Behav. 2022;213:173338. [DOI] [PubMed] [Google Scholar]
- [38].Bi YH, Wang J, Guo Z-J, Jia K-N. Characterization of ferroptosis-related molecular subtypes with immune infiltrations in neuropathic pain. J Pain Res. 2022;15:3327–48. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [39].Riaz M, Abbasi MH, Sheikh N, Saleem T, Virk AO. GABRA1 and GABRA6 gene mutations in idiopathic generalized epilepsy patients. Seizure. 2021;93:88–94. [DOI] [PubMed] [Google Scholar]
- [40].Arslan A. Pathogenic variants of human GABRA1 gene associated with epilepsy: a computational approach. Heliyon. 2023;9:e20218. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [41].Li Y, Zhang L, Mao M, et al. Multi-omics analysis of a drug-induced model of bipolar disorder in zebrafish. iScience. 2023;26:106744. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [42].Faustmann TJ, Corvace F, Faustmann PM, Ismail FS. Effects of lamotrigine and topiramate on glial properties in an astrocyte-microglia co-culture model of inflammation. Int J Neuropsychopharmacol. 2022;25:185–96. [DOI] [PMC free article] [PubMed] [Google Scholar]










