Abstract
Abstract
Ischemic stroke (IS) is the most common subtype of stroke. However, reliable blood biomarkers for early diagnosis remain unavailable. This study developed a predictive model based on peripheral blood (PB) biomarkers. PB samples from two independent cohorts including IS patients and healthy controls (CTR) were analyzed by RNA sequencing (RNA-seq). 69 mRNAs were consistently and significantly dysregulated in IS patients. Functional enrichment analysis revealed that the IS phenotype was negatively associated with NK cell-mediated cytotoxicity and single-sample gene set enrichment analysis (ssGSEA) revealed a significant reduction in Cd56bright NK cells, Cd56dim NK cells, and NKT cells in IS patients. A four-gene diagnostic model-BCL2A1, FAM200B, IGJ, and TXN-was identified and exhibited high diagnostic accuracy across derivation, validation, and external cohorts (AUCs: 0.94, 0.91, and 0.96, respectively). Additionally, potential small molecule compounds were screened using Enrichr database, among which cytochalasin D may represent a novel candidate drug for IS treatment.
Graphical Abstract
Supplementary Information
The online version contains supplementary material available at 10.1007/s12265-025-10635-w.
Keywords: Ischemic stroke, RNA sequencing, NK cells, Biomarkers, Machine learning
Introduction
Ischemic stroke (IS) is the most prevalent type of stroke, accounting for approximately 87% of all stroke cases [1, 2]. According to the 2021 Global Burden of Disease (GBD) report, although the age-standardized mortality rate of stroke declined from 1990 to 2021, the incidence rate has increased substantially [3]. This highlights an urgent need for improved strategies to facilitate early prevention and detection of IS.
Currently, the clinical diagnosis of IS primarily relies on neuroimaging techniques. However, approximately 50% of early-stage IS patients exhibit nonspecific findings on imaging. Consequently, diagnosis is often delayed until after irreversible neurological damage has occurred, resulting in poor clinical outcomes [4]. In addition, existing therapeutic strategies mainly focus on recanalizing the occluded artery via intravenous administration of recombinant tissue plasminogen activator (rt-PA) [5]. Despite these efforts, the 10-year stroke recurrence rate remains high [6, 7], highlighting the urgent need for improved diagnostic and therapeutic tools.
Accumulating evidence indicates that the immune system plays a pivotal role in the pathophysiology of IS [8]. Inflammation is a key contributor to the pathogenesis of IS. Following an ischemic event, the brain initiates both acute and chronic inflammatory responses that play critical roles in tissue injury and repair [9, 10]. Numerous clinical trials and systematic reviews have implicated the involvement of neutrophils, macrophages, T cells, dendritic cells (DCs), and natural killer (NK) cells in IS pathophysiology [11–13]. Therefore, we investigated the inflammatory responses involved in IS and explored the potential roles of these immune cells, which are discussed in detail in this study.
Plasma biomarkers are increasingly being used as inclusion criteria in clinical trials [14–16]. Accordingly, the identification of peripheral blood (PB) biomarkers is important for improving the clinical diagnosis of IS. With the rapid advancement of high-throughput genomics, RNA sequencing (RNA-seq) has become a powerful tool for transcriptome-wide profiling. An increasing number of differentially expressed genes (DEGs) have been identified as effective diagnostic, therapeutic, and prognostic biomarkers [17].
In this study, we analyzed mRNA expression profiles from PB samples obtained from two independent cohorts. We performed differential expression and functional enrichment analyses to identify biologically relevant transcripts. Using multiple algorithms, we developed a prediction model for IS based on PB biomarkers. Additionally, we explored immune cell characteristics in the PB to examine correlations between hub gene expression and immune infiltration. Finally, we predicted potential therapeutic candidates based on diagnostic models using the Enrichr database and molecular docking techniques.
Methods
Ethical Approval and Consent to Participate
All experiments involving human participants were reviewed and approved by the Human Ethics Committee, Fuwai Hospital (Approval No. 2016-732), and the study was conducted in accordance with the principles of Good Clinical Practice and the Declaration of Helsinki. Written informed consent was obtained from all study participants or their legal proxies. Data retrieved from the GEO database were collected from patients who provided informed consent on the basis of guidelines laid out by the GEO Ethics, Law, and Policy Group.
Study Subjects
The participants enrolled in this study were described in our previous study [18–20]. The transcriptome sequencing study included three independent cohorts, comprising a total of 128 ischemic stroke (IS) patients and 73 age-matched healthy controls (CTR). The derivation cohort included 43 IS patients and 31 CTR recruited from Cangzhou Central Hospital between 2014 and 2017. The validation cohort comprised 16 IS patients from the General Hospital of Ningxia Medical University and 19 CTR samples from Tsinghua University Hospital enrolled between 2017 and 2019. An external validation cohort (GSE58294) containing 69 IS patients and 23 CTR was used to assess the performance of the IS diagnostic model. PB samples from IS patients were collected within 48 hours after symptom onset, as previously reported [18–20]. PB samples were drawn into PAXgene Blood RNA Tubes via venipuncture of the antecubital vein and gently inverted to mix. Samples were incubated at room temperature for 2 hours to ensure complete RNA stabilization and subsequently stored at −80 °C until RNA extraction. RNA extraction and sequencing for the derivation cohort were completed in June 2018, while those for the validation cohort were conducted in August 2019. To ensure consistency over the study period (2014–2019) and minimize technical variability, all samples were processed using standardized protocols, including uniform RNA extraction procedures, library preparation kits, and sequencing platforms. IS diagnosis was confirmed by neurologists based on clinical assessment and neuroimaging (CT or MRI). Controls were matched for age, sex, and vascular risk factors. Participant demographics and clinical characteristics were obtained through structured interviews and medical records (Tables S1 and S2). The exclusion criteria were consistent with previous protocols [20].
RNA-Seq and Data Analysis
PB samples were collected from IS patients and age-matched CTR. Total RNA was extracted as previously described [18–20]. Only samples with RNA Integrity Numbers (RIN) ≥7.5 and a 28S/18S rRNA ratio ≥1.8 were used for library construction. For each sample, 3 μg of total RNA was used to construct the libraries. Gene expression quantification and normalization were carried out using the DESeq2 package [21]. DEGs were identified based on an absolute log2 fold change (|log2FC|)|> 0.5 and a false discovery rate (FDR) < 0.05, adjusted by the Benjamini–Hochberg method. To assess the robustness of our differential expression analysis, we applied three widely used algorithms—limma, edgeR, and DESeq2—to identify DEGs under varying log2FC thresholds (0.5, 1.0, and 1.5). We subsequently examined the overlap of DEGs identified by these methods at each threshold to evaluate the consistency and reliability of the results.
Functional Enrichment Analysis of DEGs
To explore the biological significance of the DEGs, functional enrichment analysis was conducted. Gene Ontology (GO) annotation and Kyoto Encyclopedia of Genes and Genomes (KEGG) pathway enrichment were performed to identify biological processes and pathways of the common DEGs. Enrichment results with a FDR < 0.05 were considered statistically significant and were corrected using the Benjamini–Hochberg method. In addition, fast gene set enrichment analysis (fGSEA) was applied to identify consistent pathways across both the derivation and validation cohorts.
Screening and Validation of Critical Signature Genes
To develop a robust diagnostic model for IS, we employed multiple machine learning algorithms, including random forest (RF), logistic regression, and receiver operating characteristic (ROC) curve analysis. The RF model was implemented via the randomForest package (version 4.7–1), while univariate and multivariate logistic regression analyses were performed via the rms package. ROC curves and area under the curve (AUC) values were calculated via the pROC package (version 1.18.5) [22]. A two-sided p < 0.05 was considered statistically significant. The nomogram model was constructed using the rms package, and its predictive performance was evaluated through calibration curves. To assess the clinical utility of the model, decision curve analysis (DCA) and clinical impact curve (CIC) analyses were performed. The discriminatory performance was further quantified by calculating AUC values. To evaluate the model’s robustness and generalizability, we implemented a 5-fold cross-validation strategy repeated 10 times. In each iteration, the dataset was randomly partitioned into five equal subsets; four subsets were used for training, and the remaining subset was used for validation. This process was repeated five times, ensuring that each subset served once as the validation set. The entire 5-fold cross-validation process was then repeated 10 times with different random partitions to account for data variability. AUC values obtained from the 10 repetitions were 0.936, 0.947, 0.954, 0.945, 0.929, 0.938, 0.940, 0.954, 0.956, and 0.927. The mean AUC was approximately 0.943 with a standard deviation of 0.010, indicating consistently high discriminatory performance and minimal risk of overfitting.
Clinical Applicability and Covariate-Adjusted Performance of the Diagnostic Model
To evaluate the clinical utility of the four-gene diagnostic model, we conducted multivariate ROC analyses. These analyses assessed the model’s performance after adjusting for potential confounding variables, including age, gender, body mass index (BMI), and diastolic blood pressure (DBP).
Determination, Evaluation, and Correlation Analysis of Infiltrated Immune Cells
To assess immune cell infiltration, we employed single-sample gene set enrichment analysis (ssGSEA) using the GSVA R package. The gene sets representing 28 immune cell types were obtained from the TISIDB database: http://cis.hku.hk/TISIDB/data/download/CellReports.txt. The ssGSEA scores were calculated for each sample to quantify the relative abundance of these immune cell types. Subsequently, we compared the ssGSEA scores between IS patients and CTR using the t test. Additionally, we utilized Spearman correlation analyses to examine the correlations between key genes and various immune cells.
Drug–Gene Interaction and Molecular Docking Analysis
Drugs and compounds were predicted by Enrichr [23], with a significance threshold of p < 0.05. Compounds were ranked based on their combined scores. The molecular structures of the ligands and target proteins were obtained from PubChem, the Protein Data Bank (PDB), CellMiner, HERB, and other databases and visualized via Cytoscape software. Molecular docking simulations were performed using AutoDock Vina to calculate the docking energies between ligands and target proteins [24]. Finally, the docking complex was visualized with Discovery Studio software.
Molecular Dynamics (MD) Simulations
MD simulations were performed using the Gromacs 2022 software package. The General Amber Force Field (GAFF) was applied for small molecules, whereas the AMBER14SB force field and the TIP3P water model were used for proteins. During the MD simulations, all hydrogen bond constraints were applied using the LINCS algorithm, with an integration time step of 2 fs. Electrostatic interactions were calculated using the Particle-Mesh Ewald (PME) method with a cutoff of 1.2 nm, whereas the cutoff for non-bonded interactions was set to 10 Å, which was updated every 10 steps. The simulation temperature was maintained at 298 K using the V-rescale temperature coupling method, and the pressure was controlled at 1 bar using the Berendsen method. The system underwent 100 ps of NVT and NPT equilibration at 298 K, followed by a 100 ns MD simulation, with conformations saved every 10 ps. To validate docking results, we conducted duplicated 100 ns molecular dynamics simulations for the top ligand-protein complex, using different initial random velocities.
Statistical Analysis
All statistical analyses were performed using SPSS version 23.0 (IBM Corp., USA). The distribution of continuous variables was evaluated using the Kolmogorov–Smirnov test. For normally distributed data, differences between groups were assessed using unpaired two-tailed Student’s t-tests. Categorical variables were compared using chi-square tests. Continuous variables are presented as means ± standard deviations (SD) for normally distributed data or as medians with interquartile ranges (IQR) for non-normally distributed data. Statistical significance was defined as p < 0.05.
Results
mRNA Expression Profiles are Significantly Altered in IS Patients in Both the Derivation and Validation Cohorts
To explore transcriptomic alterations associated with IS, RNA-seq was performed on PB samples from two independent cohorts. The derivation cohort included 43 IS patients and 31 CTR, while the validation cohort included 16 IS patients and 19 CTR. Detailed demographic and clinical characteristics of the participants are summarized in Tables S1 and S2. In the derivation cohort, a total of 521 DEGs were identified, consisting of 64 upregulated and 457 downregulated genes (Fig. 1A, C and Table S3). Similarly, in the validation cohort, a total of 378 DEGs were detected, including 229 upregulated and 149 downregulated genes (Fig. 1B, D and Table S4). Comparative analysis across the two datasets identified 69 consistently dysregulated DEGs, which were selected for further investigation (Fig. 1E). The expression profiles of these 69 common DEGs in both the derivation and validation datasets are illustrated in heatmaps (Fig. 1F, G). To evaluate the robustness of our differential expression analysis, we conducted a sensitivity analysis using three different statistical methods across multiple log2FC thresholds. Notably, the greatest overlap in identified DEGs was observed at a log2FC threshold of 0.5, suggesting that this cutoff provides greater consistency and stability in capturing biologically relevant signals (Supplemental Fig. 1A-C). Next, we performed functional enrichment analysis of the 69 common DEGs. Gene Ontology (GO) analysis revealed that these DEGs were significantly enriched in several biological process (BP) terms, including leukocyte-mediated immunity, leukocyte proliferation, and positive regulation of interleukin-6 production (Supplemental Fig. 2A). Subsequently, the DEGs were subsequently subjected to Kyoto Encyclopedia of Genes and Genomes (KEGG) pathway enrichment analysis, which highlighted the natural killer (NK)-mediated cytotoxicity as one of the most significantly enriched pathways (Supplemental Fig. 2B). Additionally, functional gene set enrichment analysis (fGSEA) revealed that glycolysis/gluconeogenesis was positively associated with the IS phenotype, whereas NK cell-mediated cytotoxicity, cytokine-cytokine receptor interactions were negatively associated in both the derivation and validation cohorts (Fig. 1H, I).
Fig. 1.
DEGs between IS patients and CTR in the derivation and validation cohorts. A, B Volcano plots of DEGs distributions in the (A) derivation and (B) validation cohorts. Nodes in yellow represent upregulated genes (Up), nodes in green represent downregulated genes (Down), and gray dots represent no significantly changes genes (NS). C, D Number of DEGs in the (C) derivation and (D) validation cohorts. E Venn diagram of DEGs from the two cohort datasets. F, G Heatmaps of the 69 common DEGs in the (F) derivation and (G) validation cohort datasets. H Derivation cohort and (I) validation cohort fGSEA KEGG enrichment analysis. DEGs, Differentially Expressed Genes; IS, ischemic stroke; CTR, healthy controls; fGSEA, fast Gene Set Enrichment Analysis; KEGG, Kyoto Encyclopedia of Genes and Genomes
Altered Immune Cell Composition in IS Patients
The immune cell compositions of IS patients and CTR were analyzed to explore immune alterations associated with IS. The distributions of 28 immune cell types in each sample from both the derivation and validation cohorts are presented in Fig. 2A, B. Comparisons of immune cell types between IS patients and CTR in both cohorts are shown in Fig. 2C, D. Notably, three immune cell subsets - CD56bright NK cells, CD56dim NK cells, and NKT cells - exhibited consistent and significant reductions in IS patients compared with CTR across both cohorts (all p < 0.05). These results underscore the pivotal role of immune cell dysregulation, particularly involving NK cells and their subsets, in the pathogenesis of IS.
Fig. 2.
Immune infiltration analysis in IS and CTR. A, B The heatmap of the levels of 28 immune cells in IS and control samples from (A) derivation and (B) validation cohorts. C, D Wilcoxon test determines the differences of immune cells infiltrating levels between IS and CTR in (C) derivation and (D) validation cohorts. IS, ischemic stroke; CTR, healthy controls. *p < 0.05, **p < 0.01, ***p < 0.001
Identification of Potential Diagnostic Biomarkers for IS
We employed four different algorithms to identify potential IS biomarkers. First, a RF model was used to rank gene feature importance on the basis of gene expression levels across all samples, and the top 25 genes were selected for further analysis (Fig. 3A, B). To evaluate the diagnostic potential of the 25 genes, ROC curve analysis was performed, and the AUC values were calculated to determine their ability to discriminate between IS patients and CTR (Fig. 3C, Table S5). Among the 25 genes, FAM200B, TXN, TAF7, SAMD9, CLEC2B, BCL2A1, IGJ achieved AUC values greater than 0.70, indicating high sensitivity and specificity for IS diagnosis and supporting their potential as diagnostic biomarkers. In addition, univariate logistic regression analysis revealed that these seven genes were significantly associated with the occurrence of IS (Table S6). To improve the accuracy and robustness of the diagnostic model, we further performed multivariable logistic regression analysis to assess the independent effects of each gene. Model optimization was guided by the Akaike Information Criterion (AIC), with the combination yielding the lowest AIC value considered the optimal model. Ultimately, a combination of four candidate genes (BCL2A1, FAM200B, IGJ, and TXN) presented the lowest AIC value (AIC = 55.86, Table 1), suggesting that this combination was the most stable and robust. Consequently, these four genes were identified as hub ischemic stroke-related genes (HISRGs). Correlation analysis revealed that the four genes were significantly positively correlated, suggesting that they may play a synergistic regulatory role in IS (Fig. 3D).
Fig. 3.
Screening of critical signatures via multiple machine learning algorithms. A, B Identification of signatures via RF algorithm. A Distribution of the OOB error rate at various values of trees. B Variable importance, as measured by the mean decrease in the Gini coefficient, is computed via the OOB error. The genes are shown in descending order of importance. C ROC analysis of hub genes in the derivation cohort. D Correlation chord diagram of the four hub genes. RF, random forest; OOB, out-of-band; ROC, Receiver operating characteristic
Table 1.
Final diagnostic models with the AIC in the multivariate logistic regression analysis
| Model | Hub genes investigated | AIC |
|---|---|---|
| 1 | FAM200B + BCL2A1 + IGJ + TXN | 55.86 |
| 2 | FAM200B + SAMD9 + BCL2A1 + IGJ + TXN | 57.38 |
| 3 | FAM200B + SAMD9 + BCL2A1 + IGJ + TXN + CLEC2B | 58.65 |
| 4 | FAM200B + SAMD9 + BCL2A1 + IGJ + TAF7 + TXN + CLEC2B | 60.65 |
Diagnostic Marker Validation and Nomogram Construction
The diagnostic power of HISRGs was confirmed across multiple cohorts, including our derivation cohort, validation cohort, and an independent external validation cohort GSE58294. As shown in Fig. 4A, the AUC values of all HISRGs (BCL2A1, FAM200B, IGJ, and TXN) exceeded 0.70, with the AUC value of the combined predictive model in the derivation cohort was greater than 0.90. Similar performance was observed in both the validation cohort and GSE58294, where BCL2A1, FAM200B, and TXN maintained AUC values above 0.70, although IGJ exhibited a slightly lower value (Fig. 4B, C). Importantly, the overall model retained high accuracy, with AUC values exceeding 0.90 across all cohorts, demonstrating its strong diagnostic potential. To further evaluate the clinical utility of the four-gene diagnostic model, we performed multivariate ROC analysis to assess its performance after adjusting for potential confounding factors, including age, gender, BMI, and DBP. In the derivation cohort, the unadjusted model yielded an AUC of 0.94, which slightly decreased to 0.91 after adjustment (Supplemental Fig. 3A). Conversely, the validation cohort showed a reduction in AUC from 0.91 to 0.76 following adjustment. The relatively modest decrease in AUC, particularly in the derivation cohort, suggests that the model retains strong discriminative power independent of these clinical variables, underscoring its potential as a reliable diagnostic tool for IS (Supplemental Fig. 3B). Furthermore, the differential expression of these hub genes was analyzed between the CTR and IS groups. The results confirmed that HISRGs expression levels were significantly upregulated in the IS samples within the derivation cohort (Fig. 4D). Moreover, these expression trends were consistent in both the validation cohort and the external dataset (Fig. 4E, F), supporting the robustness of their diagnostic value. Based on the expression of HISRGs, a nomogram was constructed to estimate an individual’s probability of developing IS (Fig. 4G). The calibration curve demonstrated excellent agreement between the predicted and observed outcomes, confirming the high diagnostic accuracy of the model (Fig. 4H). DCA was performed to assess the clinical utility of the predictive model. As shown in Fig. 4I, the nomogram model demonstrated a greater net clinical benefit than the “treat all” or “treat none” strategies, indicating its superior clinical applicability. Notably, the nomogram model outperformed each individual gene in terms of predictive power, further supporting its robustness in predicting IS risk. Additionally, the CIC analysis revealed that when the high-risk threshold ranged from 0.6 to 1.0, the number of patients classified as high-risk closely matched the number of true events, indicating that the nomogram model can provide meaningful clinical benefit (Fig. 4J). Collectively, these results indicate that HISRGs represent reliable diagnostic biomarkers for IS. The nomogram model, validated across multiple independent cohorts, exhibited high predictive accuracy and strong clinical utility, reinforcing its potential for implementation in clinical IS diagnosis and risk stratification.
Fig. 4.
Validation of potential diagnostic biomarkers for IS patients and CTR. A-C ROC analysis of the ability of FAM200B, BCL2A1, IGJ, and TXN to diagnose IS in the (A) derivation, (B) validation, and (C) external validation cohorts in GSE58294. D-F Differential expression of FAM200B, BCL2A1, IGJ, and TXN in the (D) derivation, (E) validation, and (F) external validation cohorts in GSE58294. G A nomogram of FAM200B, BCL2A1, IGJ, and TXN was constructed for diagnosing IS. An illustrative example demonstrating the assignment of risk scores based on gene expression levels and the calculation of a total risk score: For each gene, points are assigned by drawing a vertical line from the specific expression value to the “Points” axis in the nomogram. The points corresponding to each of the four genes are then summed and located on the “Total Points” axis. For example, a patient with the following gene expression levels–FAM200B: 4.886 (26 points), BCL2A1: 22.327 (53 points), IGJ: 21.401 (2 points), and TXN: 7.656 (8 points)–the total score is 89 points. This total corresponds to a low predicted risk of the event. H Calibration curve was used to evaluate the nomogram model. I DCA curve suggesting that decision-making based on the nomogram may benefit IS patients because the pink lines are consistently maintained above the gray and black lines from 0 to 1. J Clinical impact curves of the combination of the four genes for discriminating IS patients from CTR. IS, ischemic stroke; CTR, healthy controls; ROC, Receiver operating characteristic; DCA, Decision Curve Analysis; CIC, Clinical impact curves. *p < 0.05, **p < 0.01, ***p < 0.001
Correlation Between HISRGs Expression and Immune Cell Infiltration in IS
We aimed to investigate the relationships between HISRGs and infiltrating immune cells in IS. The results revealed that BCL2A1, FAM200B, IGJ, and TXN were negatively correlated with CD56dim NK cells (Fig. 5A). Additionally, the correlations between individual hub genes and specific immune cell subsets were visualized using a lollipop chart (Fig. 5B). Furthermore, combined insights from the fGSEA and immune infiltration analyses suggest that NK cells, particularly CD56dim NK cells, may serve as pivotal immune components in the pathophysiology of IS.
Fig. 5.
Correlation between BCL2A1, FAM200B, IGJ, TXN with infiltrating immune cells. A Spearman correlation analysis between the expression levels of HISRGs and the infiltration levels of 28 immune cell types in IS samples. Green and red colors indicate negative and positive correlations, respectively. B Lollipop plot displaying the correlation coefficients between HISRGs and specific immune cell subsets. The size of the dots represents the strength of the correlation between genes and immune cells. The color of dots represents the p-value. *p < 0.05, **p < 0.01, ***p < 0.001
Molecular Docking and MD Simulations
Among the four HISRGs, TXN had the highest AUC value (AUC = 0.95), indicating that it had the strongest diagnostic performance for IS (Fig. 4C). Given its diagnostic relevance and regulatory importance, we further focused on TXN for drug screening to explore its potential as a therapeutic target. The Enrichr package was used to screen for TXN-related compounds. A total of 157 compounds with p < 0.05 were identified (Table S7). The top ten compounds with the highest combined scores were selected, and their p values associated with thioredoxin, the protein encoded by TXN, were calculated (Fig. 6A). Among them, cytochalasin D (Cyto D) demonstrated the lowest docking score (−6.4 kcal/mol) with thioredoxin, suggesting a strong potential interaction (Table 2). A surface diagram illustrating the binding of Cyto D to thioredoxin is shown in Fig. 6B, while the results of the molecular docking simulation are presented in Fig. 6C. To further confirm the precise binding mechanism and interaction stability, MD simulations of the ligand-receptor complex were conducted (Supplemental Fig. 4). The MD simulations demonstrated stable interactions between the ligand and active site residues. In two independent 100 ns simulations, we analyzed the stability of key interactions between the ligand and the active site residues. The distance between the ligand’s carbon atom and the sulfur atom (SG) of CYS-69 remained consistently within the range of 0.5–1.0 nm, indicating a stable hydrophobic interaction (Fig. 6D). Similarly, the distance between the ligand’s oxygen atom and the terminal nitrogen atom (NZ) of LYS-8 was maintained between 0.6–1.5 nm throughout the simulations, suggesting a persistent hydrogen bond interaction (Fig. 6E). To ensure the reproducibility of the MD simulation results, an independent replicate simulation was performed using a different initial velocity seed, further validating the robustness of the binding interactions. (Supplemental Fig. 5).
Fig. 6.
Identification of potential drugs that target TXN. A Drugs and compounds were predicted via Enrichr. B 3D interaction diagram of the compound Cyto D with the thioredoxin protein. C 2D interaction diagram of Cyto D and the thioredoxin protein. The colors of the heteroatoms in the compounds and amino acid residues are shown according to the element type. The hydrogen bonds are represented by the green dotted line, and the hydrophobic interactions are shown as the pink dotted line. D Distance between the ligand’s carbon atom and the SG of residue CYS-69 in the first MD simulation. E Distance between the ligand’s oxygen atom and the terminal NZ of residue LYS-8 in the first MD simulation. Cyto D, cytochalasin D; SG, sulfur atom; NZ, nitrogen atom; MD, molecular dynamics
Table 2.
Docking scores and interactions of the five drugs with the TXN
| Compound | AutoDock Vina Score | H-Bond Interactions | Hydrophobic Interactions |
|---|---|---|---|
| cytochalasin_D | −6.4 | GLY83, LYS85 | LYS8, VAL65, CYS69, PHE80, LEU15 |
| Manumycin_A | −5.9 | GLN84 | LEU104, PHE89 |
| Diphenylcyclopropenone | −5.8 | LYS8, VAL65, LYS85 | |
| selegiline | −4.7 | VAL65, LYS8, PHE11, LYS85 | |
| salicylic acid | −4.7 | GLU88,LYS72,THR76, CYS73,VAL71 |
Discussion
In this study, we conducted transcriptomic profiling of PB samples from IS patients and CTR, identifying a distinct mRNA signature associated with disease status. Functional enrichment analysis revealed significant enrichment of immune-related pathways, with NK cell-mediated cytotoxicity emerging as the most prominent. These findings align with results from the MARVEL trial, which emphasized the therapeutic potential of targeting neuroinflammatory processes in acute IS [25]. By integrating molecular biomarkers with clinical interventions, both our study and the MARVEL trial underscore the importance of addressing neuroinflammation to improve patient outcomes. Further immune cell profiling demonstrated a marked reduction in three types of immune cell subsets in IS patients, including CD56dim NK cells, CD56bright NK cells, and NKT cells. Using an integrated approach combining multiple feature selection algorithms, we identified a robust IS diagnostic model based on four hub genes—BCL2A1, FAM200B, IGJ, and TXN—collectively referred to as HISRGs. These genes exhibited excellent discriminatory performance, with AUC values of 0.94, 0.91, and 0.96 in the derivation, validation, and external cohorts, respectively. Further correlation analysis revealed a significant negative association between HISRGs expression and CD56dim NK cell abundance. Finally, through enrichment-based drug screening and molecular docking analysis, we identified Cyto D as a promising candidate compound targeting TXN, providing a potential therapeutic avenue for IS treatment.
In China, an estimated 3.4 million new cases of stroke and approximately 2.3 million stroke-related deaths occur annually, highlighting the substantial public health burden imposed by the disease [26]. Despite its global burden, IS still lacks reliable blood-based biomarkers approved for clinical use, underscoring the urgent need for effective molecular tools to facilitate early diagnosis. In our derivation and validation cohorts, fGSEA revealed that NK cell-mediated cytotoxicity was negatively associated with the IS phenotype. These findings align with previous studies demonstrating that brain ischemia compromises NK cell-mediated immune defenses [27]. For instance, miR-1224 has been identified as a negative regulator of NK cell activation in an Sp1-dependent manner, leading to suppressed NK cell activity following ischemic injury [28]. Immune infiltration analysis further demonstrated significant alterations in circulating immune cells among IS patients, with NK cells being particularly affected. In line with earlier studies [27, 29, 30], we observed marked reductions in CD56dim NK cells, CD56bright NK cells, and NKT cells in both the derivation and validation cohorts. Previous research using ssGSEA also reported significant reductions in CD56dim NK cells IS patients compared to CTR [31]. In addition, flow cytometry analysis revealed that the proportion of NKT cells in the PB of IS patients is significantly reduced on days 1, 3, and 7 after stroke onset [32]. Cumulative evidence has established the important role of NK cells in CNS diseases [33]. However, their role appears to be dualistic and context-dependent. On the one hand, NK cell-mediated cytotoxic activity may exacerbate neuronal injury by promoting cell death and inflammation [29]. On the other hand, NK cells may contribute to preserving immune defense and preventing post-stroke infections, particularly by supporting peripheral immune surveillance and maintaining immune homeostasis within the brain [28]. A causal role for NK cells in worsening ischemic brain injury has been demonstrated in animal models, where NK cell depletion prior to middle cerebral artery occlusion (MCAO) significantly reduced infarct size and improved neurological outcomes [29, 30]. Together, these findings support our transcriptomic observations and suggest that the observed reduction in NK cells may reflect a compensatory or protective response in the acute phase of IS. Future mechanistic studies involving NK cells depletion or adoptive transfer models will be essential to delineate their functional role in IS pathogenesis and clinical prognosis.
Given the observed reduction in NK cell subsets in IS patients, we further explored whether the identified HISRGs were associated with NK cell infiltration and function. Correlation analysis revealed that BCL2A1, FAM200B, IGJ, and TXN were negatively correlated with the abundance of CD56dim NK cells, consistent with the results from fGSEA and immune infiltration analyses. These findings further support the role of NK cell–related immune dysregulation in the pathogenesis of IS. NK cells are generally divided into two major subsets based on CD56 surface expression: CD56bright and CD56dim NK cells. CD56bright NK cells predominantly reside in secondary lymphoid tissues and are considered less mature. Upon activation, they secrete high levels of immunoregulatory cytokines, including interferon-gamma (IFN-γ), tumor necrosis factor-alpha (TNF-α), and interleukin-10 (IL-10) [34]. In contrast, CD56dim NK cells constitute the predominant NK cell population in PB and exhibit potent cytotoxic activity. They are characterized by elevated expression levels of perforin and granzyme B, and mediate antibody-dependent cellular cytotoxicity (ADCC) through the expression of CD16 (FcγRIII) [35]. The expression of HISRGs may be influenced by the cytokine milieu shaped by different NK cell subsets. CD56bright NK cells, through their robust cytokine production, can modulate the activity of other immune cells, potentially affecting the expression of HISRGs involved in inflammatory responses [36]. IFN-γ produced by CD56bright NK cells has been shown to activate monocytes and enhance antigen presentation, leading to the upregulation of inflammation-related genes [37]. Conversely, during activation or exhaustion, CD56dim NK cells may downregulate activating receptors such as NKG2D and CD94, while upregulating inhibitory receptors including KIRs and LIR-1, thereby altering gene expression profiles related to cytolytic potential [38].
In this study, we identified 69 DEGs consistently shared across both the derivation and validation cohorts. Using four independent feature selection algorithms, we further selected four candidate hub genes-BCL2A1, FAM200B, IGJ, and TXN-as IS associated signatures. All four genes were significantly upregulated in IS samples and demonstrated strong diagnostic performance, with AUC values exceeding 0.90 in the derivation, validation, and external independent cohorts. Functionally, these hub genes are involved in immune and inflammatory signaling pathways relevant to IS pathogenesis. Among them, TXN, which encodes thioredoxin, is particularly notable for its immunomodulatory functions. TXN regulates immune response by modulating chemotaxis and inhibiting leukocyte infiltration at sites of inflammation [39, 40]. Previous studies have identified serum TXN levels as an independent diagnostic marker for IS and suggested that incorporating TXN into existing clinical scoring systems may improve diagnostic accuracy [41]. Moreover, TXN has been proposed as a novel blood biomarker with both diagnostic and prognostic value in IS [42]. Our findings provide additional support for the clinical potential of TXN and underscore its promise as a therapeutic target in IS.
BCL2A1 has been identified as a core gene involved in immune regulation in the PB following IS [43]. Previous studies have reported elevated blood concentrations of Bcl-2 in non-surviving IS patients compared with survivors, suggesting a potential association between high Bcl-2 expression and increased mortality risk in IS [44]. In addition, BCL2A1 overexpression has been observed in intracranial melanomas, where it promotes tumor cell survival and may contribute to disease progression [45]. IGJ, also known as the J chain, is a crucial component of the humoral immune system. It facilitates the polymerization of immunoglobulin molecules to form dimeric IgA and pentameric IgM complexes. These polymeric antibodies are essential for mucosal immunity and the primary immune response, respectively [46]. However, a direct link between IGJ and IS pathophysiology remains to be elucidated. FAM200B is poorly characterized in the current literature, and its precise biological function remains largely unknown. Nevertheless, its identification as one of the HISRGs suggests potential involvement in IS-related immune regulation, warranting further investigation.
Although several biomarkers for IS have been reported in previous studies, such as glial fibrillary acidic protein (GFAP) and S100β, these markers have inherent limitations. Both GFAP and S100β are primarily derived from astrocytes and are released into the bloodstream following cellular injury and blood-brain barrier (BBB) disruption after IS. Their elevated levels typically reflect structural brain damage rather than early-stage molecular events, and their response kinetics tend to lag behind the onset of pathological processes [47]. In contrast, the expression of the HISRGs identified in our study reflects transcriptomic alterations in PB cells during the early phase of IS. These genes are involved in immune regulation, apoptosis, and oxidative stress pathways, offering valuable insights into early molecular responses associated with IS pathogenesis. Moreover, different biomarkers exhibit distinct diagnostic time windows. GFAP levels rise as early as 3 hours post-onset, peak at approximately 4.5 hours, and gradually decline by 24 hours [48]. S100β levels increase within 24 hours and correlate with infarct size and neurological outcomes [49]. However, the utility of both markers for ultra-early diagnosis remains limited. Notably, several studies have demonstrated that widespread transcriptomic changes occur in PB cells within 3 hours of IS onset [50], suggesting that HISRGs may serve as earlier indicators of disease. The diagnostic performance of these biomarkers has been evaluated using ROC curve analysis. Previous studies reported an AUC of 0.915 for GFAP [48], while S100β displayed only modest diagnostic accuracy, with reported AUCs around 0.70 [51]. In contrast, the four-gene diagnostic model developed in our study, comprising BCL2 A1, FAM200B, IGJ, and TXN, demonstrated robust diagnostic accuracy, with AUCs of 0.94, 0.91, and 0.96 in the derivation, validation, and external cohorts, respectively.
GFAP and S100β primarily reflect structural damage to brain tissue, providing information on the extent of injury but offering limited insight into dynamic changes within the immune system. Recent literature has increasingly emphasized the importance of understanding the interplay between metabolism, inflammation, and cardiovascular diseases. For instance, Huang and Sun highlighted the complex interactions between metabolic reprogramming and inflammatory signaling in the pathogenesis of cardiovascular disorders [52]. Our findings align with this perspective, as the identified four key biomarkers were significantly negatively correlated with CD56dim NK cells in IS patients, highlighting the pivotal role of immune dysregulation in IS pathogenesis. The altered expression patterns of HISRGs may reflect functional changes in immune cell populations, offering novel mechanistic insights into the immunological basis of IS.
The diagnostic biomarkers identified in this study exhibit strong potential for clinical translation. Their expression levels can be accurately quantified using RT-qPCR, a widely adopted technique in clinical diagnostics owing to its high sensitivity, specificity, cost-effectiveness, and compatibility with PB samples [53]. To further enhance clinical applicability, future researches should focus on developing rapid, blood-based detection techniques, such as point-of-care testing (POCT) platforms, based on these gene signatures. POCT technologies enable timely and accurate diagnosis at or near the patient site, which is particularly beneficial in emergency or resource-limited settings. Integrating these biomarkers into routine diagnostic workflows could facilitate early detection and real-time monitoring of IS, ultimately improving patient stratification and clinical outcomes.
Although recent advances in treatment have improved prognosis in some patients, significant limitations remain, including a narrow therapeutic time window, limited efficacy, and increased risk of hemorrhagic complications [54, 55]. Given these challenges, we conducted targeted drug screening on the basis of diagnostic genes and proposed a novel therapeutic approach aimed at expanding treatment options for IS. One promising therapeutic candidate identified in this study is Cyto D. It has been shown to reduce infarct size in a mouse model of MCAO by modulating the expression of endothelial nitric oxide synthase (eNOS) [56]. Interestingly, Cyto D has been found to upregulate CD56 expression on the surface of NK cells, thereby enhancing their proliferation and cytokine secretion capacity [57]. These findings suggest that Cyto D may facilitate NK cell activation and could potentially exert neuroprotective and immunomodulatory effects in the context of IS. A study by Endres et al. further confirmed the potential of Cyto D, reporting a 25% to 35% reduction in infarct volume in 129/SV wild-type mice following cerebral ischemia. Notably, similar protective effects were observed in gsn–/– mice, in which Cyto D treatment reduced infarct volume by nearly 50%, comparable to its effects in gsn+/+ mice [58]. In vitro, Wang et al. demonstrated that Cyto D significantly attenuated autophagy and reduced cell injury in PC12 cells subjected to oxygen-glucose deprivation/reoxygenation (OGD/R). Furthermore, Cyto D enhanced cell viability in primary cortical neurons under the same conditions, reinforcing its neuroprotective profile [59].
Despite these encouraging findings, evaluating the safety of Cyto D is essential for its clinical translation. Although direct evidence of its ability to cross the BBB remains limited, studies have shown that Cyto D can affect BBB integrity. For instance, Mentor et al. reported that Cyto D disrupts F-actin organization in endothelial cells, leading to a dose-dependent decrease in transendothelial electrical resistance (TEER) and a significant increase in BBB permeability at all tested concentrations [60]. These results suggest that Cyto D can modulate BBB properties, although its ability to penetrate an intact BBB in vivo remains to be fully elucidated. Furthermore, Cyto D exhibits dose-dependent cytotoxicity. Increasing concentrations of Cyto D have been shown to inhibit endothelial cells damage [60], while higher doses have been associated with prolonged survival in leukemia-challenged mice in a dose-dependent manner [61]. However, compared with its analogs, Cyto D demonstrates significantly higher toxicity in mice [62]. To enhance the clinical applicability of Cyto D, current efforts are focusing on developing less toxic derivatives or incorporating targeted delivery systems, such as nanocarriers. These approaches aim to improve BBB penetration, reduce systemic toxicity, and ultimately maximize the therapeutic potential of Cyto D in IS treatment.
In summary, this study presents a novel four-gene diagnostic model—comprising BCL2A1, FAM200B, IGJ, and TXN—that demonstrates high diagnostic accuracy for IS across multiple independent cohorts. The identified HISRGs reflect early transcriptomic alterations in PB cells and provide insights into immune dysregulation underlying IS pathogenesis. Notably, the significant negative correlation between HISRGs expression and CD56dim NK cell abundance underscores the critical involvement of immune dysregulation in IS. Moreover, the identification of Cyto D as a potential therapeutic agent through small-molecule screening underscores the translational potential of our findings. Collectively, this study advances molecular-level understanding of IS and provides a foundation for developing early diagnostic biomarkers and targeted therapeutic strategies, with the potential to improve clinical management and patient outcomes. Despite these promising findings, several limitations should be acknowledged. First, although key diagnostic marker genes were identified and validated across independent cohorts, larger, multicenter prospective studies are required to further verify the robustness and generalizability of the diagnostic model. Second, while we observed marked reductions in immune cell subsets (such as CD56dim NK cells) in IS patients in both derivation and validation cohorts, additional experimental validation (e.g., flow cytometry or immunohistochemistry) is needed to confirm these findings. Finally, although Cyto D was identified as a promising therapeutic agent and has been reported to inhibit TXN nuclear translocation and NF-κB activation [63], further investigation is required to elucidate its regulatory effects on TXN expression under ischemic conditions, such as in the OGD model.
Supplementary Information
(PDF 1399 kb)
(XLSX 9 kb)
(XLSX 9 kb)
(XLSX 33 kb)
(XLSX 26 kb)
(XLSX 9 kb)
(XLSX 10 kb)
(XLSX 21 kb)
Acknowledgements
We would like to thank the creators of the GEO database for the valuable and publicly available data used in this research.
Abbreviations
- AIC
Akaike’s information criterion
- AUC
Area under the receiver operating characteristic curve
- BP
Biological Process
- CC
Cellular Component
- CIC
Clinical impact curve
- CNS
Central nervous system
- CT
Computed Tomography
- CTR
Healthy controls
- DCA
Decision Curve Analysis
- DCs
Dendritic cells
- DEGs
Differentially expressed genes
- eNOS
endothelial nitric oxide synthase
- FDR
False Discovery Rate
- fGSEA
Fast gene set enrichment analysis
- FPKM
Fragments Per Kilobase of transcript per Million mapped reads
- GAFF
General amber force field
- GBD
Global Burden of Disease, Injuries, and Risk Factors Study
- GEO
Gene Expression Omnibus
- GO
Gene Ontology
- HISRGs
Hub IS-Related Genes
- IQR
Interquartile ranges
- IS
Ischemic stroke
- KEGG
Kyoto Encyclopedia of Genes and Genomes
- MCAO
Middle cerebral artery occlusion
- MD
Molecular dynamics
- MF
Molecular Functions
- MRI
Magnetic Resonance Imaging
- NK
Natural Killer
- NZ
Nitrogen atom
- PB
Peripheral Blood
- PDB
Protein Data Bank
- PME
Particle-mesh ewald
- RF
Random Forests
- RIN
RNA Integrity Numbers
- RNA-seq
RNA sequencing
- ROC
Receiver operating characteristic
- SDs
Standard Deviations
- SG
Sulfur atom
- ssGSEA
single-sample Gene Set Enrichment Analysis
- TXN
Thioredoxin
Author Contributions
All the authors contributed to the study conception and design. JL, CB and HY collected the data. LS and YS conducted the experiments. HX handled the data and performed the bioinformatics analysis. MS and ZG were responsible for data acquisition and analysis. JL drafted the manuscript. HL, FW and JC provided scientific supervision and revised the article. All the authors read and approved the final manuscript for submission.
Funding
This study was supported by the National Natural Science Foundation of China (82130013 to Dr Chen, 82170448 to Dr Li), Chinese Academy of Medical Sciences Innovation Fund for Medical Sciences (2021-I2M-1-007 and 2023-I2M-2-003 to Dr Chen), and Henan Cardiovascular Disease Center (Central China Subcenter of National Center for Cardiovascular Diseases) (2023-FZX03 to Dr Chen and 2023-FZX09 to Dr Li).
Data Availability
The RNA-seq data have been deposited in the Genome Sequence Archive (GSA) at the National Genomics Data Center, China National Center for Bioinformation/Beijing Institute of Genomics, Chinese Academy of Sciences, and are publicly accessible at: https://ngdc.cncb.ac.cn/gsa-human/browse/HRA001807. Publicly available datasets used in this study were obtained from the NCBI Gene Expression Omnibus (GEO) under accession number: https://www.ncbi.nlm.nih.gov/geo/query/acc.cgi?acc=GSE58294. In addition, the analysis codes used in this study are available at the following GitHub repository: https://github.com/Bon-jour/identification-of-blood-based-biomarkers-for-ischemic-stroke.git.
Declarations
Ethics Approval
All experiments involving human participants were reviewed and approved by the Human Ethics Committee, Fuwai Hospital (Approval No. 2016-732), and the study was conducted in accordance with the principles of Good Clinical Practice and the Declaration of Helsinki. Written informed consent was obtained from all study participants or their legal proxies. Data retrieved from the GEO database were collected from patients who provided informed consent on the basis of guidelines laid out by the GEO Ethics, Law, and Policy Group.
Consent to Participate
All individuals were informed about the purposes of the study and signed their consent for publishing the data.
Consent to Publish
After reviewing the manuscript, all the authors agreed with its publication in the current form.
Animal Studies
No animal studies were carried out by the authors for this article
Conflict of interest
The authors declare that they have no conflict of interest.
Clinical Trial Number
Not applicable.
Footnotes
Publisher’s Note
Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
Contributor Information
Hao Li, Email: lihao@bjmu.edu.cn.
Feng Wang, Email: cjr.wangfeng@vip.163.com.
Jingzhou Chen, Email: chendragon1976@aliyun.com.
References
- 1.Capirossi C, et al. Epidemiology, organization, diagnosis and treatment of acute ischemic stroke. Eur J Radiol Open. 2023;11:100527. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 2.Benjamin EJ, et al. Heart Disease and Stroke Statistics-2019 Update: A Report From the American Heart Association. Circulation. 2019;139(10):e56–e528. [DOI] [PubMed] [Google Scholar]
- 3.Global, regional, and national burden of stroke and its risk factors, 1990-2019: a systematic analysis for the Global Burden of Disease Study 2019. Lancet Neurol. 2021;20(10):795–820. [DOI] [PMC free article] [PubMed]
- 4.Campbell BCV, Khatri P. Stroke. Lancet. 2020;396(10244):129–42. [DOI] [PubMed] [Google Scholar]
- 5.Xiong Y, et al. Chinese stroke association guidelines on reperfusion therapy for acute ischaemic stroke 2024. Stroke Vasc Neurol. 2025. 10.1136/svn-2024-003977. [DOI] [PMC free article] [PubMed]
- 6.Rostanski SK, et al. Infarct on Brain Imaging, Subsequent Ischemic Stroke, and Clopidogrel-Aspirin Efficacy: A Post Hoc Analysis of a Randomized Clinical Trial. JAMA Neurology. 2022;79(3):244–50. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 7.Li L, et al. Incidence, outcome, risk factors, and long-term prognosis of cryptogenic transient ischaemic attack and ischaemic stroke: a population-based study. Lancet Neurol. 2015;14(9):903–13. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 8.Iadecola C, Anrather J. The immunology of stroke and dementia. Immunity. 2025;58(1):18–39. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 9.Alsbrook DL, et al. Neuroinflammation in Acute Ischemic and Hemorrhagic Stroke. Curr Neurol Neurosci Rep. 2023;23(8):407–31. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 10.Simats A, et al. Innate immune memory after brain injury drives inflammatory cardiac dysfunction. Cell. 2024;187(17):4637–4655.e26. [DOI] [PubMed] [Google Scholar]
- 11.Faura J, et al. Stroke-induced immunosuppression: implications for the prevention and prediction of post-stroke infections. J Neuroinflammation. 2021;18(1):127. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 12.Xu Z, Yang F, Zheng L. Uncovering the dual roles of peripheral immune cells and their connections to brain cells in stroke and post-stroke stages through single-cell sequencing. Front Neurosci. 2024;18. 10.3389/fnins.2024.1443438. [DOI] [PMC free article] [PubMed]
- 13.Planas AM. Role of Immune Cells Migrating to the Ischemic Brain. Stroke. 2018;49(9):2261–7. [DOI] [PubMed] [Google Scholar]
- 14.Gundersen JK, et al. Neuronal plasma biomarkers in acute ischemic stroke. J Cereb Blood Flow Metab. 2025;45(1):77–84. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 15.Becktel DA, et al. Discovering novel plasma biomarkers for ischemic stroke: Lipidomic and metabolomic analyses in an aged mouse model. J Lipid Res. 2024;65(9):100614. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 16.Wu M-H, et al. Identification of a potential prognostic plasma biomarker of acute ischemic stroke via untargeted LC-MS metabolomics. PROTEOMICS – Clinical Applications. 2023;17(6):e2200081. 10.1002/prca.202200081. [DOI] [PubMed] [Google Scholar]
- 17.García-Berrocoso T, et al. Cardioembolic Ischemic Stroke Gene Expression Fingerprint in Blood: a Systematic Review and Verification Analysis. Transl Stroke Res. 2020;11(3):326–36. [DOI] [PubMed] [Google Scholar]
- 18.Bai C, et al. Machine learning-based identification of the novel circRNAs circERBB2 and circCHST12 as potential biomarkers of intracerebral hemorrhage. Front Neurosci. 2022;16:1002590. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 19.Bai C, et al. Identification of circular RNA expression profiles and potential biomarkers for intracerebral hemorrhage. Epigenomics. 2021;13(5):379–95. [DOI] [PubMed] [Google Scholar]
- 20.Bai C, et al. Identification of immune-related biomarkers for intracerebral hemorrhage diagnosis based on RNA sequencing and machine learning. Front Immunol. 2024;15:1421942. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 21.Love MI, Huber W, Anders S. Moderated estimation of fold change and dispersion for RNA-seq data with DESeq2. Genome Biol. 2014;15(12):550. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 22.Robin X, et al. pROC: an open-source package for R and S+ to analyze and compare ROC curves. BMC Bioinformatics. 2011;12:77. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 23.Kuleshov MV, et al. Enrichr: a comprehensive gene set enrichment analysis web server 2016 update. Nucleic Acids Res. 2016;44(W1):W90–7. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 24.Trott O, Olson AJ. AutoDock Vina: improving the speed and accuracy of docking with a new scoring function, efficient optimization, and multithreading. J Comput Chem. 2010;31(2):455–61. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 25.Yang Q, et al. Methylprednisolone as Adjunct to Endovascular Thrombectomy for Large-Vessel Occlusion Stroke: The MARVEL Randomized Clinical Trial. Jama. 2024;331(10):840–9. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 26.Tu W-J, et al. Estimated Burden of Stroke in China in 2020. JAMA Netw Open. 2023;6(3):e231455–5. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 27.Liu Q, et al. Brain Ischemia Suppresses Immunity in the Periphery and Brain via Different Neurogenic Innervations. Immunity. 2017;46(3):474–87. [DOI] [PubMed] [Google Scholar]
- 28.Feng Y, et al. miR-1224 contributes to ischemic stroke-mediated natural killer cell dysfunction by targeting Sp1 signaling. J Neuroinflammation. 2021;18(1):133. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 29.Gan Y, et al. Ischemic neurons recruit natural killer cells that accelerate brain infarction. Proc Natl Acad Sci U S A. 2014;111(7):2704–9. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 30.Zhang Y, et al. Accumulation of natural killer cells in ischemic brain tissues and the chemotactic effect of IP-10. J Neuroinflammation. 2014;11:79. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 31.Zhang H, et al. Identification of hypoxia- and immune-related biomarkers in patients with ischemic stroke. Heliyon. 2024;10(4) [DOI] [PMC free article] [PubMed]
- 32.张慧颖, et al., 缺血性脑卒中患者急性期外周血中NKT细胞比例变化的意义. 继续医学教育, 2017. 31(11): p. 91-93.
- 33.Wu SY, et al. Natural killer cells in cancer biology and therapy. Mol Cancer. 2020;19(1):120. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 34.Seymour F, et al. NK cells CD56bright and CD56dim subset cytokine loss and exhaustion is associated with impaired survival in myeloma. Blood Adv. 2022;6(17):5152–9. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 35.Fauriat C, et al. Regulation of human NK-cell cytokine and chemokine production by target cell recognition. Blood. 2010;115(11):2167–76. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 36.Fehniger TA, et al. CD56bright natural killer cells are present in human lymph nodes and are activated by T cell–derived IL-2: a potential new link between adaptive and innate immunity. Blood. 2003;101(8):3052–7. [DOI] [PubMed] [Google Scholar]
- 37.Dalbeth N, et al. CD56bright NK cells are enriched at inflammatory sites and can engage with monocytes in a reciprocal program of activation. J Immunol. 2004;173(10):6418–26. [DOI] [PubMed] [Google Scholar]
- 38.Marcenaro E, et al. Editorial: NK Cell Subsets in Health and Disease: New Developments. Front Immunol. 2017;8:1363. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 39.Nakamura H, et al. Elevation of plasma thioredoxin levels in HIV-infected individuals. Int Immunol. 1996;8(4):603–11. [DOI] [PubMed] [Google Scholar]
- 40.Hoshino T, et al. Redox-active protein thioredoxin prevents proinflammatory cytokine- or bleomycin-induced lung injury. Am J Respir Crit Care Med. 2003;168(9):1075–83. [DOI] [PubMed] [Google Scholar]
- 41.Qi AQ, et al. Thioredoxin is a novel diagnostic and prognostic marker in patients with ischemic stroke. Free Radic Biol Med. 2015;80:129–35. [DOI] [PubMed] [Google Scholar]
- 42.Oraby MI, Rabie RA. Blood biomarkers for stroke: the role of thioredoxin in diagnosis and prognosis of acute ischemic stroke. Egypt J Neurol Psychiatry Neurosurg. 2019;56(1):1. [Google Scholar]
- 43.Lin W, et al. Role of Calcium Signaling Pathway-Related Gene Regulatory Networks in Ischemic Stroke Based on Multiple WGCNA and Single-Cell Analysis. Oxid Med Cell Longev. 2021;2021:8060477. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 44.Lorente L, et al. Serum Levels of B-cell Lymphoma-2 Anti-Apoptotic Protein and Malignant Middle Cerebral Artery Infarction Mortality. J Stroke Cerebrovasc Dis. 2021;30(5):105717. [DOI] [PubMed] [Google Scholar]
- 45.Cruz-Muñoz W, et al. Roles for endothelin receptor B and BCL2A1 in spontaneous CNS metastasis of melanoma. Cancer Res. 2012;72(19):4909–19. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 46.Byron MJ, et al. Immunoglobulin J chain as a non-invasive indicator of pregnancy in the cheetah (Acinonyx jubatus). PLoS One. 2020;15(2):e0225354. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 47.Janigro D, et al. GFAP and S100B: What You Always Wanted to Know and Never Dared to Ask. Front Neurol. 2022;13:835597. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 48.Li Q, et al. Multi-Level Biomarkers for Early Diagnosis of Ischaemic Stroke: A Systematic Review and Meta-Analysis. International Journal of Molecular Sciences. 2023;24(18):13821. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 49.Jauch EC, et al. Association of serial biochemical markers with acute ischemic stroke: the National Institute of Neurological Disorders and Stroke recombinant tissue plasminogen activator Stroke Study. Stroke. 2006;37(10):2508–13. [DOI] [PubMed] [Google Scholar]
- 50.Amini H, et al. Early peripheral blood gene expression associated with good and poor 90-day ischemic stroke outcomes. J Neuroinflammation. 2023;20(1):13. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 51.Park SY, et al. Plasma heart-type fatty acid binding protein level in acute ischemic stroke: comparative analysis with plasma S100B level for diagnosis of stroke and prediction of long-term clinical outcome. Clin Neurol Neurosurg. 2013;115(4):405–10. [DOI] [PubMed] [Google Scholar]
- 52.Huang Z, Sun A. Metabolism, inflammation, and cardiovascular diseases from basic research to clinical practice. Cardiol Plus. 2023;8(1):4–5. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 53.Browne, D.J., et al., A high-throughput screening RT-qPCR assay for quantifying surrogate markers of immunity from PBMCs. Front Immunol, 2022. Volume 13 - 2022. [DOI] [PMC free article] [PubMed]
- 54.Fugate JE, Rabinstein AA. Absolute and Relative Contraindications to IV rt-PA for Acute Ischemic Stroke. Neurohospitalist. 2015;5(3):110–21. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 55.Iglesias-Rey R, et al. Worse Outcome in Stroke Patients Treated with rt-PA Without Early Reperfusion: Associated Factors. Transl Stroke Res. 2018;9(4):347–55. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 56.Laufs U, et al. Neuroprotection mediated by changes in the endothelial actin cytoskeleton. J Clin Invest. 2000;106(1):15–24. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 57.Lee D, Song M, Kwon S. Enhanced Natural Killer Cell Proliferation by Stress-Induced Feeder Cells. Biotechnol Bioeng. 2025; [DOI] [PubMed]
- 58.Endres M, et al. Neuroprotective effects of gelsolin during murine stroke. J Clin Invest. 1999;103(3):347–54. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 59.Wang G, et al. NMMHC IIA triggers neuronal autophagic cell death by promoting F-actin-dependent ATG9A trafficking in cerebral ischemia/reperfusion. Cell Death Dis. 2020;11(6):428. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 60.Mentor S, Makhathini KB, Fisher D. The Role of Cytoskeletal Proteins in the Formation of a Functional In Vitro Blood-Brain Barrier Model. Int J Mol Sci. 2022;23(2) [DOI] [PMC free article] [PubMed]
- 61.Trendowski M, et al. Chemotherapy with cytochalasin congeners in vitro and in vivo against murine models. Invest New Drugs. 2015;33(2):290–9. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 62.Walling EA, Krafft GA, Ware BR. Actin assembly activity of cytochalasins and cytochalasin analogs assayed using fluorescence photobleaching recovery. Arch Biochem Biophys. 1988;264(1):321–32. [DOI] [PubMed] [Google Scholar]
- 63.Go YM, Orr M, Jones DP. Increased nuclear thioredoxin-1 potentiates cadmium-induced cytotoxicity. Toxicol Sci. 2013;131(1):84–94. [DOI] [PMC free article] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
(PDF 1399 kb)
(XLSX 9 kb)
(XLSX 9 kb)
(XLSX 33 kb)
(XLSX 26 kb)
(XLSX 9 kb)
(XLSX 10 kb)
(XLSX 21 kb)
Data Availability Statement
The RNA-seq data have been deposited in the Genome Sequence Archive (GSA) at the National Genomics Data Center, China National Center for Bioinformation/Beijing Institute of Genomics, Chinese Academy of Sciences, and are publicly accessible at: https://ngdc.cncb.ac.cn/gsa-human/browse/HRA001807. Publicly available datasets used in this study were obtained from the NCBI Gene Expression Omnibus (GEO) under accession number: https://www.ncbi.nlm.nih.gov/geo/query/acc.cgi?acc=GSE58294. In addition, the analysis codes used in this study are available at the following GitHub repository: https://github.com/Bon-jour/identification-of-blood-based-biomarkers-for-ischemic-stroke.git.







