Abstract
Purpose
This study investigated the involvement of disulfidptosis in the pathophysiology of sepsis by applying a bioinformatics analysis assisted by large language models (LLMs).
Methods
Based on DeepSeek R1 and retrieval-augmented generation technology, a deep retrieval architecture was developed for extracting disulfidptosis-related genes. An intersection of genes from LLM extraction, manual extraction, and datasets was included for bioinformatics analyses. Using DeepSeek R1, we synthesized a multi-step bioinformatics protocol from prior publications. The analyses were then performed according to the protocol. Key gene candidates were identified using multiple machine learning models, and validation was performed in a cecal ligation and puncture mouse model of sepsis.
Results
A total of 21 disulfidptosis-related genes were included for bioinformatics analyses. Nine bioinformatics techniques were integrated based on the LLM summarization of two key references. Thirteen disulfidptosis-related differentially expressed genes (DEGs) were identified in sepsis. Based on these DEGs, sepsis patients were classified into two molecular subgroups with distinct immune profiles. Among the machine learning models evaluated, the support vector machine achieved the highest classification performance (AUC = 0.989). Five hub genes—FSTL1, SELP, PPBP, ITGA2B, and PF4—were selected as key biomarkers. Experimental validation confirmed significantly elevated expression of these genes in the sepsis mice compared with their sham counterparts.
Conclusion
Our study assisted bioinformatics analysis with large language models and revealed a critical role for disulfidptosis in sepsis. A high-performance diagnostic model was developed, and five genes were validated as potential biomarkers for the diagnosis and treatment of sepsis.
Supplementary Information
The online version contains supplementary material available at 10.1007/s13755-025-00385-z.
Keywords: Large language model, Disulfidptosis, Sepsis, Machine learning, Diagnosis model
Introduction
Sepsis is characterized as a life-threatening condition marked by organ dysfunction resulting from an abnormal and exaggerated immune response to infection [1]. During 2017, an estimated 48.9 million people worldwide developed sepsis, resulting in 11 million fatalities [2]. China bears a particularly high sepsis burden, with hospital prevalence and mortality rates of 3.8% and 26%, respectively, and the proportion of patients admitted to the ICU can reach 25.5% and 40% [3]. However, the early clinical manifestations of sepsis are often nonspecific and heterogeneous, complicating timely diagnosis and contributing to high mortality. Nowadays, early and accurate diagnostic tools remain a pressing need in clinical practice all around the world.
Programmed cell death (PCD) is critically involved in the pathophysiology of sepsis [4]. Disulfidptosis is a newly characterized form of PCD, triggered by abnormal intracellular disulfide accumulation under glucose-deprived conditions, particularly in cells highly expressing SLC7A11 [5, 6]. This accumulation induces disulfide stress in actin cytoskeletal proteins, leading to excessive disulfide bond formation, cytoskeletal collapse, and cell death ultimately [6, 7]. Although disulfidptosis has been primarily studied in cancer [8, 9], its role in sepsis remains poorly defined. However, emerging evidence suggests the relevance—Zhang et al. identified two disulfidptosis-associated genes, ACSL4 and MYL6, as critical mediators in sepsis-induced acute lung injury [10], highlighting a potential mechanistic role for disulfidptosis in sepsis.
Large language models (LLMs) have recently gained traction in medical research and healthcare for their capacity to synthesize knowledge and automate analytical tasks [11]. While most LLM applications to date have focused on medical knowledge retrieval, their potential for workflow generation and protocol optimization remains underexplored [12]. Among natural language processing and natural language understanding tasks, summarization tasks account for only 8.9%.
In this study, we developed a deep retrieval architecture based on DeepSeek R1 to assist us in extracting disulfidptosis-related genes (DRGs) from previous studies. Besides, we employed DeepSeek R1 to derive a structured bioinformatics protocol from two previously published studies [13, 14]. By applying this protocol, we conducted a comprehensive analysis of DRGs in sepsis. We compared gene expression between septic patients and healthy controls, performed clustering and immune profiling, developed diagnostic models using machine learning, and validated our findings in a murine model of sepsis. Our approach demonstrates how LLM-guided bioinformatics can accelerate the process of discovering and refining biomarker identification in complex diseases such as sepsis.
Materials and methods
Large language model architecture
This architecture is designed to achieve precise extraction of DEGs in sepsis, empowering medical research. The workflow begins when users upload sepsis literature and submit task instructions through the interaction layer. The system access layer then handles protocol conversion, permission control, and flow scheduling for these requests. Subsequently, the capability support and research service layer executes literature preprocessing and orchestrates task workflows based on large models and specialized service modules.
The specialized service modules work collaboratively to enhance extraction accuracy, while the quality assurance layer establishes a closed-loop quality control system through process traceability auditing and gene origin verification. Ultimately, the research output layer delivers three types of results: standardized gene sets, accuracy validation reports, and gene analysis reports. Through multi-module collaboration and closed-loop management, this architecture effectively addresses the inefficiency and accuracy limitations inherent in traditional gene extraction methods.
Literature preprocessing
The adapted LLM optimizes rules for extracting textual information from various document formats such as PDF, XML, and Word. For example, for PDFs, the LLM identifies structures such as the header, footer, and chart nested text and develops rules like ‘preserving body paragraph logic and extracting chart title + critical labeling text’. When processing XML, the LLM parses the label hierarchy and extracts a plain text stream according to the rules of ‘ < Text Content > Text Priority under Labels and Associated < Gene-related > Label Supplement’.
In case of blurred PDF scanning text and confusing XML custom labels, the rules are automatically adjusted by comparing historical success cases with the LLM. If the link of ‘semantic complement after OCR identification’ is added for low-definition PDF, the complex format is continuously adapted to ensure the quality of text conversion.
Intent identification and preliminary screening of genes
The LLM relies on prompt words (such as ‘extracting gene name and function description in human lung cancer research literature’), turns on the ‘semantic detector’ mode, disassembles subtasks such as ‘recognizing gene symbols (such as EGFR) and grasping function description statements (such as ‘regulating cell proliferation signaling pathways’)’, traverses the preprocessed text, and marks gene entities and association descriptions based on gene knowledge. For example, the LLM accurately extracted the gene name ‘EGFR’ and the description ‘overexpression in lung cancer cells and promoting tumor angiogenesis’ from ‘EGFR gene overexpression in lung cancer cells and promoting tumor angiogenesis’ to complete the primary gene screening.
Multiple rounds of enhancement with retrieval-augmented generation technology
A closed loop of ‘knowledge search-validation-supplement’ was constructed using retrieval-augmented generation (RAG) technology (Supplementary Fig. S1), with primary screening genes (such as ‘EGFR’) as search terms. The GEO database was called to match official functional annotations with study conclusions: new paper descriptions were grasped for cutting-edge content. If the returned information of the knowledge base is different from the primary screening, the LLM initiates the ‘conflict check’ and re-analyzes whether the original context judgment is an expression difference or misidentification. If it is an omission, the knowledge will be supplemented, and the genetic information will be continuously iterated until the results meet the requirements.
Result integration
Finally, the LLM organizes gene names, multi-dimensional descriptions, literature sources, and other information according to preset templates (such as JSON format: ‘gene name’: ‘EGFR’, ‘functional description’: ‘regulation of cell proliferation signaling pathways, involved in lung cancer angiogenesis’, ‘literature source’: ‘XXX research paper’) and unifies field format.
Data collection and preprocessing
The keyword ‘sepsis’ was used to search the GEO database for gene expression datasets containing sepsis patients and healthy controls. The dataset GSE65682 (GPL13667 platform, peripheral blood) was downloaded for primary analysis, comprising 760 sepsis samples and 42 healthy samples. For validation, the GSE185263 dataset (GPL13607 platform, salivary gland) was utilized, which included 348 sepsis samples and 44 healthy samples. All data were log2-transformed before analysis.
Extraction of disulfidptosis-related genes
Given that most studies on disulfidptosis are tumor-related and that the lung is one of the most important target organs of sepsis, we chose lung adenocarcinoma as the disease type to test the ability of gene extraction by our adapted LLM. We searched PubMed for articles containing ‘disulfidptosis’ and ‘lung adenocarcinoma’ in the title, limiting the time to 2024 and before, obtaining 33 articles. The first article on the PubMed list was downloaded for LLM training and manual extraction of fields [15]. Manual extraction was performed by two independent reviewers, and the distractions were negotiated by consulting with a senior reviewer. For LLM extraction, we let the large model learn the full text of the first literature as well as automatically grab relevant abstracts in PubMed before the output of DRGs. The gene set used for bioinformatics analyses was the intersection of genes from manual extraction, LLM extraction, and the GSE datasets.
Development of the analysis protocol
All bioinformatics analyses were guided by protocols derived from DeepSeek R1. Two peer-reviewed articles involving bioinformatic workflows were uploaded into DeepSeek R1, accompanied by the query ‘summarize the bioinformatics approaches used in the two articles’. The bioinformatic analyses were performed following the protocols generated by the model.
Bioinformatics analyses
The detailed methods of bioinformatics analyses were similar to the process described in previous studies [13, 14]. In brief, unsupervised clustering was performed on the sepsis samples from dataset GSE65682 based on the expression of DRGs. Immune cell proportions were estimated using the CIBERSORT algorithm, and immune cell distributions were visualized. Spearman correlation coefficients were computed to assess associations between DRG expression levels and immune cell fractions. Weighted gene co-expression network analysis (WGCNA) was applied to the top 25% most variable genes to identify co-expressed gene modules.
Gene set variation analysis (GSVA) was performed to show pathway enrichment in the two clusters of sepsis patients. The KEGG pathway gene sets were obtained to serve as the reference dataset. Pathway enrichment scores were computed using the ssGSEA algorithm implemented in the R GSVA package. These GSVA scores, representing the absolute enrichment level of each gene set, were subsequently compared between the two clusters using the limma package.
The predictive genes used for model construction were identified by the intersection of hub genes and DEGs in sepsis patients and healthy controls. Four machine learning approaches—random forest (RF), support vector machine (SVM), generalized linear models (GLMs), and eXtreme Gradient Boosting (XGB)—were employed to identify key predictors of sepsis. The diagnostic performance of the models was analyzed using the receiver operating characteristic (ROC) curves, with the area under curve (AUC) serving as the evaluation metric. Model performances were further validated using the GSE185263 dataset.
Establishment of a mouse sepsis model
Six- to eight-week-old male C57BL/6 mice were randomly allocated to either undergo cecal ligation and puncture (CLP) or serve as sham counterparts. The CLP model was established according to well-characterized protocols from prior studies [16]. In brief, the cecum of mice in the CLP group was ligated at the midpoint (50%) and punctured once with a 21G needle. Sham-operated mice underwent laparotomy and wound closure without cecal ligation or puncture. Mice from both groups were euthanized 24 h post-procedure, and tissues from the heart, liver, lung, and kidney were harvested for subsequent analyses. All animal procedures received approval from the Institutional Review Board of the Fourth Medical Center of PLA General Hospital.
Histological analysis
The tissues from sepsis and sham mice were fixed in 4% paraformaldehyde for at least 24 h before paraffin embedding. Tissue sections were cut from paraffin-embedded tissues and routinely stained with hematoxylin and eosin for pathological evaluation. Following dehydration and sealing, stained slides were digitally scanned, and representative injury regions were captured for histological evaluation.
Quantitative real-time PCR
The expression levels of SLC7A11 (a core gene of disulfidptosis) and five other key genes identified in our study were assessed using quantitative real-time PCR. We extracted total RNA with an RNA extraction kit (CWBIO, China) and performed reverse transcription with the HiScript III 1st Strand cDNA Synthesis Kit (Vazyme, China). SYBR Green-based qPCR was conducted using AceQ Universal SYBR qPCR Master Mix (Vazyme, China).
Statistical analysis
Bioinformatics analyses were performed using R software (v4.3.1), while experimental data from animal studies were statistically analyzed with GraphPad Prism (v9.5.1). We considered results statistically significant with a p-value less than 0.05 and used the following asterisk notation: *p < 0.05, **p < 0.01, ***p < 0.001, ****p < 0.0001.
Results
Bioinformatic analysis protocol summarized by DeepSeek R1
To establish a comprehensive bioinformatics framework, we utilized DeepSeek R1 to extract and summarize the analytical methodologies from two selected studies (PMID: 37984252 and 39051056), which investigated the roles of PANoptosis and Cuproptosis in sepsis and primary Sjögren’s syndrome, respectively. DeepSeek R1 identified nine key analytic strategies: differential expression analysis, immune cell infiltration analysis, unsupervised clustering, weighted gene co-expression network analysis, functional enrichment analysis, machine learning-based model development, nomogram and survival analysis, external dataset validation, and experimental validation. For each method, a concise methodological description was provided, followed by an integrated summary, offering a structured reference for our bioinformatic approach. An overview of the bioinformatic analyses was presented in Fig. 1.
Fig. 1.
Flowchart illustrating the bioinformatics analyses
Gene extraction assisted by a large language model
The schema of our adapted large language model is shown in Fig. 2. We selected six disulfidptosis-related articles to test the time of manual extraction and LLM extraction of DEGs. The result showed that LLM extraction was significantly faster than manual extraction (37.88 ± 10.24 s vs. 239.2 ± 74.84 s, p < 0.0001). With the help of the LLM, 23 DRGs were extracted from the given literature and abstracts. The manually extracted genes were completely consistent with the genes extracted from the LLM. After intersection with genes from the GSE65682 and GSE185263 datasets, 21 genes were finally included for further analysis (Fig. 3a).
Fig. 2.
Schema of the large language model
Fig. 3.
Identification of differentially expressed DRGs in sepsis. a Identification of 21 DRGs by the intersection of manually extracted genes, LLM extracted genes, and genes in the dataset. b Boxplot of DRG expression between sepsis patients and healthy individuals. c Heatmap of the 13 DEGs between sepsis patients and healthy individuals. d Chromosomal localization of the DEGs. e Chord plot shows the gene–gene correlation network
Differential expression and immune infiltration of DRGs in sepsis
Analysis of the GSE65682 transcriptomic dataset revealed 13 DEGs among 21 examined DRGs when comparing septic patients and healthy individuals. Among them, SLC7A11, PDLIM1, TLN1, MYL6, ACTB, GYS1, NCKAP1, INF2, and DSTN were significantly upregulated in sepsis, while NDUFA11, RPN1, NDUFS1, and LRPPRC were downregulated (Fig. 3b, c). Chromosomal locations of these DEGs are shown in Fig. 3d, and their pairwise correlations are illustrated in Fig. 3e, with line thickness representing correlation strength (positive in red, negative in green).
Identification of sepsis subtypes based on DRG expression
Unsupervised clustering based on consensus cluster revealed two distinct molecular subtypes (k = 2) based on DRG expression (Fig. 4a). This classification showed optimal stability, with minimal cumulative distribution function (CDF) curve fluctuation (Fig. 4b), the highest consensus score (> 0.9) (Fig. 4c), and well-separated clusters according to principal component analysis (PCA) (Fig. 4d).
Fig. 4.
Identification of DRG-based sepsis subtypes. a Consensus clustering matrix at k = 2. b Consensus CDF curves with k = 2–9. c Bar plot demonstrates the consensus score among different key parameters. d Distinct distributions of cluster C1 and C2 in the PCA plot
Immune infiltration profiles of DRG-based sepsis subtypes
DRG expression profiles revealed subtype-specific signatures: NDUFA11 and LRPPRC were elevated in Cluster 1, while PDLIM1, TLN1, MYL6, ACTB, GYS1, NDUFA11, NCKAP1, RPN1, INF2, and DSTN were significantly upregulated in Cluster 2 (Fig. 5a, b). Immune profiling based on Cibersort demonstrated cluster-specific infiltration patterns across 12 cell types spanning adaptive (B/T cell subsets) and innate (macrophages, NK cells, neutrophils) immunity (Fig. 5c, d). We found that Cluster 2 showed preferential enrichment of immune system pathways, while Cluster 1 exhibited stronger metabolic pathway activity through GSVA (Fig. 5e).
Fig. 5.
Immune and molecular features of DRG-based clusters. a Boxplot of DRGs expression between two sepsis clusters. b Heatmap of DRGs expression between two sepsis clusters. c Relative distribution of 22 immune infiltrating cell types (IICs) analyzed by immune cell infiltration analysis. d Boxplots comparing IIC populations between two sepsis clusters. e GSVA based on KEGG showing pathway enrichment: immune activity in C2, metabolic activity in C1
Gene module identification and co-expression network construction
Using WGCNA, we constructed signed gene co-expression networks to identify gene modules significantly associated with DRG-defined sepsis subtypes. A soft threshold of 5 was selected, as it achieved a satisfactory scale-free topology fit (R2 > 0.85) and optimized module detection (Fig. 6a). Dynamic tree cutting revealed eight co-expression modules, visualized through gene dendrograms and TOM heatmaps (Fig. 6b–d). Among these modules, the black one showed the strongest correlation with the DRG-based sepsis subtypes (Fig. 6e, f). Given the heightened immune activity in Cluster 2, we selected the 100 genes from the black module in this cluster for further exploration. Fifty-two hub genes were subsequently identified based on thresholds of |MM|> 0.8 and |GS|> 0.5, and these were prioritized for downstream analysis.
Fig. 6.
WGCNA-based co-expression analysis. a Determination of the optimal soft threshold. b Gene dendrogram with module color assignments. c Feature gene clustering tree. d TOM heatmap visualizing module correlations. e Association between module eigengenes and clinical traits. f Scatter plot analyzing gene significance and module membership in the black module
Construction and evaluation of machine learning models
A total of 2,510 DEGs were identified between healthy controls and sepsis patients in dataset GSE65682 (Fig. 7a, b). Immune-related biological processes, particularly leukocyte activation, cytokine production, lymphocyte differentiation, and mononuclear cell differentiation, emerged as the most significantly enriched gene ontology (GO) terms among the DEGs (Fig. 7c). Kyoto Encyclopedia of Genes and Genomes (KEGG) pathway analysis further revealed that these genes participated in key immune pathways, notably those linked to coronavirus disease (COVID-19), Th1 and Th2 cell differentiation, efferocytosis, and T cell receptor signaling (Fig. 7d).
Fig. 7.
Differential gene expression between sepsis and control samples. a Volcano plot of DEGs. b Heatmap displaying each of the 50 most significant genes from both up- and down-regulated groups. c Functional annotation of cluster-related DEGs through GO enrichment analysis. d KEGG pathway enrichment profile of DEGs associated with the identified cluster
By intersecting the 52 hub genes from disulfidptosis-associated co-expression modules with the 2,510 DEGs, we identified 30 characteristic Disulfidptosis-related DEGs specific to sepsis (Fig. 8a). Based on these 30 genes, four machine learning models (RF, SVM, GLM, and XGB) were constructed. Residual distributions for each model were analyzed using reverse cumulative distribution and boxplots, providing an overview of model errors (Fig. 8b, c). Ranking the models by Root Mean Square Error (RMSE) allowed us to identify the ten most influential variables for each algorithm (Fig. 8d). ROC analysis showed the SVM model outperformed others, exhibiting the highest accuracy, lowest residuals, and superior AUC (Fig. 8e).
Fig. 8.
Construction and performance evaluation of machine learning models. a Intersection of DEGs and WGCNA hub genes. b Distribution of model residuals across four prediction models. c Distribution of model residuals across prediction methods. d Consensus feature importance ranking from multiple algorithms. e ROC analysis of model discrimination performance in the testing cohort. f External validation of candidate biomarkers in the GSE185263 dataset
To validate the generalizability of the SVM model, we tested five key DSP-related genes—FSTL1, SELP, PPBP, ITGA2B, and PF4—using an external dataset (GSE185263). The resulting ROC analysis (AUC = 0.871) confirmed the model’s robust discriminative capability between sepsis and control samples (Fig. 8f). Collectively, these findings underscore the diagnostic potential of our machine learning framework for identifying sepsis based on DSP gene expression.
Nomogram construction and survival analysis
To further assess the prognostic utility of the five key genes in sepsis, we developed a nomogram model (Fig. 9a). The calibration curve demonstrated high agreement between predicted and observed outcomes, indicating low prediction error (Fig. 9b). Decision curve analysis (DCA) confirmed the clinical value of our model, suggesting strong net benefit across a range of threshold probabilities (Fig. 9c). Correlation analysis demonstrated significant positive associations between five candidate genes and immune cell populations, particularly plasma cells and macrophage subsets M0 and M1, highlighting their potential immunological relevance in the pathogenesis of sepsis (Fig. 10).
Fig. 9.
Nomogram-based prediction of sepsis risk using five DSP-related genes. a Nomogram showing relationships of DRGs and risk. b High agreement between predicted and observed outcomes is revealed by the calibration curve. c DCA curves supported the clinical utility of the prediction model
Fig. 10.
Correlations between the five key model genes and immune cell infiltration. Lollipop plots showed the correlation patterns between model genes and 22 immune cell subtypes: a FSTL1, b ITGA2B, c PF4, d PPBP, and e SELP
Expression of DSP-related genes in a mouse sepsis model
Histopathological examination confirmed sepsis-induced tissue damage in key target organs—heart, liver, lung, and kidney—in the CLP mouse model (Fig. 11a). PCR results demonstrated comparable mRNA levels of FSTL1 and ITGA2B in all the tissues from septic mice and sham mice. However, PF4 was significantly upregulated in both hepatic and pulmonary tissues of septic mice, while SELP was elevated in the liver, lung, and kidney (Fig. 11b). Notably, PPBP expression was consistently increased across all analyzed tissues in septic animals, supporting its potential as a systemic biomarker.
Fig. 11.
Validation of DSP-related genes in a murine sepsis model. a H&E staining of sliced tissues from septic and sham mice. b PCR quantification of gene expression in each tissue
Discussion
Disulfidptosis is a newly identified type of programmed cell death that differs mechanistically from known pathways such as apoptosis, ferroptosis, and others. It arises from disulfide stress triggered by intracellular cystine overload. In 2020, Gan et al. discovered that the reduction of imported cystine to cysteine, mediated by SLC7A11, depends heavily on NADPH generated through the glucose-pentose phosphate pathway [17]. Under glucose-deprived conditions, NADPH levels rapidly decline, leading to the accumulation of disulfides, particularly cystine, in cells overexpressing SLC7A11. This imbalance induces lethal disulfide stress and rapid cell death. In 2023, Gan’s team formally defined this phenomenon as a distinct form of cell death and termed it disulfidptosis [18].
In the present study, we developed a deep retrieval architecture based on DeepSeek R1 for gene extraction and annotation from previous studies—a novel approach in medical research. The genes identified by the LLM were completely consistent with those extracted manually, indicating that the LLM can be used as a new method to assist scientific researchers in gene screening and greatly save time in scientific research work. Among the 23 genes identified by the LLM in this study, ACTB, FLNB, NCKAP1, SLC3A2, and SLC7A11 were identified as model genes. They played a clear role in chromosome localization, cell expression type, and prognosis association with lung adenocarcinoma. The other 18 genes had expression differences between lung adenocarcinoma and normal tissues but were not included in the model. However, the genes identified by LLM were not consistent with key genes generated by bioinformatic analysis, suggesting that LLM still cannot replace traditional analytical methods.
Besides, we also used DeepSeek R1 to generate the workflows of bioinformatic workflows. The use of DeepSeek not only provided a clear research framework but also offered detailed methodological guidance. As LLMs continue to evolve, they may enable researchers without technical backgrounds to independently perform fundamental bioinformatics analyses, significantly reducing reliance on bioinformatics specialists and streamlining interdisciplinary workflows.
To investigate the link between sepsis and disulfidptosis, we examined the expression patterns of DRGs using publicly available transcriptomic data from the GEO database. Comparative analysis of peripheral blood samples from septic patients and healthy controls identified 13 DRGs with significant expression changes. Subsequent unsupervised clustering identified three distinct disulfidptosis-related molecular subtypes. These clusters were further validated through PCA, confirming their relevance for sepsis classification.
To develop a robust diagnostic tool, we constructed four prediction models and found the SVM model exhibited the best performance, achieving an AUROC of 0.966 in dataset GSE65682 and 0.871 in dataset GSE185263. Five key DRGs closely related to sepsis, including FSTL1, SELP, PPBP, ITGA2B, and PF4, were obtained from the SVM model. A diagnostic nomogram incorporating these genes was subsequently developed and validated. The model demonstrated excellent diagnostic utility, with a C-index of 0.988. Decision curve and clinical impact analyses further supported the nomogram’s potential for clinical application. Additionally, correlation analyses revealed strong associations between five candidate genes and levels of immune cell infiltration in sepsis.
Sepsis-3.0 criteria define sepsis as a potentially fatal condition marked by organ dysfunction arising from the host’s maladaptive response to infection [19]. The pathophysiology of sepsis involves complex cellular and metabolic disturbances, including multiple forms of cell death. Our findings reinforce this concept by demonstrating differential expression of DRGs in sepsis, with distinct upregulation and downregulation patterns that likely reflect shifts in redox balance and mitochondrial function.
Immune dysregulation in sepsis is dynamic and multifaceted. Depending on disease stage, patients may exhibit either hyperinflammation or profound immunosuppression, both of which disrupt immune homeostasis. Immune infiltration analysis demonstrated elevations in multiple lymphocyte populations in septic patients relative to healthy controls, including memory B cells, regulatory T cells (Tregs), γδ T cells, and activated dendritic cells. Memory B cells, formed in germinal centers following initial antigen exposure, play a crucial role in rapid secondary immune responses. However, memory B cell depletion contributes to sepsis-induced immunosuppression, indicating that the function of these cells may vary with the progressive change of sepsis. Similarly, an increase in Tregs, known for their immunosuppressive roles, further suggests a shift toward an immunosuppressive state in septic patients. Two sepsis subtypes were identified based on DSP-related gene expression. According to the GSVA results, one subtype showed enrichment of immune system pathways, and the other subtype showed stronger metabolic pathway activity. This result suggested the immune status may differ between different septic patients or at different periods of sepsis, and these findings should be deeply investigated in future studies.
Early diagnosis and timely intervention are critical to improving outcomes in septic patients. While prior studies have identified disulfidptosis-related genes and constructed predictive models for sepsis, our study builds on this foundation with a comprehensive machine learning approach. Zhang et al. identified MYL6 and ACSL4 as key markers of sepsis-induced acute lung injury, whereas Zou et al. pinpointed five hub genes, including MYH10, FLNA, ACTN4, MYH9, and IQGAP1. Through comparative analysis of two predictive models, He et al. discovered six intersecting genes—LRPPRC, SLC7A11, GLUT, MYH9, NUBPL, and GYS1. Expanding on these findings, we developed four machine learning models and found that the SVM algorithm yielded the highest predictive accuracy. SVM classifies data by constructing a maximum-margin hyperplane and efficiently handles nonlinear data through kernel functions. Its strengths lie in its robustness in high-dimensional spaces, strong generalizability with small datasets, and flexibility in modeling nonlinear relationships.
Using the SVM model, we identified five disulfidptosis-related genes—FSTL1, PPBP, PF4, SELP, and ITGA2B—as key predictors of sepsis. These genes were validated using an independent external dataset. To further assess their clinical utility, we constructed a nomogram for sepsis risk prediction based on these five genes. The model demonstrated excellent predictive performance and holds promise for clinical translation. Beyond predictive value, these genes may also serve as potential therapeutic targets in sepsis management.
FSTL1 encodes Follistatin-like 1, a secreted glycoprotein involved in immune modulation, inflammation, and tissue repair [20]. Elevated serum levels of FSTL1 have been reported in sepsis and proposed as a potential biomarker for systemic inflammation [21]. PPBP encodes CXCL7, a chemokine released from activated platelets that contributes to inflammation and tissue regeneration [22]. PF4 (CXCL4) is another platelet-derived chemokine [23] involved in coagulation, immune regulation, and angiogenesis, particularly under inflammatory conditions [24]. SELP encodes P-selectin, a key cell adhesion molecule implicated in thrombosis, leukocyte recruitment, and vascular inflammation [25, 26]. ITGA2B encodes the integrin αIIb subunit (CD41), which pairs with integrin β3 (CD61) to form the GPIIb/IIIa complex, a central regulator of platelet aggregation, thrombosis, angiogenesis, and tumor metastasis [27].
In our murine sepsis model, we observed the upregulation of PPBP, PF4, and SELP, all of which are platelet-associated genes. This suggests a pivotal role for platelet dysfunction in sepsis pathophysiology. These findings imply that disulfidptosis-related genes may modulate the inflammatory response through effects on platelet activation and vascular endothelial permeability. However, the precise molecular mechanisms remain to be elucidated in future studies.
Correlation analysis revealed strong associations between the five key DRGs and immune cell infiltration, particularly macrophage polarization, which is a known determinant of sepsis progression. Macrophages exhibit dynamic phenotypic plasticity during sepsis. In the early hyperinflammatory phase, M1 macrophages dominate, producing pro-inflammatory cytokines and mediating pathogen clearance. While this response is essential for controlling infection, excessive M1 activation can trigger cytokine storms and organ failure. As sepsis progresses, M2 macrophages become more prevalent, promoting tissue repair through anti-inflammatory cytokine secretion and suppression of T cell responses. However, excessive M2 polarization contributes to immune paralysis, heightening susceptibility to secondary infections and increasing mortality.
While our study provides valuable insights, certain limitations merit consideration. The bioinformatics analyses were based on publicly available datasets, and prospective clinical studies are needed to validate our findings. Furthermore, tissue-specific expression differences between peripheral blood and other compartments, such as salivary glands, may reflect distinct microenvironments. Functional validation in animal models will be essential to clarify the roles of these genes in vivo in the future.
Conclusion
This study applies the LLMs to bioinformatics analysis and trains a model of literature processing and gene grasping. Our results highlight the involvement of disulfidptosis in sepsis and identify five key predictive genes that demonstrated strong diagnostic performance in both training and validation datasets. The sepsis mouse model confirmed increased expression of key genes, further supporting their relevance. Our findings advance the application of LLMs in medical research and highlight the role of disulfidptosis in sepsis.
Supplementary Information
Below is the link to the electronic supplementary material.
Author contributions
Tian Liu engaged in conceptualization, investigation, data curation, formal analysis, and writing the original draft. Zhi Mao engaged in resources, investigation, validation, and data curation. Jiake Chai engaged in conceptualization and supervision. Hui Zhou engaged in resources and software. Yirui Qu engaged in validation. Chengfeng Xu engaged in methodology. Yunfei Chi engaged in methodology, project administration, and writing (review and editing).
Funding
The authors have not disclosed any funding.
Data availability
Our data can be accessed by contacting the corresponding author.
Declarations
Conflict of interest
All authors declare no competing interests.
Footnotes
Publisher's Note
Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
Tian Liu, and Zhi Mao have equally contributed to this work.
Contributor Information
Jiake Chai, Email: cjk304@126.com.
Yunfei Chi, Email: chi_yunfei@126.com.
References
- 1.Singer M, Deutschman CS, Seymour CW, Shankar-Hari M, Annane D, Bauer M, et al. The third international consensus definitions for sepsis and septic shock (Sepsis-3). JAMA. 2016;315(8):801–10. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 2.Rudd KE, Johnson SC, Agesa KM, Shackelford KA, Tsoi D, Kievlan DR, et al. Global, regional, and national sepsis incidence and mortality, 1990–2017: analysis for the Global Burden of Disease Study. Lancet. 2020;395(10219):200–11. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 3.Lei S, Li X, Zhao H, Xie Y, Li J. Prevalence of sepsis among adults in China: a systematic review and meta-analysis. Front Public Health. 2022;10:977094. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 4.Li JY, Yao RQ, Xie MY, Zhou QY, Zhao PY, Tian YP, et al. Publication trends of research on sepsis and programmed cell death during 2002–2022: a 20-year bibliometric analysis. Front Cell Infect Microbiol. 2022;12:999569. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 5.Zheng T, Liu Q, Xing F, Zeng C, Wang W. Disulfidptosis: a new form of programmed cell death. J Exp Clin Cancer Res. 2023;42(1):137. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 6.Wang X, Zheng C, Yao H, Guo Y, Wang Y, He G, et al. Disulfidptosis: six riddles necessitating solutions. Int J Biol Sci. 2024;20(3):1042–4. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 7.Liu X, Zhuang L, Gan B. Disulfidptosis: disulfide stress-induced cell death. Trends Cell Biol. 2024;34(4):327–37. [DOI] [PubMed] [Google Scholar]
- 8.Zheng P, Zhou C, Ding Y, Duan S. Disulfidptosis: a new target for metabolic cancer therapy. J Exp Clin Cancer Res. 2023;42(1):103. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 9.Xiao F, Li HL, Yang B, Che H, Xu F, Li G, et al. Disulfidptosis: a new type of cell death. Apoptosis. 2024;29(9–10):1309–29. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 10.Zhang A, Wang X, Lin W, Zhu H, Pan J. Identification and verification of disulfidptosis-related genes in sepsis-induced acute lung injury. Front Med (Lausanne). 2024;11:1430252. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 11.Shool S, Adimi S, SabooriAmleshi R, Bitaraf E, Golpira R, Tara M. A systematic review of large language model (LLM) evaluations in clinical medicine. BMC Med Inform Decis Mak. 2025;25(1):117. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 12.Bedi S, Liu Y, Orr-Ewing L, Dash D, Koyejo S, Callahan A, et al. Testing and evaluation of health care applications of large language models: a systematic review. JAMA. 2025;333(4):319–28. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 13.Wei B, Wang A, Liu W, Yue Q, Fan Y, Xue B, et al. Identification of immunological characteristics and cuproptosis-related molecular clusters in primary Sjogren’s syndrome. Int Immunopharmacol. 2024;126:111251. [DOI] [PubMed] [Google Scholar]
- 14.Xu J, Zhu M, Luo P, Gong Y. Machine learning screening and validation of PANoptosis-related gene signatures in sepsis. J Inflamm Res. 2024;17:4765–80. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 15.Ni L, Yang H, Wu X, Zhou K, Wang S. The expression and prognostic value of disulfidptosis progress in lung adenocarcinoma. Aging (Albany NY). 2023;15(15):7741–59. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 16.Wang JF, Wang YP, Xie J, Zhao ZZ, Gupta S, Guo Y, et al. Upregulated PD-L1 delays human neutrophil apoptosis and promotes lung injury in an experimental mouse model of sepsis. Blood. 2021;138(9):806–10. [DOI] [PubMed] [Google Scholar]
- 17.Liu X, Olszewski K, Zhang Y, Lim EW, Shi J, Zhang X, et al. Cystine transporter regulation of pentose phosphate pathway dependency and disulfide stress exposes a targetable metabolic vulnerability in cancer. Nat Cell Biol. 2020;22(4):476–86. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 18.Liu X, Nie L, Zhang Y, Yan Y, Wang C, Colic M, et al. Actin cytoskeleton vulnerability to disulfide stress mediates disulfidptosis. Nat Cell Biol. 2023;25(3):404–14. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 19.Zhu Y, Zhang R, Ye X, Liu H, Wei J. SAPS III is superior to SOFA for predicting 28-day mortality in sepsis patients based on Sepsis 30 criteria. Int J Infect Dis. 2022;114:135–41. [DOI] [PubMed] [Google Scholar]
- 20.Mattiotti A, Prakash S, Barnett P, van den Hoff MJB. Follistatin-like 1 in development and human diseases. Cell Mol Life Sci. 2018;75(13):2339–54. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 21.Chaly Y, Fu Y, Marinov A, Hostager B, Yan W, Campfield B, et al. Follistatin-like protein 1 enhances NLRP3 inflammasome-mediated IL-1beta secretion from monocytes and macrophages. Eur J Immunol. 2014;44(5):1467–79. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 22.Graca FA, Stephan A, Minden-Birkenmaier BA, Shirinifard A, Wang YD, Demontis F, et al. Platelet-derived chemokines promote skeletal muscle regeneration by guiding neutrophil recruitment to injured muscles. Nat Commun. 2023;14(1):2900. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 23.Warkentin TE. Platelet-activating anti-PF4 disorders: an overview. Semin Hematol. 2022;59(2):59–71. [DOI] [PubMed] [Google Scholar]
- 24.Hoeft K, Schaefer GJL, Kim H, Schumacher D, Bleckwehl T, Long Q, et al. Platelet-instructed SPP1(+) macrophages drive myofibroblast activation in fibrosis in a CXCL4-dependent manner. Cell Rep. 2023;42(2):112131. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 25.Purdy M, Obi A, Myers D, Wakefield T. P- and E- selectin in venous thrombosis and non-venous pathologies. J Thromb Haemost. 2022;20(5):1056–66. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 26.Leiter O, Brici D, Fletcher SJ, Yong XLH, Widagdo J, Matigian N, et al. Platelet-derived exerkine CXCL4/platelet factor 4 rejuvenates hippocampal neurogenesis and restores cognitive function in aged mice. Nat Commun. 2023;14(1):4375. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 27.Middleton EA, Rowley JW, Campbell RA, Grissom CK, Brown SM, Beesley SJ, et al. Sepsis alters the transcriptional and translational landscape of human and murine platelets. Blood. 2019;134(12):911–23. [DOI] [PMC free article] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Data Availability Statement
Our data can be accessed by contacting the corresponding author.











