Skip to main content
PeerJ logoLink to PeerJ
. 2026 Feb 13;14:e20770. doi: 10.7717/peerj.20770

Risk prediction models for sepsis-associated encephalopathy: a systematic evaluation and meta-analysis

Ting ting He 1, Tuo quan Jiao 1, Xue mei An 2,3,✉
Editor: Nicole Nogoy
PMCID: PMC12908571  PMID: 41704240

Abstract

Background

The number of risk prediction models for sepsis-associated encephalopathy (SAE) is increasing, while the quality and applicability of these models in clinical practice and future research remain uncertain.

Objective

To systematically review published studies on SAE risk prediction models.

Design

Systematic review and meta-analysis of observational studies.

Methods

A systematic search of PubMed, Web of Science, Embase, Wanfang, VIP, and CNKI databases was conducted from inception to April 2, 2025, to identify studies on SAE risk prediction models. Two independent reviewers screened the studies and extracted data. The Prediction model Risk Of Bias Assessment Tool (PROBAST) was applied to evaluate the risk of bias and applicability of the included studies.

Results

A total of 1,994 studies were identified, and 10 were included after screening. The reported incidence of SAE ranged from 15.16% to 63.3%. Age and Sequential Organ Failure Assessment (SOFA) score are the most frequently adopted factors with significant predictive value, both of which were incorporated into five models. Both the SOFA score and age were significantly associated with SAE. In studies with available data, the odds ratio (OR) for age ranged from 1.084 to 1.018, while that for SOFA score ranged from 1.246 to 2.416. The area under the receiver operating characteristic curve (AUC) for the 10 studies ranged from 0.743 to 0.975. All studies were found to have a high risk of bias, primarily due to inappropriate data sources and deficiencies in the analysis domain. The pooled AUC for the six validated models was 0.83 (95% confidence interval [0.77–0.89]), indicating fair discrimination.

Conclusion

Although the included studies reported some discrimination in the SAE prediction models, all were found to have a high risk of bias according to the PROBAST checklist.

Registration

This study protocol was registered on PROSPERO (registration number: CRD420251012485).

Keywords: Sepsis-Associated Encephalopathy, Meta analysis, Systematic review, Risk, Prediction model

Introduction

Sepsis, defined as life-threatening organ dysfunction caused by a dysregulated host response to infection, remains one of the leading causes of mortality in the intensive care unit (ICU) and continues to pose a substantial global health burden, particularly in low- and middle-income countries (Aljefri et al., 2023; Hotchkiss & Karl, 2003). Sepsis-associated encephalopathy (SAE), a diffuse brain dysfunction occurring in the absence of direct central nervous system infection, affects approximately 8% to 70% of patients with sepsis and is associated with 28-day mortality rates of up to 46% (Lei & Wu, 2025; Li et al., 2025; Mazeraud et al., 2020).

The absence of standardised diagnostic criteria, combined with confounding factors such as sedation and mechanical ventilation, poses major challenges for recognising SAE (Dumbuya et al., 2023; Mazeraud et al., 2020; Shirodkar et al., 2025). Its non-specific symptoms (e.g., delirium, coma) can be misattributed to other encephalopathies, risking misdiagnosis. Delayed or missed diagnosis often contributes to higher mortality and an increased risk of neurological complications. Therefore, early identification of patients at risk for SAE is crucial, as it enables timely interventions to mitigate long-term complications.

Despite the growing number of SAE prediction models, their methodological quality and clinical utility remain uncertain. This study systematically reviews and meta-analyses SAE risk prediction models to evaluate their risk of bias, performance, and applicability, providing an evidence-based foundation for future model development and clinical implementation.

Methods

The study protocol was registered on PROSPERO (registration number: CRD420251012485). Literature screening was conducted on April 2, 2025, and data extraction was conducted on April 10, 2025.

Search strategy

The literature search was conducted in accordance with the Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA) 2020 guidelines and was independently verified by two reviewers (TT and TQ) to ensure the comprehensiveness and reproducibility of the search results (see Table S1 for details). The PubMed, Web of Science, Embase, Wanfang, VIP, and CNKI databases were searched for studies on SAE risk prediction models, with a search timeframe from inception to April 2, 2025, for all databases. The search terms included “sepsis-associated encephalopathy, sepsis-associated psychosis, sepsis encephalopathy prediction, early warning, predictor, influencer, influencing factor, risk assessment, risk prediction, modeling, tool, column-line graph, nomogram”. We also identified other relevant studies through the reference lists of retrieved studies and review articles.

For systematic evaluation, we adopted the PICOTS framework, and the key items of our systematic review are described below:

P (Population): Adult patients (≥18 years old) with sepsis, as defined by the Sepsis 3.0 international consensus criteria (Singer et al., 2016). The setting included Intensive Care Unit (ICU), emergency departments, and general wards. Patients with other primary causes of encephalopathy were excluded from the study.

I (Intervention model): Risk prediction model for SAE.

C (Comparator): No competing model.

O (Outcome): Occurrence of SAE, rather than subgroup outcomes such as sepsis-associated delirium.

T (Timing): Assessed within 24 h of hospital admission.

S (Setting): Intended use was to predict SAE.

Inclusion and exclusion criteria

Inclusion criteria: (1) Study subjects: Adult patients (≥18 years old) diagnosed with sepsis according to the Sepsis 3.0 consensus criteria were included. Specifically, studies were eligible if sepsis was defined as a suspected or documented infection accompanied by an increase in the Sequential Organ Failure Assessment (SOFA) score of ≥2 points (Singer et al., 2016); (2) Study types: cohort studies, case-control studies and cross-sectional studies; (3) Study content: development of a predictive model for the risk of SAE; (4) Outcome indicators: the occurrence of SAE. Exclusion criteria: (1) Animal or cell-based experiments, reviews, and conference papers; (2) Studies that only analyse the predictive capability of influencing factors without constructing SAE prediction models; (3) Not written in English or Chinese.

Given the absence of a universally accepted diagnostic gold standard, studies were included if SAE was defined as brain dysfunction occurring in the context of sepsis, after excluding other causes such as metabolic, structural, or drug-induced encephalopathy. In most included studies, SAE diagnosis was based on a Glasgow Coma Scale (GCS) score <15 and/or the presence of delirium as assessed by the Confusion Assessment Method for the ICU (CAM-ICU).

Literature screening and data extraction

To ensure the study’s objectivity, two reviewers (TT and TQ) independently screened the literature and extracted the data. In case of disagreement, a third party (XM) was asked to discuss or consult to resolve the issue. EndNote was used to manage retrieved citations. Data extraction followed the Critical Appraisal and Data Extraction for Systematic Reviews of Prediction Modelling Studies (CHARMS) checklist (Moons et al., 2014). Extracted items included:

  • (1)

    Basic information: including authors, year of publication, study design, study population, data source and sample size.

  • (2)

    Model information: including handling of missing data, variable selection methods, model development methods, model validation methods, model performance measures, final predictor variables, and model presentation.

Quality evaluation of models

Two independent researchers (TT and TQ) assessed the risk of bias, applicability, and overall quality of the included studies using the Prediction model Risk Of Bias Assessment Tool (PROBAST) (Moons et al., 2019) and the Grading of Recommendations, Assessment, Development, and Evaluation (GRADE) system (Guyatt et al., 2008). In the event of disagreement, a third party (XM) was asked to resolve the issue through discussion or consultation. PROBAST evaluates four domains—participants, predictors, outcomes, and analysis—each rated as low, high, or unclear risk of bias; applicability concerns were judged across the first three domains. GRADE classifies the certainty of evidence as high, moderate, low, or very low, taking into account study design, risk of bias, inconsistency, indirectness, imprecision, and publication bias.

Data synthesis and statistical analysis

We performed a meta-analysis of AUC values from validated models using R. To assess heterogeneity among the included studies, we used the I2 statistic. If I2 < 50% and p > 0.10, it means that the heterogeneity among studies is not significant, and the fixed-effects model was used for the merger; if I2 ≥ 50% or p ≤ 0.1, it means that the heterogeneity among studies is significant, in which case sensitivity analysis will be performed (Lehrer et al., 2023).

After excluding the literature article by article, if heterogeneity persisted, the random-effects model was used for meta-analysis. If the merged effect value did not change significantly, the results of the meta-analysis were deemed stable. Egger’s test (Egger et al., 1997) was used to identify publication bias. p > 0.05 indicated a low probability of publication bias.

Results

Literature screening process and results

A total of 1,994 records were retrieved from all databases. After rigorous screening and eligibility assessment, 10 studies met the inclusion criteria, of which six provided model data suitable for the meta-analysis (Fig. 1).

Figure 1. PRISMA flowchart of the study screening process.

Figure 1

Characteristics of included studies

Table 1 summarises the design and subject characteristics of the 10 included studies. The 10 included studies, published between 2021 and 2024, comprised five based on Chinese cohorts and five using U.S. databases. Eight studies were retrospective, and only two were prospective. Regarding the study population, nine studies enrolled adult patients with sepsis, while one focused on elderly patients (≥65 years). Among the ten included studies, the sample size ranged from 67 to 22,361 participants (median: 1,587.5; IQR: 130–893.7). The reported incidence of SAE across these studies ranged from 15.16% to 63.3% (median: 44.7%).

Table 1. Overview of basic data of the included studies.

Author (year) Country of data source Study design Participants Data source Main outcome SAE cases/ sample size (%)
Wang, Ziwena (2023) China Retrospective cohort study Patients with sepsis admitted to the ICU for the first time ICU of a hospital SAE 97/640 (15.16%)
Zhou, Hangxianga (2023) China Retrospective cohort study Patients with sepsis > 18 years of age ICU of a hospital SAE 84/213 (39.44%)
Zhao, Qing (2023) United States Retrospective cohort study ICU sepsis patients ≥ 65 years of age MIMIC-IV SAE 8,290/22,361 (37.1%)
Xiao, Lu (2022) United States Retrospective cohort study Patients with sepsis ≥ 18 years of age MIMIC-IV SAE 4,684/8,935 (52.4%)
Zhang, Lia (2024) China Retrospective cohort study Sepsis patients ICU of a hospital SAE 52/130 (40.00%)
Jin, Jun (2024) United States Retrospective cohort study Patients with sepsis ≥ 18 years of age MIMIC-IV SAE 2781/4476 (62.1%)
Zhao, Lina (2021) United States Retrospective cohort study Patients with sepsis ≥ 18 years of age MIMIC III SAE 1,055/2,535 (41.6%)
Liu, Xiaoyua (2021) China Prospective cohort study Patients with sepsis ≥ 18 years of age ICU and emergency department at a hospital SAE 57/90 (63.3%)
Mei, Jiangjun (2024) China Prospective cohort study Patients with sepsis ≥ 18 years of age ICU of a hospital SAE 32/64 (47.8%)
Ge, Chenglong (2022) United States Retrospective cohort study Patients with sepsis ≥ 18 years of age MIMIC III SAE 6,284/12,460 (50.4%)

Table 2 presents details of the included prediction models. Most studies (n = 9) constructed multivariable logistic regression models, while two applied machine-learning (ML) algorithms. The most frequently used predictors included age (5 models), SOFA score (5 models), serum sodium (Na+) (4 models), and body temperature, SpO2, Acute Physiology and Chronic Health Evaluation II (APACHE II) score, and serum albumin (each in 3 models).

Table 2. Overview of the information of the included prediction models.

Author
(year)
Missing data handing Variable selection Model development method Calibration method Validation method Final predictors Model performance Model presentation
Wang, Ziwen (2023) – Forward LR Logistic regression model Hosmer– Lemeshow test Internal validation Age, Use Boosters, Albumin, SpO2, S100β A,0.810
(0.763–0.857)
B,0.813
(0.740–0.885)
Nomogram model
Zhou,
Hang
xiang (2023)
– Stepwise regression analysis Logistic regression model Hosmer– Lemeshow test Internal validation APACHEII, SOFA, Middle cerebral artery PI, Arterial blood lactate, ALT, rScO2,
Albumin
B,0.831
(0.773–0.889)
Nomogram model
Zhao,
Qing (2023)
Direct exclusion – Logistic regression model Calibration curve analysis Internal validation Age, SOFA, Na+, HR, T A, 0.802
B, 0.809
Nomogram model
Xiao, Lu (2022) Multiple interpolation – GBDT
XGBoost
Light-GBM
SVM
DT
RF
– Internal validation The top 5 factors ranked by importance are as follows: mechanical ventilation, duration of mechanical ventilation, serum phosphorus level, SOFA, vasopressor A,
0.883 (0.847–0.896)
0.902 (0.883–0.919)
0.879 (0.864–0.887)
0.832 (0.824–0.857)
0.849 (0.838–0.868)
0.886 (0.871–0.890)
B,
0.872 (0.859–0.885)
0.884 (0.865–0.898)
0.877 (0.869–0.888)
0.818 (0.808–0.839)
0.847 (0.839–0.855)
0.874 (0.868–0.881)
Script
Zhang,
Li (2024)
– – Logistic regression model – – APACHE II,
SOFA, CCI,
lung infection, Hb
A,0.975
(0.931–1)
–
Jin, Jun (2024) Multiple interpolation LASSO regression Multivariable Logistic regression model Hosmer– Lemeshow test
Calibration curve analysis
Internal validation Gender, Age, BMI, MAP, T, Platelet count, Na+, Midazolam use, SOFA A,0.751
(0.734–0.768)
B,0.766
(0.74–0.793)
Nomogram model
Zhao, Li
na (2021)
Mean value interpolation LASSO regression Multivariable Logistic regression model Calibration curve analysis Internal validation Age, Carbapenem, antibiotics, Quinolone antibiotics, qSOFA, midazolam, H antagonists, Steroids, Phenylephrine, hydrochloride,
Heparin sodium injection
A,0.743
(0.72–0.766)
B,0.762
(0.716–0.807)
Nomogram model
Liu, Xiaoyu (2021) Direct exclusion – Multivariable Logistic regression model – – APACHE II, CD86MFI, Albumin A,0.894
(0.817–0.97)
Formulas
Mei, Jiang jun (2024) – – Multivariable Logistic regression model Hosmer– Lemeshow test
Calibration curve analysis
Internal validation CCT, PI, S100β B,0.924
(0.833–0.975)
Nomogram model
Ge, Cheng long (2022) Single interpolation – Logistic regression model
SVM
DT
RF
GBM
MLP
XGBoost
LGBM
Calibration curve analysis Internal validation GCS, Glucose, Age, Mean arterial pressure, Mean heart rate, Hemoglobin, Length of stay in hospital, Length of stay in ICU, Platelet, WBC, PO2, Weight, Liver disease, PH, Resprate mean A,0.74
0.72
0.71
0.75
0.77
0.76
0.91
B,0.74
0.71
0.72
0.86
0.87
0.69
0.85
0.87
Script

Notes.

SVC
Support vector machine
DT
Decision Tree
RF
Random Forest
GBM
Gradients Boosting Machine
MLP
Multiple Layer Perception
GBDT
Gradient Boosting Decision Tree
SVM
Support Vector Machine
S100β
S100 Calcium-Binding Proteinβ
ALT
Alanine Aminotransferase
HR
Heart Rate
T
Temperature
LOS
ICU Stay Time
CCI
Charlson Comorbidity Index
Hb
Hemoglobin.
BMI
Body Mass Index
MAP
Mean Arterial Pressure
qSOFA
Quick Sepsis Related Organ Failure Assessment
CD86MFI
CD86 Mean Fluorescence Intensity
WBC
White Blood Cell

A, development cohort. B, validation cohort.

Model validation

Among the included studies, eight conducted internal validation using different approaches. Three studies employed the bootstrap method, five used random split validation, and two utilised cross-validation. None of the models underwent external validation. Reported AUC values ranged from 0.743 to 0.975, with 9 models exceeding 0.75, indicating strong discriminatory performance. Seven studies assessed model calibration, primarily using calibration curves. Among these, four also reported Hosmer–Lemeshow test p-values > 0.05 (range: 0.126–0.944), indicating good agreement between predicted and observed risks. Four studies further evaluated the clinical utility of the models via Decision-Curve Analysis (DCA), demonstrating that the models provided a positive net benefit across clinically relevant probability thresholds.

Results of quality assessment

Table 3 summarises the quality level, risk of bias, and applicability of the included studies. All studies were assessed as having a high risk of bias, mainly due to issues identified in the participant and analysis domains, as detailed below.

Table 3. PROBAST results of the included studies.

Author (year) Study ROB Applicability Overall
Participants Predictors Outcome Analysis Participants Predictors Outcome ROB Applicability
Wang, Ziwen (2023) B – + – – + + + – +
Zhou, Hang
xiang (2023)
B – + + – + + + – +
Zhao, Qing (2023) B – + – – + + + – +
Xiao, Lu (2022) B – + – – + + + – +
Zhang, Li (2024) B – + + – + + + – +
Jin, Jun (2024) B – + – + + + + – +
Zhao, Lina (2021) B – + – + + + + – +
Liu, Xiaoyu (2021) B + + + – + + + – +
Mei, Jiangjun (2024) B + – – – + + + – +
Ge, Chenglong (2022) B – + – – + + + – +

Notes.

PROBAST
Prediction model Risk Of Bias Assessment Tool
B
moderate
ROB
risk of bias

+indicates low ROB/low concern regarding applicability.

-indicates high ROB/high concern regarding application.

In the participant domain, the majority of studies were at high risk of bias, mainly attributable to their retrospective design (Ge et al., 2022; Jin et al., 2024; Lu et al., 2022; Wang, Zhao & Chao, 2023; Zhang et al., 2024; Zhao et al., 2021; Zhao et al., 2023; Zhou et al., 2023). In the predictor domain, one study was deemed high risk due to potential information bias in its predictors (Mei et al., 2024). In the outcome domain, seven studies were at high risk of not adequately separating predictor variables from the outcome definition (Ge et al., 2022; Jin et al., 2024; Lu et al., 2022; Mei et al., 2024; Wang, Zhao & Chao, 2023; Zhao et al., 2021; Zhao et al., 2023).

The analysis domain presented the most widespread concerns, with all studies exhibiting a high risk of bias. Key methodological shortcomings included an insufficient sample size relative to the number of predictor variables, insufficient reporting of missing-data handling, omission of calibration performance measures, lack of model validation, and reliance on univariate screening.

Regarding applicability, all models were rated low across domains, indicating strong relevance to the target clinical scenario. For instance, the machine learning model developed by Lu et al. demonstrated not only high performance but also enhanced interpretability through SHAP analysis and clinical review, underscoring its potential utility.

Meta-analysis

Four studies were excluded from the meta-analysis due to insufficient reporting of model validation details. Two studies involved multiple models, and all methods were based on the same samples; therefore, only the best-performing XGBoost and Light Gradient Boosting Machine (LGBM) models were included. Consequently, six studies that provided adequate validation data were synthesized (Jin et al., 2024; Lu et al., 2022; Mei et al., 2024; Wang, Zhao & Chao, 2023; Zhao et al., 2021; Zhou et al., 2023). Using a random-effects model, the combined AUC was (95% CI [0.77–0.89]) (Fig. 2). The I2 value was 93.3% (p < 0.001), indicating a high degree of heterogeneity between the studies. Sensitivity analysis, performed by sequentially excluding each study, showed minimal change in the overall results, suggesting the meta-analysis was robust. The Egger’s test yielded a value of 0.2699 (p > 0.05), indicating a low likelihood of publication bias (Fig. 3).

Figure 2. AUC’s forest diagram.

Figure 2

Figure 3. Funnel diagram.

Figure 3

Discussion

The number of SAE prediction models has been steadily increasing; however, their clinical practice, quality, and applicability are still unknown. This study is the first to systematically evaluate the methodological quality of SAE prediction models and provide evidence-based recommendations for model optimization.

During the process of constructing the SAE model, several noteworthy strengths are worth learning from. For example, Jin et al. (2024) included a large sample size their retrospective design increased the potential for bias. However, their study excelled in analysis by using multiple imputation for missing data, an approach often neglected in similar studies. Mei et al. (2024) employed Least Absolute Shrinkage and Selection Operator (LASSO) regression for variable selection, which offers notable advantages in terms of predictive accuracy and model interpretability (Emmert-Streib & Dehmer, 2019). In contrast, two studies performed direct deletion of missing data; however, this approach has been shown to perform poorly in terms of calibration and predictive accuracy. Future studies could utilize optimal performance methods, such as multiple imputation for missing values (Deforth, Heinze & Held, 2024). Notably, Lu et al. (2022) and Ge et al. (2022) applied machine learning (ML) methods during model development. It has been shown that machine learning methods tend to achieve higher accuracy than traditional logistic regression (Churpek et al., 2016), owing to their superior ability to capture nonlinear relationships, handle high-dimensional data, and automate variable selection. However, one of the drawbacks of machine learning models is their lack of interpretability, and many ML scientists agree that “black boxes” are one of the main barriers to the adoption of ML in medicine (Vellido, 2020). Therefore, the development of interpretable ML models is an urgent need nowadays. Lu et al. (2022) employed the SHAP method to interpret the outputs of the constructed ML models and invited clinicians to rate them, thereby further enhancing the models’ interpretability. This combined strategy—SHAP-based interpretation coupled with expert evaluation—represents a promising framework for improving transparency in ML-driven prediction models.

The existing predictive models reported in this review also have important clinical implications. The high-frequency predictors shown are informative for future nursing practice and clinical diagnostic studies. The SOFA score is a commonly used method for assessing the degree of organ failure in patients with sepsis; although SOFA does not directly diagnose SAE, the severity of sepsis as assessed by the SOFA has been associated with the presence or absence of SAE, higher SOFA scores have been associated with SAE occurrence (Leventogiannis et al., 2022). The study by Zhao et al. (2021) included both SOFA scores and qSOFA scores; ultimately, qSOFA was used to construct the model, whereas SOFA did not emerge as a significant predictor—an observation that warrants further investigation to inform model development. The SOFA score, as a robust and readily available clinical metric, shows consistent predictive value for SAE and should be considered a cornerstone variable in future risk stratification tools. Age is a well-established susceptibility factor for SAE, and the incidence of SAE gradually increases after the aged ≥50 years, which is associated with a decline in brain reserve, a physiological decline in the function of most organs, and concomitant complications (Schütze et al., 2023). Therefore, healthcare professionals should be alert to septic patients aged 50 years and older.

Sodium (Na+) is critical for plasma osmolality regulation. Hypernatremia increases plasma osmolality and can drive intracellular water shifts in neurons, precipitating neurological dysfunction; Yang et al. (2020) reported a significant association between hypernatremia and SAE (Yang et al., 2020). APACHE II score and body temperature were used as predictors in three studies. Fever commonly reflects systemic infection and may disrupt cerebral metabolism; prolonged hyperthermia can exacerbate neuronal injury and compromise the blood–brain barrier (Sun et al., 2013). Higher APACHE II scores reflect greater overall disease severity and have been linked to increased mortality and SAE risk (Liu, 2021; Sonneville et al., 2017). Future model development should integrate these key predictors and further examine their mechanistic roles and interactions to improve predictive performance and clinical applicability.

The clinical significance of these individual predictors ultimately manifests in how they are selected and integrated into multivariable models. An examination of the six predictive models included in this meta-analysis (Jin et al., 2024; Lu et al., 2022; Mei et al., 2024; Wang, Zhao & Chao, 2023; Zhao et al., 2021; Zhou et al., 2023) revealed substantial discrepancies in their construction strategies. Specific models integrated organ dysfunction scores (SOFA/APACHE II), routine laboratory parameters, and demographic data to synthesise commonly available clinical information. By contrast, other models focused on pathophysiological mechanisms, emphasising cerebral hemodynamics or immune biomarkers. Notably, Zhao et al. (2021) incorporated treatment-related variables, highlighting the potential impact of clinical interventions on the risk of sepsis-associated encephalopathy (SAE). The diversity in predictor selection not only corroborates the multifactorial nature of SAE but also partially accounts for the significant statistical heterogeneity observed in this meta-analysis. Simultaneously, it underscores the current lack of a consensus regarding the core variables for SAE prediction.

This review also identified several significant issues with existing SAE models. We included 10 studies and conducted a meta-analysis of six predictive models (Jin et al., 2024; Lu et al., 2022; Mei et al., 2024; Wang, Zhao & Chao, 2023; Zhao et al., 2021; Zhou et al., 2023). The ten included studies reported AUC values ranging from 0.786 to 0.988. However, all studies were deemed high-risk according to the PROBAST checklist, limiting the practical applicability of predictive models. The pooled AUC of the six validated models was 0.83 (95% CI [0.77–0.89]), demonstrating moderate predictive performance but substantial heterogeneity. Such heterogeneity likely arises from differences in study design (predominantly retrospective), patient characteristics, and data quality. Addressing such heterogeneity and improving model generalizability requires systematic multicenter external validation.

All included studies were authored by Chinese researchers, potentially reflecting China’s growing concern regarding SAE (Mazeraud et al., 2020). This heightened concern arises from the high prevalence and clinical burden of sepsis in the Chinese population, where SAE frequently occurs as one of its most severe complications. The reported incidence of sepsis in China remains persistently high, mainly due to its large population, aging demographics, and high rates of comorbidities (Liao et al., 2016; Weng et al., 2023). Elderly patients, in particular, are more vulnerable to SAE due to diminished physiological reserves and pre-existing conditions (Gu et al., 2025). China’s issues with antibiotic resistance and uneven distribution of medical resources further contribute to the high incidence of sepsis (Luo et al., 2024; Weng et al., 2023).

Consequently, Chinese researchers have been actively developing early diagnostic models to mitigate the risk of misdiagnosis and improve clinical outcomes. Furthermore, the availability of international open-access databases, such as Medical Information Mart for Intensive Care (MIMIC), has facilitated the development of SAE prediction models. However, most of these models remain reliant on retrospective, single-centre datasets and lack external validation across diverse populations and healthcare systems, which limits their generalizability and clinical applicability.

In summary, the two major issues identified in this review—high heterogeneity and high risk of bias—suggest that future studies must undergo multicenter external validation and adopt rigorous research designs. External validation is essential to determine whether the observed heterogeneity arises from model overfitting or from clinical differences in patient populations and healthcare settings. Validating models across centres in different regions and healthcare environments allows direct assessment of their generalisation capabilities in patient populations and settings distinct from the development cohort. Moreover, validation across diverse clinical settings can assess the robustness of predictive performance and thereby determine broad clinical applicability (Debray et al., 2015). Importantly, external validation also identifies meaningful variations across environments, driving the development of more universal models or guiding adjustments tailored to specific contexts.

High bias risk primarily stems from suboptimal study designs, underscoring the need for rigorous research methodologies in future work. Currently, most studies are retrospective, which may introduce bias in participant selection and data quality. Future research should therefore prioritize prospective study designs, which can mitigate bias and ensure more comprehensive data collection. In addition, a priori sample size calculations should be performed to ensure reliability (Talari & Goyal, 2020). To maintain data integrity, advanced techniques such as multiple imputation should be employed to handle missing values appropriately. Furthermore, reliance on univariate analysis for variable selection may omit important predictors or overemphasize less relevant ones (Chang & Chen, 2025). Techniques such as the least absolute shrinkage and selection operator regression should be employed for more reliable variable selection, particularly in high-dimensional datasets.

Limitations

This study has several limitations. First, it only included literature in Chinese and English, which may have introduced language bias. Due to significant heterogeneity among the studies, quantitative analysis was not performed. Furthermore, most of the included studies were retrospective, single-centre investigations lacking external validation, which may limit the generalizability and clinical applicability of the models. Adjustments may be necessary when applying these models in different regions.

Conclusions

This systematic review evaluated 10 studies. For the six models subjected to validation, the pooled AUC was 0.83 (95% CI [0.77–0.89]), reflecting moderate discriminatory capability. However, according to PROBAST and GRADE assessments, all studies were classified as having a high risk of bias and low methodological quality. To translate this predictive potential into clinical utility, researchers should strengthen study design, perform multicenter external validation, and adhere to the PROBAST framework to enhance the credibility and clinical utility of future models.

Supplemental Information

Supplemental Information 1. PRISMA checklist.
peerj-14-20770-s001.docx (266.3KB, docx)
DOI: 10.7717/peerj.20770/supp-1
Supplemental Information 2. Pubmed history.
peerj-14-20770-s002.xlsx (10.6KB, xlsx)
DOI: 10.7717/peerj.20770/supp-2

Funding Statement

This work was supported by the Deyang City Science and Technology Program (2024SZY009). The funders had no role in study design, data collection and analysis, decision to publish, or preparation of the manuscript.

Additional Information and Declarations

Competing Interests

The authors declare there are no competing interests.

Author Contributions

Ting ting He conceived and designed the experiments, performed the experiments, analyzed the data, prepared figures and/or tables, authored or reviewed drafts of the article, and approved the final draft.

Tuo quan Jiao conceived and designed the experiments, analyzed the data, authored or reviewed drafts of the article, and approved the final draft.

Xue mei An conceived and designed the experiments, performed the experiments, analyzed the data, prepared figures and/or tables, and approved the final draft.

Data Availability

The following information was supplied regarding data availability:

This is a systematic review/meta-analysis.

References

  • Aljefri et al. (2023).Aljefri AA, Almutairi LH, Alraddadi MH, Alahmadi SA, Sufyani AQ, Almarzooq JN, Nashri AHA, Al Wosaibi FA, Boukhamssein NA, Almahasnah MS, Alyami SD. Definition, epidemiology and characterization of sepsis. International Journal of Community Medicine and Public Health. 2023;11:376–380. doi: 10.18203/2394-6040.ijcmph20233852. [DOI] [Google Scholar]
  • Chang & Chen (2025).Chang T-E, Chen A. Variable selection using relative importance rankings. 2025. Undefined. [DOI]
  • Churpek et al. (2016).Churpek MM, Yuen TC, Winslow C, Meltzer DO, Kattan MW, Edelson DP. Multicenter comparison of machine learning methods and conventional regression for predicting clinical deterioration on the wards. Critical Care Medicine. 2016;44:368–374. doi: 10.1097/ccm.0000000000001571. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • Debray et al. (2015).Debray TPA, Vergouwe Y, Koffijberg H, Nieboer D, Steyerberg EW, Moons KGM. A new framework to enhance the interpretation of external validation studies of clinical prediction models. Journal of Clinical Epidemiology. 2015;68:279–289. doi: 10.1016/j.jclinepi.2014.06.018. [DOI] [PubMed] [Google Scholar]
  • Deforth, Heinze & Held (2024).Deforth M, Heinze G, Held U. The performance of prognostic models depended on the choice of missing value imputation algorithm: a simulation study. Journal of Clinical Epidemiology. 2024;176:111539. doi: 10.1016/j.jclinepi.2024.111539. [DOI] [PubMed] [Google Scholar]
  • Dumbuya et al. (2023).Dumbuya JS, Li S, Liang L, Zeng Q. Paediatric sepsis-associated encephalopathy (SAE): A comprehensive review. Missouri Medicine. 2023;29:27. doi: 10.1186/s10020-023-00621-w. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • Egger et al. (1997).Egger M, Smith GDavey, Schneider M, Minder C. Bias in meta-analysis detected by a simple, graphical test. Bmj. 1997;315:629–634. doi: 10.1136/bmj.315.7109.629. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • Emmert-Streib & Dehmer (2019).Emmert-Streib F, Dehmer M. High-dimensional LASSO-based computational regression models: regularization, shrinkage, and selection. Machine Learning and Knowledge Extraction. 2019;1:359–383. doi: 10.3390/make1010021. [DOI] [Google Scholar]
  • Ge et al. (2022).Ge C, Deng F, Chen W, Ye Z, Zhang L, Ai Y, Zou Y, Peng Q. Machine learning for early prediction of sepsis-associated acute brain injury. Frontiers in Medicine. 2022;9:962027. doi: 10.3389/fmed.2022.962027. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • Gu et al. (2025).Gu W, Zhong J, Lyu C, Zhang G, Xie M, Ma Y, Guo W. An approach for the emergency diagnosis and treatment of sepsis-associated encephalopathy in elderly individuals: a literature review. World Journal of Emergency Medicine. 2025;16:415. doi: 10.5847/wjem.j.1920-8642.2025.0101. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • Guyatt et al. (2008).Guyatt GH, Oxman AD, Vist GE, Kunz R, Falck-Ytter Y, Alonso-Coello P, Schünemann HJ. GRADE: an emerging consensus on rating quality of evidence and strength of recommendations. Bmj. 2008;336:924–926. doi: 10.1136/bmj.39489.470347.AD. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • Hotchkiss & Karl (2003).Hotchkiss RS, Karl IE. The pathophysiology and treatment of sepsis. New England Journal of Medicine. 2003;348:138–150. doi: 10.1056/NEJMra021333. [DOI] [PubMed] [Google Scholar]
  • Jin et al. (2024).Jin J, Yu L, Zhou Q, Zeng M. Improved prediction of sepsis-associated encephalopathy in intensive care unit sepsis patients with an innovative nomogram tool. Frontiers in Neurology. 2024;15:1344004. doi: 10.3389/fneur.2024.1344004. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • Lehrer et al. (2023).Lehrer EJ, Wang M, Sun Y, Zaorsky NG. An introduction to meta-analysis. International Journal of Radiation Oncology, Biology, Physics. 2023;115:564–571. doi: 10.1016/j.ijrobp.2022.07.1831. [DOI] [PubMed] [Google Scholar]
  • Lei & Wu (2025).Lei Z, Wu X. Research progress of sepsis-related encephalopathy. Medical Journal of Wuhan University. 2025;46:1358–1363. doi: 10.14188/j.1671-8852.2024.0549. [DOI] [Google Scholar]
  • Leventogiannis et al. (2022).Leventogiannis K, Kyriazopoulou E, Antonakos N, Kotsaki A, Tsangaris I, Markopoulou D, Grondman I, Rovina N, Theodorou V, Antoniadou E, Koutsodimitropoulos I, Dalekos G, Vlachogianni G, Akinosoglou K, Koulouras V, Komnos A, Kontopoulou T, Prekates A, Koutsoukou A, Van der Meer JWM, Dimopoulos G, Kyprianou M, Netea MG, Giamarellos-Bourboulis EJ. Toward personalized immunotherapy in sepsis: the PROVIDE randomized clinical trial. Cell Reports Medicine. 2022;3:100817. doi: 10.1016/j.xcrm.2022.100817. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • Li et al. (2025).Li J, Jia Q, Yang L, Wu Y, Peng Y, Du L, Fang Z, Zhang X. Sepsis-associated encephalopathy: Mechanisms, diagnosis, and treatments update. International Journal of Biological Sciences. 2025;21:3214–3228. doi: 10.7150/ijbs.102234. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • Liao et al. (2016).Liao X, Du B, Lu M, Wu M, Kang Y. Current epidemiology of sepsis in mainland China. Annals of Translational Medicine. 2016;4:324–324. doi: 10.21037/atm.2016.08.51. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • Liu (2021).Liu X. Master’s thesis. 2021. The expression of CD86 in CD3+CD56+NKT cell is associated with sepsis-associated encephalopathy in sepsis patients. [Google Scholar]
  • Lu et al. (2022).Lu X, Kang H, Zhou D, Li Q. Prediction and risk assessment of sepsis-associated encephalopathy in ICU based on interpretable machine learning. Scientific Reports. 2022;12:22621. doi: 10.1038/s41598-022-27134-6. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • Luo et al. (2024).Luo Q, Lu P, Chen Y, Shen P, Zheng B, Ji J, Ying C, Liu Z, Xiao Y. ESKAPE in China: Epidemiology and characteristics of antibiotic resistance. Emerging Microbes & Infections. 2024;13:2317915. doi: 10.1080/22221751.2024.2317915. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • Mazeraud et al. (2020).Mazeraud A, Righy C, Bouchereau E, Benghanem S, Bozza FA, Sharshar T. Septic-associated encephalopathy: a comprehensive review. Neurotherapeutics. 2020;17:392–403. doi: 10.1007/s13311-020-00862-1. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • Mei et al. (2024).Mei J, Zhang X, Sun X, Hu L, Song Y. Optimizing the prediction of sepsis-associated encephalopathy with cerebral circulation time utilizing a nomogram: a pilot study in the intensive care unit. Frontiers in Neurology. 2024;14:1303075. doi: 10.3389/fneur.2023.1303075. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • Moons et al. (2014).Moons KG, De Groot JA, Bouwmeester W, Vergouwe Y, Mallett S, Altman DG, Reitsma JB, Collins GS. Critical appraisal and data extraction for systematic reviews of prediction modelling studies: The CHARMS checklist. PLOS Medicine. 2014;11:e1001744. doi: 10.1371/journal.pmed.1001744. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • Moons et al. (2019).Moons KGM, Wolff RF, Riley RD, Whiting PF, Westwood M, Collins GS, Reitsma JB, Kleijnen J, Mallett S. PROBAST: a tool to assess risk of bias and applicability of prediction model studies: Explanation and elaboration. Annals of Internal Medicine. 2019;170:W1–w33. doi: 10.7326/m18-1377. [DOI] [PubMed] [Google Scholar]
  • Schütze et al. (2023).Schütze S, Drevets DA, Tauber SC, Nau R. Septic encephalopathy in the elderly - Biomarkers of potential clinical utility. Frontiers in Cellular Neuroscience. 2023;17:1238149. doi: 10.3389/fncel.2023.1238149. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • Shirodkar et al. (2025).Shirodkar R, Bourgeois IJ, Kim M, Kimchi EY, Liotta EM, Maas MB. Covert critical illness encephalopathy: Impairments that escape detection by guideline recommended, protocolized assessments. Critical Care Medicine. 2025;53:e613–e618. doi: 10.1097/ccm.0000000000006558. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • Singer et al. (2016).Singer M, Deutschman CS, Seymour CW, Shankar-Hari M, Annane D, Bauer M, Bellomo R, Bernard GR, Chiche JD, Coopersmith CM, Hotchkiss RS, Levy MM, Marshall JC, Martin GS, Opal SM, Rubenfeld GD, Poll Tvander, Vincent JL, Angus DC. The third international consensus definitions for sepsis and septic shock (sepsis-3) Jama. 2016;315:801–810. doi: 10.1001/jama.2016.0287. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • Sonneville et al. (2017).Sonneville R, De Montmollin E, Poujade J, Garrouste-Orgeas M, Souweine B, Darmon M, Mariotte E, Argaud L, Barbier F, Goldgran-Toledano D, Marcotte G, Dumenil AS, Jamali S, Lacave G, Ruckly S, Mourvillier B, Timsit JF. Potentially modifiable factors contributing to sepsis-associated encephalopathy. Intensive Care Medicine. 2017;43:1075–1084. doi: 10.1007/s00134-017-4807-z. [DOI] [PubMed] [Google Scholar]
  • Sun et al. (2013).Sun G, Qian S, Jiang Q, Liu K, Li B, Li M, Zhao L, Zhou Z, von Deneen KM, Liu Y. Hyperthermia-induced disruption of functional connectivity in the human brain network. PLOS ONE. 2013;8:e61157. doi: 10.1371/journal.pone.0061157. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • Talari & Goyal (2020).Talari K, Goyal M. Retrospective Studies –Utility and Caveats. Journal of the Royal College of Physicians of Edinburgh. 2020;50:398–402. doi: 10.4997/jrcpe.2020.409. [DOI] [PubMed] [Google Scholar]
  • Vellido (2020).Vellido A. The importance of interpretability and visualization in machine learning for applications in medicine and health care. Neural Computing and Applications. 2020;32:18069–18083. doi: 10.1007/s00521-019-04051-w. [DOI] [Google Scholar]
  • Wang, Zhao & Chao (2023).Wang Z, Zhao W, Chao Y. Establishment and validation of a predictive model for sepsis patient-associated encephalopathy. China Emergency Medicine. 2023;43:434–439. doi: 10.3969/j.issn.1002-1949.2023.06.002. [DOI] [Google Scholar]
  • Weng et al. (2023).Weng L, Xu Y, Yin P, Wang Y, Chen Y, Liu W, Li S, Peng JM, Dong R, Hu XY, Jiang W, Wang CY, Gao P, Zhou MG, Du B. National incidence and mortality of hospitalized sepsis in China. Critical Care. 2023;27:84. doi: 10.1186/s13054-023-04385-x. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • Yang et al. (2020).Yang J, Li Y, Liu Q, Li L, Feng A, Wang T, Zheng S, Xu A, Lyu J. Brief introduction of medical database and data mining technology in big data era. Journal of Evidence-Based Medicine. 2020;13:57–69. doi: 10.1111/jebm.12373. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • Zhang et al. (2024).Zhang L, Yu X, Ma L, Wang Y, Li X, Yang Y. Construction and analysis of early warning and prediction model for risk factors of sepsis-associated encephalopathy. Zhonghua Wei Zhong Bing Ji Jiu Yi Xue. 2024;36:124–130. doi: 10.3760/cma.j.cn121430-20231008-00847. [DOI] [PubMed] [Google Scholar]
  • Zhao et al. (2021).Zhao L, Wang Y, Ge Z, Zhu H, Li Y. Mechanical learning for prediction of sepsis-associated encephalopathy. Frontiers in Computational Neuroscience. 2021;15:739265. doi: 10.3389/fncom.2021.739265. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • Zhao et al. (2023).Zhao Q, Xiao J, Liu X, Liu H. The nomogram to predict the occurrence of sepsis-associated encephalopathy in elderly patients in the intensive care units: A retrospective cohort study. Frontiers in Neurology. 2023;14:1084868. doi: 10.3389/fneur.2023.1084868. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • Zhou et al. (2023).Zhou H, Yuan J, Zhang Q, Tao J, Liu L. Factors influencing the occurrence of sepsis-associated encephalopathy and the construction of its column-line diagram risk model. Journal of Difficult Diseases. 2023;22:1245–1250. doi: 10.3969/j.issn.1671-6450.2023.12.003. [DOI] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Supplemental Information 1. PRISMA checklist.
peerj-14-20770-s001.docx (266.3KB, docx)
DOI: 10.7717/peerj.20770/supp-1
Supplemental Information 2. Pubmed history.
peerj-14-20770-s002.xlsx (10.6KB, xlsx)
DOI: 10.7717/peerj.20770/supp-2

Data Availability Statement

The following information was supplied regarding data availability:

This is a systematic review/meta-analysis.


Articles from PeerJ are provided here courtesy of PeerJ, Inc

RESOURCES