Skip to main content
PLOS One logoLink to PLOS One
. 2026 Jun 2;21(6):e0348676. doi: 10.1371/journal.pone.0348676

The prognostic value of the early neutrophil-to-lymphocyte ratio for 28-day mortality in sepsis patients: A machine learning-based investigation of the MIMIC database

Jiyang Liao 1,☯, Qianwen Xiang 2,☯, Xingwang Chen 1,☯, Long Wu 1, Houwang Chen 1, Zhijun Yao 1, Huachu Wu 1,*, Jianbo Lai 1,*
Editor: Chiara Lazzeri3
PMCID: PMC13229304  PMID: 42228737

Abstract

Background

The neutrophil-to-lymphocyte ratio (NLR) has shown inconsistent prognostic value in individuals with sepsis. This study aimed to clarify its ability to predict 28-day mortality via a machine learning-based analysis of a large ICU database.

Methods

This retrospective analysis employed data from the MIMIC-IV database (v3.1). The Boruta algorithm combined with XGBoost was used for two-stage feature selection. Patients were stratified by NLR quartiles into three groups (low: <4.34, intermediate: 4.34–14.70, and high: >14.70). This study defined 28-day mortality as the primary outcome. Associations between the NLR and mortality were evaluated by using multivariable logistic regression (progressively adjusted for demographic, clinical, and machine learning-derived features), along with restricted cubic splines. Sensitivity analyses included quantifying NLR feature importance via machine learning and performing subgroup analyses across clinical strata.

Results

This cohort study included 4,376 patients with a 28-day mortality rate of 18.4%. Compared with the SOFA and SAPS II scores, the prediction performance of XGBoost was superior (ROC-AUC 0.875; 95% CI 0.854–0.896; PR-AUC 0.603). Although the NLR ranked 14th in SHAP-based feature importance, multivariable analysis confirmed its independent association with elevated mortality risk: 28-day (OR 1.16; 95% CI 1.06–1.27; p < 0.001), in-hospital (OR 1.13; 95% CI 1.03–1.24; p < 0.001), and ICU (OR 1.14; 95% CI 1.03–1.25; p = 0.008). Stratified analyses indicated consistent mortality associations for the NLR, with enhanced predictive value being observed in patients aged >45 years and those with SOFA scores ≤4 or SAPS II scores >29. Specifically, patients ≥65 years of age demonstrated a 17% increase in 28-day mortality risk (p = 0.019), and patients with a SOFA score ≤4 exhibited a greater than 20% elevated risk across all of the endpoints (p < 0.001), whereas no significant association was observed in the SOFA ≥9 subgroup (p = 0.369).

Conclusions

The NLR effectively identifies inflammation-driven mortality risk for early sepsis patients but fails to predict outcomes for patients with terminal organ failure. This biphasic predictive pattern highlights the unique value of the NLR in moderate sepsis risk stratification but cautions against its use in cases of advanced disease. Its value lies in dynamic monitoring rather than static risk assessment.

Introduction

Sepsis is a critical medical state involving organ failure resulting from a maladaptive systemic reaction to infection that impacts millions of patients worldwide annually [1]. Despite recent advancements in sepsis recognition and treatment strategies, mortality rates remain alarmingly high and range from 16.7% to 33.3% [2,3], imposing a substantial burden on global healthcare economics. Nevertheless, early identification and intervention have led to significant improvements in clinical outcomes.

Current assessment methods predominantly rely on manual documentation or electronic medical records, which often incorporate nursing process-dependent parameters such as vital sign monitoring [4–7]. Variations in measurement and recording protocols across healthcare institutions may lead to substantial discrepancies in evaluation results. Although the levels of routinely available inflammatory markers such as procalcitonin and C-reactive protein are easily determined in clinical practice, they exhibit limited prognostic value for sepsis because of their suboptimal specificity and sensitivity [1,8]. Although conventional scoring systems such as sequential organ failure assessment (SOFA) and acute physiology and chronic health assessment scoring system II (APACHE II) demonstrate high prognostic value, their complex assessment procedures and poor operational feasibility limit their widespread clinical application [9–11]. Consequently, the pursuit of more objective and convenient prognostic indicators remains important.

Neutrophils (NEUT) and lymphocytes (LYMPH), as the principal cellular components of innate and adaptive immunity, respectively, play indispensable roles as first-line immune defenses during systemic inflammatory responses [12,13]. Bacterial or fungal infections typically induce isolated NEUT elevation with concomitant lymphopenia, whereas viral infections or malignancies may predominantly increase LYMPH counts [14]. Consequently, the neutrophil-to-lymphocyte ratio (NLR) generally increases during inflammatory states. Although clinically favorable because of its accessibility and low cost, some alterations in the NLR are not inflammation-specific; for example, acute trauma, stroke, myocardial infarction, and postoperative complications can similarly modify this ratio, which compromises its diagnostic specificity and contributes to persistent clinical skepticism regarding its utility [15]. The COVID-19 pandemic has reinvigorated research interest in the NLR [16,17], particularly as a result of a recent meta-analysis by Wu et al. [18], which reported reliable prognostic performance of the NLR in sepsis despite considerable between-study heterogeneity (I² = 87.2%, 95% CI 79.5–92; p < 0.0001). Moreover, a retrospective study conducted by Schupp et al. [19] failed to demonstrate a discriminative capacity for 30-day mortality. These contradictory prognostic interpretations in sepsis underscore the need for machine learning (ML)-based reevaluations of the ability of the NLR to predict mortality.

Machine learning algorithms have emerged in recent years as powerful tools in clinical prediction models. The 2021 sepsis guidelines highlight the superior discriminative capacity of ML algorithms for predicting in-hospital mortality in sepsis patients [1]. These algorithms offer dual advantages—the efficient management of missing data and the ability to integrate weak predictive features into robust models through data-driven learning. Among the various ML techniques, extreme gradient boosting (XGBoost) has demonstrated exceptional data mining and predictive modeling capabilities. Its dominance is evidenced by its adoption in 17 of 29 winning solutions in the 2015 Kaggle competitions and its exclusive use by all top 10 teams that participated in the 2015 KDD Cup. Several studies have validated the clinical utility of XGBoost: Dong et al. developed a ML model that predicts pediatric acute kidney injury 48 hours in advance by analyzing premorbid physiological measurements and incorporating current guidelines; Zhang et al. established an XGBoost model that effectively distinguishes between fluid-responsive and nonresponsive sepsis patients; and Yue et al. created a ML model for the early detection of sepsis-associated acute kidney injury and identified XGBoost as the optimal predictive algorithm. These findings collectively support the theoretical potential of ML algorithms to enhance predictive model development and validation in critical care. However, no studies have yet employed XGBoost modeling with large datasets to systematically evaluate the ability of the NLR to predict 28-day mortality, which could provide valuable insights into its clinical relevance under resource-constrained conditions.

Methods

Database

The patient data utilized for this study were derived from the Medical Information Mart for Intensive Care IV database version 3.1 (MIMIC-IV v3.1) [20,21]. This dataset, which includes deidentified health data from more than 65,000 Intensive Care Unit (ICU) hospitalizations and over 200,000 emergency department visits at Beth Israel Deaconess Medical Center (Boston, MA) from 2008–2022, offers broad clinical coverage. The dataset includes demographic characteristics, physiological parameters validated by ICU nursing staff, laboratory test results, therapeutic intervention orders, standardized data dictionaries, and diagnostic codes (International Classification of Diseases, 9th and 10th Revisions; ICD-9 and ICD-10). As this study used a deidentified public database with all protected health information removed, this study was exempted from the requirement to obtain individual patient informed consent from the Institutional Review Board (IRB) of the Massachusetts Institute of Technology (MIT). Author JY Liao (certification ID: 63692240) completed the required training and was authorized to access the database for research purposes. This study was performed in accordance with the Declaration of Helsinki. Due to the deidentified nature and public availability of the MIMIC database (which contains no personally sensitive information), the study was ethically exempt from review.

Study population

Inclusion criteria for this study were as follows: (1) aged between 18 and 85 years; (2) an ICU stay duration that exceeded 24 hours; and (3) suspected or confirmed infection with a SOFA score ≥2. For patients with multiple ICU admissions, only data obtained from the first ICU stay were included in the analysis. Patients were excluded if they met any of the following conditions: (1) had missing laboratory data, including NEUT, LYMPH and platelet (PLT) data; or (2) had a documented diagnosis of malignant tumors or autoimmune diseases. The study design is schematically illustrated in Fig 1.

Fig 1. Overview of the study design and workflow.

Fig 1

BIDMC, Beth Israel Deaconess Medical Center; MIMIC-IV, Medical Information Mart for Intensive Care IV; XGBoost, extreme gradient boosting; NLR, neutrophil-to-lymphocyte ratio; RCS, restricted cubic splines; ROC, receiver operating characteristic; PR, precision-recall; SHAP, Shapley additive explanations.

Data extraction

Using structured query language (SQL) within the Navicat Premium environment (version 16.1.12), we extracted the requisite patient data from the MIMIC-IV database. The analysis focused exclusively on each patient’s first ICU admission record. Due to the extensive temporal data entries in the MIMIC-IV database, strict time windows were applied during data extraction to minimize potential confounding from different treatment phases. Demographic characteristics and clinical variables, such as sex, age, weight, SOFA score, simplified acute physiology score II (SAPS II) score and comorbidities, were extracted from the earliest available clinical records during the first 24 hours following ICU admission. Vital signs were collected within the first 2 hours after ICU admission. The crystalloid fluid intake volume was obtained within the first 3 hours and was used to calculate the crystalloid-to-weight ratio. The first available laboratory test results obtained within the initial 6 hours of ICU admission were extracted for analysis. Cases with NEUT, LYMPH, or PLT counts of zero were excluded to ensure valid calculation of the systemic immune-inflammation index (SII), which was derived using the following formula: absolute NEUT count × PLT count/absolute LYMPH count [22]. Fluid intake and output volumes were recorded during the first 24 hours of ICU admission, after which the fluid intake-to-weight ratio was calculated. The vasoactive inotropic score (VIS) and antibiotic initiation time were concurrently determined. Upon ICU discharge, the types of mechanical ventilation (MV) and continuous renal replacement treatment (CRRT) were recorded. Patients who received noninvasive ventilation, invasive ventilation, or tracheostomy were classified as mechanically ventilated. The primary outcome was defined as all-cause mortality within 28 days following ICU admission. Secondary outcomes included in-hospital mortality, ICU mortality, length of hospital stay, and length of ICU stay. Follow-up was conducted from ICU admission until 28 days thereafter.

To address missing data, a 40% missingness threshold was set. Consequently, any parameter with missing values exceeding this percentage was excluded from analysis under the assumption that the data were missing at random (MAR). [23]. For variables where fewer than 40% of the values were missing, we performed imputation using the “missForest” package (version 1.5) in R Studio [24]. To assess the robustness of the imputation, we compared the distribution of key variables before and after imputation by using the missForest algorithm. Post-imputation values remained within clinically plausible ranges, and no substantial shifts in central tendency or dispersion were observed (S1 and S2 Tables).

Feature selection

Before the association between the NLR and 28-day mortality was assessed, we implemented a rigorous two-stage ML-driven feature selection protocol. In the first phase, we performed 10 iterative runs of the Boruta algorithm to evaluate feature selection stability. During each run, features were classified as “confirmed” if their importance significantly exceeded the maximum shadow feature importance (p < 0.05) [25]. Feature stability was quantified by calculating the selection frequency (SF), which is defined as the ratio of confirmed designations to total runs. Only features that demonstrated high stability (SF ≥ 0.9) were retained for subsequent XGBoost analysis [26]. The XGBoost feature selection phase employed Pearson correlation-based collinearity analysis to remove redundant features, followed by hyperparameter optimization via grid search with three repetitions of 5-fold cross-validation, learning curve analysis, and early stopping to prevent overfitting—all of which were implemented to enhance clinical interpretability and enable SHAP analysis. To address class imbalance in sepsis mortality outcomes, the precision-recall area under the curve (PR-AUC) was employed as a more informative metric compared to the conventional receiver operating characteristic AUC (ROC-AUC) for evaluating model performance, with particular emphasis being focused on the accurate detection of the minority class. The model’s performance was then comprehensively assessed by using a suite of metrics, including accuracy, sensitivity, specificity, and F1-score metrics, in order to ensure a robust evaluation of its clinical utility. The final model was interpreted using Shapley additive explanations (SHAP) to assess the magnitude and direction of the individual feature influences [27,28]; additionally, the top-ranking NLR-associated variables were subsequently entered into multivariable regression analysis for clinical validation.

Statistical analysis

Statistical analyses and visualizations were performed by using the R programming language (version 4.4.3), with the exception of XGBoost modeling, which was implemented in Python (version 3.12). A two-sided p value <0.05 was considered to be statistically significant. Categorical variables were analyzed by using chi-square tests or Fisher’s exact tests and are presented as frequencies (percentages). To control the false discovery rate (FDR) due to multiple comparisons across categorical covariates (e.g., antibiotic initiation, dialysis type and ventilation status), p values were adjusted by using the Benjamini–Hochberg procedure. For continuous variables, we assessed normality by histograms and skewness analysis (absolute skewness ≥2 with evident left or right skewing in histograms was considered to indicate a nonnormal distribution). Normally distributed continuous variables were compared using Student’s t tests or one-way analysis of variance, whereas nonnormally distributed variables were assessed via Wilcoxon or Kruskal‒Wallis tests, with the results being expressed as the means ± standard deviations (SDs) or medians (interquartile ranges, or IQRs). We examined potential nonlinear associations between the NLR and 28-day mortality using restricted cubic splines (RCS) regression, which tests the statistical significance of nonlinear terms. We simultaneously generated PR and ROC curves for the NLR, XGBoost model, and both the SOFA and SAPS II scoring systems and calculated the Youden index to determine the optimal cutoff values and evaluate their ability to predict 28-day mortality. Multivariate regression models were used to assess the associations between the NLR and 28-day mortality, hospital mortality, and ICU mortality. All included continuous variables were standardized using Z scores to minimize scale differences. Three sequential models were constructed: Model 1 included only the standardized NLR, Model 2 added sex, standardized age, heart rate (HR), respiratory rate (RR), weight, and mean arterial pressure (MAP) to Model 1, and Model 3 further incorporated clinically relevant variables, previously established predictors, and features selected through the Boruta algorithm and XGBoost modeling. Subgroup analyses were performed to further examine the NLR–mortality relationship, with subgroups defined according to clinical relevance and data distribution patterns. Interaction terms were included to assess subgroup heterogeneity, and Firth’s penalized likelihood regression was applied to prevent false-positive results in small sample sizes. Interaction p values were also adjusted for the FDR to comprehensively evaluate potential effect modifiers and the robustness of the NLR predictive value. Survival was analyzed by using Kaplan‒Meier curves with log-rank tests for between-group comparisons and Cox proportional hazards models, with the results being reported as hazard ratios with 95% confidence intervals.

Results

Baseline characteristics

In this study, 4,376 eligible patients from the MIMIC-IV database (with 28-day, in-hospital, and ICU mortality rates of 18.40%, 17.50%, and 14.12%, respectively) were ultimately enrolled. The demographic characteristics, vital signs, laboratory tests, organ dysfunction assessments, treatments and other outcomes are summarized in Table 1. Compared with nonsurvivors, survivors were significantly younger (59.77 ± 14.97 vs. 63.37 ± 15.10; p < 0.05), were more likely to be male (63.93% vs. 59.25%), and had lower disease severity, as reflected by SOFA (3.76 ± 1.96 vs. 4.99 ± 2.71) and SAPS II scores (36.79 ± 12.90 vs. 51.74 ± 15.76). With respect to vital signs, survivors demonstrated lower HR (89.48 ± 17.68 vs. 95.55 ± 20.99 bpm; p < 0.001), lower RR (19.07 ± 5.54 vs. 22.11 ± 6.29 breaths/min; p < 0.001), and higher percutaneous oxygen saturation (SpO₂, median: 98.57% vs. 97.00%; p < 0.001) than nonsurvivors, whereas MAP did not significantly differ between the groups (81.58 ± 15.22 vs. 81.81 ± 19.81 mmHg; p = 0.755). Although survivors demonstrated marginally better vital signs (all p < 0.05 except SpO₂), the intergroup differences were clinically modest. Laboratory analyses revealed significantly elevated inflammatory markers, including the NLR (6.97 [IQR 4.11–12.59] vs. 12.70 [IQR 6.92–22.75]; p < 0.05) and the SII, alongside more pronounced acidotic patterns on blood gas analysis, in nonsurvivors. Higher creatinine and blood urea nitrogen (BUN) levels suggested increased renal dysfunction in nonsurvivors, which is consistent with the results of the organ dysfunction assessments. Erythrocyte and hemoglobin-related parameters significantly differed between the two groups. Organ dysfunction distributions revealed greater respiratory and coagulation dysfunction in survivors, whereas nonsurvivors had greater hepatic, neurological, and renal failure rates (all p < 0.05). Comorbidity analyses revealed significant differences in hypertension, heart failure, chronic obstructive pulmonary disease, chronic kidney disease, atrial fibrillation, cerebral infarction, cerebral hemorrhage, and thrombosis. Fluid resuscitation metrics indicated minimal crystalloid administration within the first 3 ICU hours in both groups, with survivors receiving less crystalloid. Although statistically significant differences were observed in weight-adjusted crystalloid ratios, these differences lacked clinical relevance. Analysis of 24-hour fluid balance revealed higher total intake and output volumes in survivors than in nonsurvivors, whereas nonsurvivors exhibited greater net positive fluid balance, suggesting potential fluid overload in this group. Nonsurvivors required more MV, CRRT, and vasopressor support (all p < 0.05), whereas survivors had higher rates of antibiotic initiation within 1 hour of ICU admission. Although survivors had shorter ICU stays, their total hospitalization duration exceeded that of nonsurvivors.

Table 1. Baseline characteristics of the survivor and nonsurvivor groups.

Categories Survivors (N = 3,571) Non-survivors (N = 805) P-value
Demographic
Age, years 59.77 ± 14.97 63.37 ± 15.10 <0.001
Gender(male), n (%) 2283 (63.93%) 477 (59.25%) 0.012
SOFA 3.76 ± 1.96 4.99 ± 2.71 <0.001
SAPS II 36.79 ± 12.90 51.74 ± 15.76 <0.001
Weight, kg 89.12 ± 23.70 87.44 ± 27.64 0.111
Vital Signs
Heart Rate, bpm 89.48 ± 17.68 95.55 ± 20.99 <0.001
Respiratory Rate, bpm 19.07 ± 5.54 22.11 ± 6.29 <0.001
MAP, mmHg 81.58 ± 15.22 81.81 ± 19.81 0.755
SpO2, % 98.57 (96.00–99.40) 97.00 (94.00–100.00) <0.001
Laboratory Tests
Platelets, K/uL 185.71 ± 102.27 192.85 ± 115.18 0.106
Neutrophils, K/uL 10.84 ± 6.05 13.53 ± 8.30 <0.001
Lymphocytes, K/uL 1.34 (0.81–2.01) 0.91 (0.52–1.50) <0.001
WBC, K/uL 13.34 ± 6.72 16.08 ± 9.60 <0.001
NLR 6.97 (4.11–12.59) 12.70 (6.92–22.75) <0.001
SII 1113.95 (579.22–2470.93) 2083.27 (913.42–4536.00) <0.001
PO2, mmHg 221.07 ± 115.04 151.38 ± 86.71 <0.001
PCO2, mmHg 41.00 (37.80–45.00) 40.92 (36.30–46.44) 0.716
pH 7.36 ± 0.08 7.31 ± 0.12 <0.001
BE, mEq/L −1.22 ± 4.19 −4.22 ± 6.29 <0.001
Bicarbonate, mEq/L 22.24 ± 4.25 20.40 ± 5.88 <0.001
BUN, mg/dL 17.00 (12.00–26.00) 31.00 (19.00–52.00) <0.001
Creatinine, mg/dL 0.90 (0.70–1.30) 1.60 (1.00–2.70) <0.001
Hemoglobin, g/dL 10.51 ± 2.29 10.91 ± 2.68 <0.001
Hematocrit, % 31.75 ± 6.83 33.71 ± 8.12 <0.001
RBC, m/uL 3.50 ± 0.78 3.61 ± 0.93 0.001
RDW, % 14.34 ± 1.91 15.91 ± 2.85 <0.001
MCH, pg 30.18 ± 2.53 30.47 ± 3.04 0.013
MCHC, g/dL 33.12 ± 1.59 32.38 ± 1.82 <0.001
Organ Dysfunction, n (%)
Respiration 1564 (43.8%) 250 (31.06%) <0.001
Coagulation 1528 (42.79%) 285 (35.4%) <0.001
Hepatic 494 (13.83%) 234 (29.07%) <0.001
Cardiovascular 2289 (64.1%) 507 (62.98%) 0.559
Neurologic 543 (15.21%) 145 (18.01%) 0.048
Kidney 1051 (29.43%) 435 (54.04%) <0.001
Comorbidities, n (%)
Diabetes 982 (27.5%) 247 (30.68%) 0.082
Hypertension 1786 (50.01%) 361 (44.84%) 0.007
Heart Failure 845 (23.66%) 252 (31.3%) <0.001
AMI 49 (1.37%) 18 (2.24%) 0.086
COPD 171 (4.79%) 74 (9.19%) <0.001
CKD 459 (12.85%) 192 (23.85%) <0.001
Atrial Fibrillation 979 (27.42%) 270 (33.54%) 0.001
Cerebral Infarction 335 (9.38%) 125 (15.53%) <0.001
Cerebral Hemorrhage 45 (1.26%) 36 (4.47%) <0.001
Thrombosis 372 (10.42%) 138 (17.14%) <0.001
Treatment
Crystalloid Volume in 3 h, mL 80.50 (0.00–350.00) 153.00 (0.00–558.40) <0.001
Fluid/Weight in 3 h, mL/kg 0.85 (0.00–4.24) 1.89 (0.00–6.74) <0.001
Fluid Input in 24 h, mL 4939.43 (2247.60–7224.11) 3224.80 (1327.82–6880.24) <0.001
Fluid Output in 24 h, mL 2580.00 (1755.00–3540.00) 1370.00 (610.00–2520.00) <0.001
Fluid Balance in 24 h, mL 2373.99 ± 4304.62 3022.37 ± 5097.93 <0.001
Fluid Input/Weight in 24 h, mL/kg 57.01 (24.69–86.47) 38.11 (16.38–79.78) <0.001
Antibiotic Initiation (less than 1 h), n (%) 1074 (30.1%) 172 (21.4%) <0.001
VIS 0.00 (0.00–5.00) 8.60 (0.00–40.06) <0.001
CRRT, n (%) 174 (4.9%) 234 (29.1%) <0.001
Mechanical Ventilation, n (%) 2422 (67.8%) 604 (75.0%) <0.001
Outcomes
LOS of Hospital, days 8.84 (5.66–15.75) 6.71 (2.93–12.98) <0.001
LOS of ICU, days 2.94 (1.62–6.20) 4.93 (2.48–9.29) <0.001

Data: Mean ± Standard Deviation or Median (Q1–Q3) or N (%). P-value < 0.05 was considered statistical significance.

SOFA, Sequential Organ Failure Assessment; SAPS II, Simplified Acute Physiology Score II; MAP, Mean Arterial Pressure; SpO2, Peripheral Oxygen Saturation; WBC, White Blood Cell; NLR, Neutrophil-to-Lymphocyte Ratio; SII, Systemic Immune-Inflammation Index; PO2, Partial Pressure of Oxygen; PCO2, Partial Pressure of Carbon Dioxide; BE, Base Excess; BUN, Blood Urea Nitrogen; RBC, Red Blood Cell; RDW, Red Cell Distribution Width; MCH, Mean Corpuscular Hemoglobin; MCHC, Mean Corpuscular Hemoglobin Concentration; AMI, Acute Myocardial Infarction; COPD, Chronic Obstructive Pulmonary Disease; CKD, Chronic Kidney Disease; CRRT, Continuous Renal Replacement Therapy; VIS, Vasoactive-Inotropic Score; LOS, Length of Stay.

On the basis of the overall NLR distribution and quartile values, patients were categorized into low (<4.34), intermediate (4.34–14.70), and high (>14.70) NLR groups (as shown in Table 2). All mortality outcomes, including 28-day, hospital, and ICU mortality, significantly increased in a stepwise manner with increasing NLR (p < 0.05) until they reached approximately 30% in the high-NLR group. Similarly, both hospital and ICU length of stay increased in a concentration-dependent manner (p < 0.05). No statistically significant differences were observed in mean corpuscular hemoglobin (MCH) levels, cardiovascular failure incidence, or the incidence of comorbidities such as diabetes, atrial fibrillation, or cerebral infarction. However, all other parameters that demonstrated significant differences between the survivor and nonsurvivor groups maintained concentration-dependent statistical significance across NLR stratifications (p < 0.05).

Table 2. Characteristics and outcomes of patients categorized by NLR.

Categories NLR < 4.34

(N = 1,094)
NLR 4.34–14.70

(N = 2185)
NLR > 4.34

(N = 1,094)
P-value
Demographic
Age, years 61.20 ± 14.44 59.87 ± 15.21 60.79 ± 15.32 0.039
Gender (male), n (%) 665 (60.79%) 1445 (66.04%) 650 (59.41%) <0.001
SOFA 3.76 ± 1.96 3.92 ± 2.10 4.36 ± 2.44 <0.001
SAPS II 36.60 ± 13.33 38.63 ± 14.17 44.30 ± 15.76 <0.001
Weight, kg 86.73 ± 22.50 89.91 ± 23.44 88.70 ± 28.02 <0.001
Vital Signs
Heart Rate, bpm 85.27 ± 16.05 90.75 ± 17.83 95.61 ± 20.49 <0.001
Respiratory Rate, bpm 17.44 ± 4.52 19.46 ± 5.69 22.16 ± 6.20 <0.001
MAP, mmHg 80.29 ± 11.95 81.83 ± 16.16 82.54 ± 19.44 <0.001
SpO2, % 98.92 (97.66–99.47) 98.52 (96.00–99.75) 97.00 (94.00–99.00) <0.001
Laboratory Tests
Platelets, K/uL 159.18 ± 81.02 189.82 ± 102.86 209.29 ± 122.19 <0.001
Neutrophils, K/uL 6.84 ± 3.81 11.12 ± 5.22 16.25 ± 7.83 <0.001
Lymphocytes, K/uL 2.09 (1.48–2.86) 1.34 (0.96–1.84) 0.57 (0.37–0.84) <0.001
WBC, K/uL 9.93 ± 5.30 13.63 ± 6.38 18.20 ± 8.71 <0.001
SII 440.61 (291.56–620.37) 1267.47 (798.13–2038.41) 4795.85 (2935.87–7612.25) <0.001
PO2, mmHg 261.23 ± 119.21 209.82 ± 111.95 152.14 ± 80.15 <0.001
PCO2, mmHg 41.34 ± 7.96 42.10 ± 8.69 42.82 ± 11.27 0.001
pH 7.38 ± 0.09 7.36 ± 0.09 7.33 ± 0.09 <0.001
BE, mEq/L 0.00 (−1.84–1.71) −0.68 (−3.67–1.00) −3.00 (−6.00–-0.35) <0.001
Bicarbonate, mEq/L 22.43 ± 3.86 22.15 ± 4.42 20.90 ± 5.56 <0.001
BUN, mg/dL 16.00 (12.00–21.00) 18.00 (13.00–28.36) 25.50 (16.00–47.00) <0.001
Creatinine, mg/dL 0.85 (0.70–1.10) 1.00 (0.70–1.50) 1.30 (0.90–2.40) <0.001
Hemoglobin, g/dL 9.98 ± 2.06 10.63 ± 2.37 11.11 ± 2.53 <0.001
Hematocrit, % 30.11 ± 6.15 32.15 ± 7.04 34.05 ± 7.64 <0.001
RBC, m/uL 3.31 ± 0.71 3.53 ± 0.81 3.71 ± 0.87 <0.001
RDW, % 13.70 (13.00–14.80) 14.00 (13.20–15.30) 14.60 (13.50–16.10) <0.001
MCH, pg 30.32 ± 2.39 30.24 ± 2.62 30.12 ± 2.87 0.225
MCHC, g/dL 33.14 ± 1.55 33.09 ± 1.67 32.63 ± 1.70 <0.001
Organ Dysfunction, n (%)
Respiration 512 (46.8%) 967 (44.2%) 335 (30.62%) <0.001
Coagulation 611 (55.85%) 875 (39.99%) 327 (29.89%) <0.001
Hepatic 121 (11.06%) 340 (15.54%) 267 (24.41%) <0.001
Cardiovascular 720 (65.81%) 1387 (63.39%) 689 (62.98%) 0.304
Neurologic 147 (13.44%) 349 (15.95%) 192 (17.55%) 0.028
Kidney 244 (22.3%) 717 (32.77%) 525 (47.99%) <0.001
Comorbidities, n (%)
Diabetes 302 (27.61%) 610 (27.88%) 317 (28.98%) 0.741
Hypertension 590 (53.93%) 1080 (49.36%) 477 (43.6%) <0.001
Heart Failure 228 (20.84%) 532 (24.31%) 337 (30.8%) <0.001
AMI 6 (0.55%) 39 (1.78%) 22 (2.01%) 0.008
COPD 45 (4.11%) 105 (4.8%) 95 (8.68%) <0.001
CKD 130 (11.88%) 302 (13.8%) 219 (20.02%) <0.001
Atrial Fibrillation 297 (27.15%) 630 (28.79%) 322 (29.43%) 0.464
Cerebral Infarction 113 (10.33%) 222 (10.15%) 125 (11.43%) 0.516
Cerebral Hemorrhage 11 (1.01%) 39 (1.78%) 31 (2.83%) 0.006
Thrombosis 78 (7.13%) 237 (10.83%) 195 (17.82%) <0.001
Treatment
Crystalloid Volume in 3 h, mL 50.00 (0.00–320.25) 98.45 (0.00–393.51) 106.38 (0.00–508.32) <0.001
Fluid/Weight in 3 h, mL/kg 0.64 (0.00–3.86) 0.98 (0.00–4.60) 1.35 (0.00–6.15) <0.001
Fluid Input in 24 h, mL 5306.60 (2871.38–7256.89) 4886.59 (2117.32–7349.11) 3394.81 (1347.47–6590.64) <0.001
Fluid Output in 24 h, mL 2700.00 (1960.50–3563.25) 2505.00 (1590.00–3483.50) 1807.50 (1015.00–2930.00) <0.001
Fluid Balance in 24 h, mL 2510.63 (347.47–4356.38) 2260.98 (−101.81–4497.61) 1325.28 (−520.00–4584.98) <0.001
Fluid Input/Weight in 24 h, mL/kg 61.92 (33.08–88.92) 55.86 (22.91–86.79) 39.45 (15.66–76.70) <0.001
Antibiotic Initiation (less than 1 h), n (%) 388 (35.47%) 625 (28.56%) 233 (21.3%) <0.001
VIS 0.00 (0.00–3.00) 0.00 (0.00–8.00) 0.00 (0.00–20.00) <0.001
CRRT, n (%) 54 (4.94%) 176 (8.04%) 178 (16.27%) <0.001
Mechanical Ventilation, n (%) 783 (71.57%) 1535 (70.16%) 708 (64.72%) 0.001
Outcomes
LOS of Hospital, days 7.10 (5.08–11.49) 8.47 (5.23–14.69) 10.84 (5.89–19.71) <0.001
LOS of ICU, days 2.32 (1.38–4.44) 3.18 (1.81–6.85) 4.75 (2.43–9.77) <0.001
Hospital Mortality, n (%) 93 (8.5%) 333 (15.22%) 340 (31.08%) <0.001
ICU Mortality, n (%) 79 (7.22%) 267 (12.2%) 272 (24.86%) <0.001
28-day Mortality, n (%) 101 (9.23%) 347 (15.86%) 357 (32.63%) <0.001

Data: Mean ± Standard Deviation or Median (Q1–Q3) or N (%). P-value < 0.05 were considered statistical significance.

SOFA, Sequential Organ Failure Assessment; SAPS II, Simplified Acute Physiology Score II; MAP, Mean Arterial Pressure; SpO2, Peripheral Oxygen Saturation; WBC, White Blood Cell; NLR, Neutrophil-to-Lymphocyte Ratio; SII, Systemic Immune-Inflammation Index; PO2, Partial Pressure of Oxygen; PCO2, Partial Pressure of Carbon Dioxide; BE, Base Excess; BUN, Blood Urea Nitrogen; RBC, Red Blood Cell; RDW, Red Cell Distribution Width; MCH, Mean Corpuscular Hemoglobin; MCHC, Mean Corpuscular Hemoglobin Concentration; AMI, Acute Myocardial Infarction; COPD, Chronic Obstructive Pulmonary Disease; CKD, Chronic Kidney Disease; CRRT, Continuous Renal Replacement Therapy; VIS, Vasoactive-Inotropic Score; LOS, Length of Stay; ICU, Intensive Care Unit.

Selection of features in the models

The results from a single Boruta run are presented in S1 Fig, whereas Fig 2A displays stability-selected features (SF ≥ 0.9) from 10 iterations; this process identified 37 confirmed and 2 tentative features, whereas the remaining features were rejected. These 39 features were subjected to XGBoost modeling, which was preceded by collinearity analysis (Pearson’s r ≥ 0.7 exclusion) and clinical relevance screening (S2 Fig and S3 Table), thus yielding 29 final parameters. As shown in S4 Table, the optimized XGBoost model achieved a test-set ROC-AUC of 0.875 (95% CI: 0.854–0.896) and a PR-AUC of 0.603 (95% CI: 0.534–0.671), thereby outperforming the individual NLR, as well as the SAPS II and SOFA scores, for 28-day mortality prediction (Fig 3A, 3B). The data in Fig 2B reveal 13 important parameters, and Model 3 ultimately incorporated SAPS II and SOFA scores, as well as 24-hour fluid output, VIS, 24-hour fluid balance, dialysis type, cerebral infarction, atrial fibrillation, red cell distribution width (RDW), partial pressure of oxygen (PO2), hematocrit, creatinine, LYMPH, base excess (BE), and MCH. S5 Table presents the predictive contribution ranking of the NLR in the XGBoost model for 28-day mortality prediction. The NLR ranked 8th in global importance but decreased to 14th in terms of the SHAP-based marginal contribution, which indicates its relatively low individual impact and potential synergistic interactions with other parameters.

Fig 2. Application of machine learning algorithm for feature selection.

Fig 2

A Feature stability selection for the relationship between the NLR and 28-day mortality was analyzed by using the Boruta algorithm. The x-axis lists parameters (actual and shadow features), whereas the y-axis shows noise-adjusted importance scores from 10 Boruta runs. Features are color-coded by stability: light blue (confirmed, SF ≥ 90%), green (tentative, 50% ≤ SF < 90%), red (rejected, SF < 50%), and gray (shadow features). A red dashed line marks the stability threshold. B SHAP analysis of XGBoost model predictions. Features are vertically ranked by the mean absolute SHAP value (light blue bars), with individual SHAP values (colored dots) distributed horizontally and colored by normalized feature magnitude (blue: low, red: high). The gray dashed line marks the baseline.

Fig 3. Comparative analysis of predictive models and survival stratification by the NLR on 28-day mortality.

Fig 3

Panel A: receiver operating characteristic (ROC) curves demonstrate the discriminative performance of the NLR, SAPS II, SOFA, and XGBoost; Panel B: precision-recall (PR) curves assessing the trade-off between precision and recall; Panel C: Kaplan-Meier survival analysis stratified by NLR quartiles.

NLR and 28-day, hospital, and ICU mortality

Multivariable logistic regression analysis demonstrated that the standardized NLR was independently associated with 28-day mortality (OR 1.16; 95% CI 1.06–1.27; p < 0.001), in-hospital mortality (OR 1.13; 95% CI 1.03–1.24; p < 0.001), and ICU mortality (OR 1.14; 95% CI 1.03–1.25; p = 0.008) after full adjustment (Model 3, Table 3). Unadjusted analysis by NLR quartiles revealed a dose‒response relationship with all of the mortality endpoints (S3A–C Fig), and stronger associations were observed with increasing NLR values. In the fully adjusted model (Model 3), patients in the highest NLR quartile exhibited significantly higher risks of 28-day mortality (OR 1.95, 95% CI 1.40–2.75), in-hospital mortality (OR 1.90, 95% CI 1.34–2.70), and ICU mortality (OR 1.79, 95% CI 1.23–2.62) compared to those in the lowest quartile, and all of the trend tests were statistically significant (p < 0.001).(Fig 3C) temporally corroborated these findings, which demonstrates early divergence among groups by day 7 (log-rank p < 0.0001), whereas the high-NLR group exhibited markedly decreased survival. Restricted cubic spline analysis (S4 Fig) demonstrated biphasic nonlinear relationships between the NLR and all of the mortality endpoints, with Wald tests confirming significant nonlinear effects across all of the models (P-nonlinear <0.05). The ROC-derived threshold (NLR = 7.3) coincided with both the median NLR value and the first inflection point, where the predicted probabilities sharply increased. This predictive probability plateaued near the second inflection point (NLR = 27.3) and increased marginally from 0.15 to 0.30, even though sample representation markedly declined in this range. Beyond the second inflection point, statistical reliability diminished because of sparse data.

Table 3. The association between NLR groups and 28-day mortality, hospital mortality and ICU mortality.

Exposure Model 1 Model 2 Model 3
OR (95% CI) P-value OR (95% CI) P-value OR (95% CI) P-value
28-day mortality
NLR as continuous 1.50 (1.40–1.60) <0.001 1.36 (1.27–1.45) <0.001 1.16 (1.06–1.27) <0.001
R1 Ref Ref Ref
R2 1.85 (1.47–2.35) <0.001 1.61 (1.27–2.05) <0.001 1.34 (1.01–1.80) 0.048
R3 4.76 (3.76–6.08) <0.001 3.34 (2.60–4.31) <0.001 1.95 (1.40–2.75) <0.001
P for trend <0.001 <0.001 <0.001
Hospital mortality
NLR as continuous 1.49 (1.39–1.59) <0.001 1.34 (1.25–1.43) <0.001 1.13 (1.03–1.24) 0.008
R1 Ref Ref Ref
R2 1.93 (1.52–2.47) <0.001 1.63 (1.28–2.10) <0.001 1.37 (1.02–1.86) 0.040
R3 4.85 (3.80–6.25) <0.001 3.28 (2.54–4.27) <0.001 1.90 (1.34–2.70) <0.001
P for trend <0.001 <0.001 <0.001
ICU mortality
NLR as continuous 1.44 (1.34–1.54) <0.001 1.30 (1.21–1.39) <0.001 1.14 (1.03–1.25) 0.008
R1 Ref Ref Ref
R2 1.79 (1.38–2.33) <0.001 1.48 (1.14–1.95) 0.004 1.31 (0.95–1.82) 0.109
R3 4.25 (3.27–5.58) <0.001 2.82 (2.14–3.75) <0.001 1.79 (1.23–2.62) 0.003
P for trend <0.001 <0.001 <0.001

NLR, Neutrophil-to-Lymphocyte Ratio; R1, low-NLR group; R2, intermediate-NLR group; R3, high-NLR group

Model 1: Unadjusted

Model 2: Adjusted for gender, standardized age, HR, RR, weight, and MAP

Model 3: Adjusted for gender, cerebral infarction, atrial fibrillation, standardized age, HR, RR, weight, MAP, SAPS II, SOFA, 24-hour fluid output, VIS, 24-hour fluid balance, dialysis type, RDW, PO2, HCT, creatinine, lymphocytes, BE, and MCH

Subgroup analysis

To further evaluate the independent association between the standardized NLR and mortality outcomes, including 28-day, in-hospital, and ICU mortality, we conducted stratified analyses and interaction tests on Model 3 covariates. As demonstrated in Fig 4 and S6 Table, the NLR maintained consistent associations with mortality, with increased predictive value being observed in patients aged >45 years, those with SOFA scores ≤4 and those with SAPS II scores > 29. Those individuals ≥65 years exhibited a 17% increased 28-day mortality risk (P = 0.019), and those with a SOFA score ≤4 exhibited a > 20% increased risk across all of the mortality endpoints (p < 0.001), whereas no significant association was observed in the subgroup of patients with SOFA scores ≥9 (p = 0.369). Significant effect modifications (interaction p < 0.05) were identified in patients with a 24-hour fluid output of 1475–3420 mL, a BE > -2, a HR < 100 bpm, or a RR of 12–20 breaths/min, and similar trends were observed in those with a 24-hour fluid balance ≤2493.27 mL, an MCH ≥ 27 pg, a MAP of 65–110 mmHg, or a creatinine concentration of 1.5–3.0 mg/dL. Although the standardized NLR demonstrated strong associations with all three mortality outcomes in patients with an HR < 60 bpm, the limited sample size within this stratified subgroup resulted in imprecise estimates, thus warranting cautious interpretation. The subgroup analyses consistently demonstrated significant associations between the NLR and all three mortality endpoints across most predefined strata, which confirms the robustness of the predictive value of this biomarker in diverse patient populations.

Fig 4. Forest plot of subgroup analyses for the association between the standardized NLR and mortality outcomes.

Fig 4

The plot displays odds ratios (ORs) with 95% confidence intervals (CIs) for 28-day, in-hospital, and ICU mortality across clinically relevant subgroups. Solid squares represent point estimates of ORs; horizontal lines indicate 95% CIs. The vertical dashed line denotes the null effect (OR = 1).

Discussion

Machine learning algorithms have increasingly been employed to refine prognostic assessment in sepsis using large-scale ICU databases. For example, Lou et al. [29] and Zheng et al. [30] recently applied XGBoost based on the MIMIC database to explore the association between the triglyceride-glucose (TyG) index and risk of death in sepsis patients, thus further confirming the potential of ML in routine biomarker re-evaluation. Within this evolving landscape, as an accessible and widely recorded inflammatory marker, the NLR has garnered renewed interest; however, its independent prognostic value for short-term mortality remains debated, and systematic ML‑driven re-evaluations remain scarce. In the present study, we leveraged a rigorous two-stage feature selection pipeline combining the Boruta algorithm with XGBoost to specifically evaluate the ability of the early NLR to predict 28‑day mortality in a large, real‑world sepsis cohort. Our optimized model achieved robust discrimination (ROC‑AUC 0.875; PR‑AUC 0.603), thus outperforming conventional severity scores. More importantly, after full multivariable adjustment, each 1‑SD increase in the standardized NLR was independently associated with a 16%, 13%, and 14% higher risk of 28‑day, in‑hospital, and ICU mortality, respectively. Dose‑response analyses confirmed a progressive mortality increase with increasing NLR, whereas RCS revealed a nonlinear, biphasic pattern; specifically, predictive capacity became increased near the first inflection point (NLR = 7.3, approximating the optimal threshold) and plateaued beyond the second inflection point (NLR = 27.3). Subgroup analyses further elucidated that the NLR-mortality association was most pronounced in hemodynamically stable patients and those with low organ failure scores (SOFA ≤ 4), whereas it lost statistical significance in the setting of advanced multi‑organ dysfunction (SOFA ≥ 9). Collectively, these findings do not merely replicate prior prognostic observations; rather, they delineate (for the first time) a phenotype‑specific, biphasic predictive pattern of the NLR that distinguishes inflammation‑driven mortality risk from organ failure‑driven mortality. By integrating state‑of‑the‑art feature selection with granular clinical stratification, this study provides a methodologically rigorous and clinically nuanced re-evaluation of the NLR’s prognostic utility in sepsis.

Our XGBoost model achieved a test-set ROC-AUC of 0.875 (PR-AUC 0.603) for 28-day mortality prediction, thus representing a substantial improvement over conventional severity scores and individual biomarkers. To contextualize this performance within the broader landscape of emerging prognostic indices in sepsis, we compared our findings with recent studies evaluating the TyG index and the creatinine-to-albumin ratio (CAR) using the MIMIC database [29–31]. Both TyG-based studies reported significant associations with mortality, with Lou et al. [29] identifying a nonlinear relationship (inflection point: TyG = 8.9) and Zheng et al. [30] describing a linear dose-response pattern. The CAR study by Lou et al. [31] demonstrated a J-shaped association with 30-day mortality in sepsis-associated acute kidney injury, with an inflection point at CAR = 1.2 mg/dL and an AUC of 0.68–0.75 being reported. Notably, our NLR-based XGBoost model achieved higher discriminative performance (AUC 0.875) than any of these reported models, which likely reflects the synergistic integration of multiple clinical and laboratory features through a robust two-stage feature selection pipeline. Beyond performance metrics, a key methodological parallel involves the nonlinear risk patterns shared across these biomarkers. The J-shaped CAR-mortality curve and the biphasic NLR-mortality relationship observed in our cohort both suggest threshold effects that demarcate distinct pathophysiological states, including renal dysfunction with systemic inflammation for CAR and innate-adaptive immune imbalance for the NLR. However, the NLR offers distinct practical advantages; specifically, it is derived from a complete blood count without the need for additional metabolic assays, incurs no extra cost, and directly reflects the hyperinflammation–immunosuppression disequilibrium that is a central process to sepsis immunopathology [15,17,18,32]. Thus, although the TyG and CAR capture insulin resistance and renal-inflammatory crosstalk, respectively, the NLR (particularly when embedded in a ML framework) provides complementary prognostic information with superior discrimination and unique biological interpretability. These findings demonstrate the NLR not as a replacement for existing biomarkers; rather, it can function as a phenotype-specific adjunct for early risk stratification in moderate sepsis, where its predictive value is most pronounced.

Recent five-year evidence has consistently demonstrated that an elevated NLR increases sepsis mortality risk by 10–40% across general populations [33–42]. More critically, when high-NLR cohorts were compared with their low-NLR counterparts, multivariable-adjusted analyses revealed that the high-NLR group had an 80% greater mortality risk, which establishes a robust dose‒response relationship. Three studies [39,41,42] employing RCS analysis confirmed nonlinear NLR-mortality relationships and demonstrated progressive risk escalation until the NLR reached a range of 15–20, beyond which the trend plateaued. Prior evaluations of the discriminative capacity of the NLR universally relied on ROC curves, with AUROC values ranging from 0.6 to 0.7, which are consistent with our findings and indicate moderate discrimination. However, given the imbalanced outcome distribution (18.4% 28-day mortality in our cohort), ROC metrics may overestimate performance because of excessive true-negative influence, which potentially masks true positive-class identification capacity [32,43,44]. We therefore performed ROC analyses with PR curves and confirmed that the NLR, SOFA score, and SAPS II score all exhibit limited mortality discrimination under outcome imbalance. XGBoost modeling that incorporated multiple parameters significantly improved both the AUC and average precision values. While the NLR demonstrated a substantial overall model contribution, SHAP analysis revealed marginal individual predictive utility, which suggests limited direct mortality risk stratification capacity. Combined NLR stratification and subgroup analyses revealed potential interaction effects, particularly in hemodynamically stable patients for whom the correlations between the NLR and mortality remained robust.

Returning to our primary research question, the limited predictive value of the NLR for 30-day mortality observed in the study by Schupp et al. [19] may stem from two methodological factors. First, the exclusively enrolled critically ill patients demonstrated uniformly high disease severity (mean SOFA score of 11 and mean APACHE II score of 23), and no significant differences were observed in illness severity between survivors and nonsurvivors. Second, the small sample size may have compromised the stability of the results. Conversely, studies [33,35,36,38,39] that support the ability of the NLR to predict sepsis mortality but that demonstrate superior discriminatory performance for 28-day mortality (with the AUROC value being observed at 0.827 in some reports) have predominantly analyzed cohorts with extreme mortality rates (30–70%) and significant intergroup differences in illness severity scores. Notably, three additional studies reported the limited predictive value of the NLR [37,40,45] after patients with critical illness or hemodynamic instability were excluded through subgroup or sensitivity analyses. These collective observations suggest that the ability of the NLR to predict mortality may be restricted to clinically stable conditions, while its utility appears to decrease in severe cases with concomitant organ dysfunction.

The context‑dependent nature of the NLR’s prognostic value observed in our study aligns with emerging evidence on other composite biomarkers that integrate inflammatory and nutritional dimensions. Recent studies investigating the NEUT count‑to‑prognostic nutritional index ratio (NPNR) in septic patients have reported J‑ or V‑shaped associations with mortality [46,47], thus mirroring the biphasic pattern that we identified for the NLR. Both the NLR and NPNR appear to capture the delicate balance between host immune activation and physiological reserve; specifically, their predictive weight diminishes when organ failure becomes the dominant driver of the outcome. This convergence across distinct biomarker constructs reinforces the notion that inflammatory markers are not universal prognostic tools; rather, they are phenotype‑specific indicators exhibiting a utility that is most pronounced in moderate disease and wanes in the setting of terminal organ dysfunction. Therefore, future research should focus on context‑aware risk stratification that accounts for the dynamic interplay between inflammation, nutrition, and organ function.

Mechanistically, the increase in the NLR in sepsis patients precisely mirrors the dual pathological processes of hyperactivated NEUT and apoptotic LYMPH depletion; thus, the NLR serves as a dynamic biomarker of the “hyperinflammation–immunosuppression” imbalance that characterizes sepsis immunopathology [15,45]. Current evidence indicates that delayed NEUT apoptosis leads to systemic accumulation, whereas accelerated CD4+/CD8+ T-cell apoptosis and impaired NK cell function collectively contribute to this dysregulated state. Notably, the magnitude of NLR elevation may reflect the depletion of anti-inflammatory reserves, and higher values may indicate a poorer capacity to counteract inflammatory insults [45]. Clinically, the NLR demonstrates robust predictive value for moderate sepsis (SOFA scores <9), as evidenced by the finding reported by Liu et al. [37] that serial NLR measurements (AUC = 0.823) outperformed single assessments, as a day-7 NLR > 4.18 predicted significantly increased 28-day mortality. Ye et al. [40] further confirmed that an NLR > 20.25 was an independent predictor of mortality (hazard ratio = 1.22) in large cohorts. However, its ability to predict critical illness (SOFA score ≥11) is diminished, which is likely due to the predominance of multiorgan failure observed in terminal pathophysiology [40]. The primary sources of interstudy variability include discrepancies in measurement timing, divergent threshold selection criteria, and confounding effects of therapeutic interventions such as blood purification. Notably, baseline NLR values inherently differ between specific disease populations and healthy individuals. Karakonstantis et al. [48] reported that multiple factors potentially contribute to nonpathological NLR elevation, including advanced age, exogenous steroid administration, elevated endogenous hormone levels, active hematologic disorders, type 2 diabetes mellitus, and acute trauma. Future research should prioritize three objectives: standardizing measurement protocols, establishing population-specific cutoffs, and investigating combination models with emerging immune biomarkers. Fundamentally, the clinical value of the NLR lies in its noninvasive nature and ability to track immune trajectory evolution rather than as a standalone prognostic determinant.

This study implemented multiple methodological innovations in its investigation of sepsis outcomes. Given the well-documented prognostic influences of key sepsis interventions, including fluid resuscitation, antibiotic therapy, MV, and blood purification [1], we developed advanced data extraction protocols to obtain and systematically categorize relevant parameters from the MIMIC-IV database. Through stepwise feature selection, the analysis consistently identified fluid resuscitation and blood purification as clinically significant factors, while numerous other variables were excluded. A planned evaluation of crystalloid fluid administration effects within the initial three hours after ICU admission was determined to be impractical because of insufficient recorded intravenous infusion volumes. The derived weight-adjusted ratios fell below clinically meaningful thresholds, demonstrated minimal predictive value for mortality outcomes, and were consequently excluded during collinearity analysis and feature selection procedures.

Several limitations must be acknowledged. First, although the MIMIC-IV database provides a large, well-curated sample of critically ill patients, its single-center origin introduces inherent selection bias and reflects the treatment protocols and case mixture of a specific healthcare system. Thus, the absence of external validation represents a substantial limitation. Although we implemented rigorous analytical methods to ensure internal validity (particularly the proposed biphasic, phenotype‑dependent predictive pattern of the NLR), our findings must be regarded as hypothesis‑generating rather than definitive. Independent validation in prospective, multi-center cohorts is urgently required before the NLR can be considered for routine clinical risk stratification. Furthermore, external validation should specifically test the generalizability of the identified inflection points (NLR = 7.3 and 27.3) and the differential predictive utility across SOFA strata observed in our study. Second, while NLR standardization was necessary to address scale disparities in multivariable regression modeling, this processing may reduce clinical interpretability. Additionally, the exclusive calculation of the NLR from the first available measurement within six hours of ICU admission captures only a static snapshot of the inflammatory state. Given the highly dynamic nature of sepsis, this approach may underestimate the prognostic value of NLR trajectories or peak values, which could better reflect the evolving host response. Future studies incorporating serial NLR measurements or longitudinal trajectory modeling are warranted to further validate our findings. Third, the dataset lacked several potentially relevant parameters, including nutritional status indicators, metabolic markers, and established inflammatory biomarkers such as procalcitonin and C-reactive protein, which may affect the generalizability of the results. Fourth, the sepsis management guidelines were revised multiple times during the extended study period spanning 2008–2022, which may have introduced temporal confounding factors that could have influenced the outcomes. Despite these limitations, this study benefits from a substantial sample size and comprehensive statistical methodology that may partially mitigate these concerns. Future research should focus on targeted analyses of specific sepsis patient subgroups to more precisely characterize the prognostic utility of the NLR.

Conclusions

This investigation revealed a positive correlation between the magnitude of NLR elevation and sepsis-related mortality, with the strongest predictive value being observed at the highest NLR values. This correlation may primarily reflect the ability of the NLR to predict the outcomes of clinically stable patients, whereas its predictive value may be reduced for patients with severe organ failure. Thus, the utility of the NLR appears to primarily involve dynamic monitoring rather than single static measurements. However, given the absence of external validation, these conclusions remain hypothesis‑generating and require confirmation in independent, prospective, multi-center studies before clinical implementation.

Supporting information

S1 Fig. Boruta algorithm feature selection results for XGBoost modeling.

Box plots display the Z scores of each parameter, with the x-axis showing their names and the y-axis showing the Z values. Parameters are color-coded by importance (blue for important, green for tentative, and red for unimportant).

(TIF)

pone.0348676.s001.tif (5.2MB, tif)
S2 Fig. Feature collinearity assessment in XGBoost modeling.

(TIF)

pone.0348676.s002.tif (4.8MB, tif)
S3 Fig. Unadjusted associations between NLR levels and mortality outcomes.

Panels A-C correspond to 28-day mortality (A), hospital mortality (B), and ICU mortality (C), respectively. Error bars represent 95% CIs, with the intermediate NLR concentration group serving as the reference category.

(TIF)

pone.0348676.s003.tif (511.8KB, tif)
S4 Fig. Restricted cubic spline plots illustrating the nonlinear relationship between the NLR and mortality.

(A) 28-day mortality, (B) in-hospital mortality, and (C) ICU mortality. The solid line represents the adjusted odds ratio; the pink band indicates the 95% confidence interval. The histogram displays the distribution of the NLR values, with the median (7.6) and the ROC-derived optimal cutoff (8.08) marked. The nonlinear association was statistically significant for all of the endpoints (P for nonlinearity < 0.05).

(TIFF)

pone.0348676.s004.tiff (1.4MB, tiff)
S1 Table. Missing rate for demographics and clinical variables extracted from the database during the observation period.

(PDF)

pone.0348676.s005.pdf (101.6KB, pdf)
S2 Table. Comparison of Key Variables Before and After Imputation Using the misForest Algorithm.

(PDF)

pone.0348676.s006.pdf (94.4KB, pdf)
S3 Table. Selected features and discarded features of collinearity assessment in XGBoost modeling.

(PDF)

pone.0348676.s007.pdf (102KB, pdf)
S4 Table. Performance evaluation of XGBoost model across training and test datasets.

(PDF)

pone.0348676.s008.pdf (118.9KB, pdf)
S5 Table. Feature importances in the XGBoost model.

(PDF)

pone.0348676.s009.pdf (105.6KB, pdf)
S6 Table. Subgroup analysis for the association of the NLR with 28-day mortality, hospital mortality and ICU mortality.

(PDF)

pone.0348676.s010.pdf (88.3KB, pdf)
S7 Table. Abbreviations.

(PDF)

pone.0348676.s011.pdf (100.3KB, pdf)

Acknowledgments

The authors appreciate all of the investigators who organized, developed, and maintained the MIMIC database.

Data Availability

The complete minimal anonymized dataset necessary to replicate all study findings is publicly available from Zenodo (https://doi.org/10.5281/zenodo.20139761).

Funding Statement

This project was supported by grants from the High-level Medical Team Project in Baoan (No. 202405) and Medical and Health Scientific Research Project of Shenzhen Baoan District in 2023 (No. 2023JD118). Funding related to No. 2023JD118 was received by JYL, who conceived and designed the study, supervised the research, performed the data collection and curation, writing-review and editing the manuscript. Funding related to 2023. Funding related to No. 202405 was received by JBL, who contributed to critical revision of the manuscript and participated in the drafting of the original manuscript.

References

  • 1.Evans L, Rhodes A, Alhazzani W, Antonelli M, Coopersmith CM, French C, et al. Surviving sepsis campaign: international guidelines for management of sepsis and septic shock 2021. Crit Care Med. 2021;49(11):e1063–143. doi: 10.1097/CCM.0000000000005337 [DOI] [PubMed] [Google Scholar]
  • 2.Fleischmann-Struzek C, Mellhammar L, Rose N, Cassini A, Rudd KE, Schlattmann P, et al. Incidence and mortality of hospital- and ICU-treated sepsis: results from an updated and expanded systematic review and meta-analysis. Intensive Care Med. 2020;46(8):1552–62. doi: 10.1007/s00134-020-06151-x [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 3.Fleischmann C, Scherag A, Adhikari NK, Hartog CS, Tsaganos T, Schlattmann P, et al. Assessment of global incidence and mortality of hospital-treated sepsis. Current estimates and limitations. Am J Respir Crit Care Med. 2016;193(3):259–72. Epub 2015/09/29. doi: 10.1164/rccm.201504-0781OC [DOI] [PubMed] [Google Scholar]
  • 4.Alberto L, Marshall AP, Walker R, Aitken LM. Screening for sepsis in general hospitalized patients: a systematic review. J Hosp Infect. 2017;96(4):305–15. Epub 2017/05/17. doi: 10.1016/j.jhin.2017.05.005 [DOI] [PubMed] [Google Scholar]
  • 5.Bhattacharjee P, Edelson DP, Churpek MM. Identifying patients with sepsis on the hospital wards. Chest. 2017;151(4):898–907. doi: 10.1016/j.chest.2016.06.020 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6.Warttig S, Alderson P, Evans DJ, Lewis SR, Kourbeti IS, Smith AF. Automated monitoring compared to standard care for the early detection of sepsis in critically ill patients. Cochrane Database Syst Rev. 2018;6(6):CD012404. doi: 10.1002/14651858.CD012404.pub2 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 7.Makam AN, Nguyen OK, Auerbach AD. Diagnostic accuracy and effectiveness of automated electronic sepsis alert systems: a systematic review. J Hosp Med. 2015;10(6):396–402. doi: 10.1002/jhm.2347 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 8.Tang A, Shi Y, Dong Q, Wang S, Ge Y, Wang C, et al. Prognostic differences in sepsis caused by gram-negative bacteria and gram-positive bacteria: a systematic review and meta-analysis. Crit Care. 2023;27(1):467. doi: 10.1186/s13054-023-04750-w [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 9.Raith EP, Udy AA, Bailey M, McGloughlin S, MacIsaac C, Bellomo R, et al. Prognostic accuracy of the SOFA score, SIRS criteria, and qSOFA score for in-hospital mortality among adults with suspected infection admitted to the intensive care unit. JAMA. 2017;317(3):290–300. doi: 10.1001/jama.2016.20328 [DOI] [PubMed] [Google Scholar]
  • 10.Moreno-Torres V, Royuela A, Múñez E, Ortega A, Gutierrez Á, Mills P, et al. Better prognostic ability of NEWS2, SOFA and SAPS-II in septic patients. Med Clin (Barc). 2022;159(5):224–9. doi: 10.1016/j.medcli.2021.10.021 [DOI] [PubMed] [Google Scholar]
  • 11.Tekin B, Kiliç J, Taşkin G, Solmaz İ, Tezel O, Başgöz BB. The comparison of scoring systems: SOFA, APACHE-II, LODS, MODS, and SAPS-II in critically ill elderly sepsis patients. J Infect Dev Ctries. 2024;18(1):122–30. doi: 10.3855/jidc.18526 [DOI] [PubMed] [Google Scholar]
  • 12.Song M, Graubard BI, Rabkin CS, Engels EA. Neutrophil-to-lymphocyte ratio and mortality in the United States general population. Sci Rep. 2021;11(1):464. doi: 10.1038/s41598-020-79431-7 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 13.Li Y, Wang W, Yang F, Xu Y, Feng C, Zhao Y. The regulatory roles of neutrophils in adaptive immunity. Cell Commun Signal. 2019;17(1):147. doi: 10.1186/s12964-019-0471-y [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 14.Lowsby R, Gomes C, Jarman I, Lisboa P, Nee PA, Vardhan M, et al. Neutrophil to lymphocyte count ratio as an early indicator of blood stream infection in the emergency department. Emerg Med J. 2015;32(7):531–4. doi: 10.1136/emermed-2014-204071 [DOI] [PubMed] [Google Scholar]
  • 15.Buonacera A, Stancanelli B, Colaci M, Malatino L. Neutrophil to lymphocyte ratio: an emerging marker of the relationships between the immune system and diseases. Int J Mol Sci. 2022;23(7):3636. doi: 10.3390/ijms23073636 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16.Parthasarathi A, Padukudru S, Arunachal S, Basavaraj CK, Krishna MT, Ganguly K, et al. The role of neutrophil-to-lymphocyte ratio in risk stratification and prognostication of COVID-19: a systematic review and meta-analysis. Vaccines (Basel). 2022;10(8):1233. doi: 10.3390/vaccines10081233 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 17.Snopkowska Lesniak SW, Maschio D, Neria F, Rey-Delgado B, Moreno Cuerda V, Henriquez-Camacho C. Novel biomarkers for SARS-CoV-2 infection: a systematic review and meta-analysis. J Pers Med. 2025;15(6):225. doi: 10.3390/jpm15060225 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 18.Wu H, Cao T, Ji T, Luo Y, Huang J, Ma K. Predictive value of the neutrophil-to-lymphocyte ratio in the prognosis and risk of death for adult sepsis patients: a meta-analysis. Front Immunol. 2024;15:1336456. doi: 10.3389/fimmu.2024.1336456 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19.Schupp T, Weidner K, Rusnak J, Jawhar S, Forner J, Dulatahu F, et al. The neutrophil-to-lymphocyte-ratio as diagnostic and prognostic tool in sepsis and septic shock. Clin Lab. 2023;69(5):10.7754/Clin.Lab.2022.220812. doi: 10.7754/Clin.Lab.2022.220812 [DOI] [PubMed] [Google Scholar]
  • 20.Goldberger A, Amaral L, Glass L, Hausdorff J, Ivanov PC, Mark R, et al. PhysioBank, PhysioToolkit, and PhysioNet: Components of a new research resource for complex physiologic signals. Circulation. 2000;101(23):e215–20. doi: 10.1161/01.CIR.101.23.e215 [DOI] [PubMed] [Google Scholar]
  • 21.Johnson AEW, Bulgarelli L, Shen L, Gayles A, Shammout A, Horng S, et al. MIMIC-IV, a freely accessible electronic health record dataset. Sci Data. 2023;10(1):1. doi: 10.1038/s41597-022-01899-x [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22.Jiang D, Bian T, Shen Y, Huang Z. Association between admission systemic immune-inflammation index and mortality in critically ill patients with sepsis: a retrospective cohort study based on MIMIC-IV database. Clin Exp Med. 2023;23(7):3641–50. doi: 10.1007/s10238-023-01029-w [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 23.Dziura JD, Post LA, Zhao Q, Fu Z, Peduzzi P. Strategies for dealing with missing data in clinical trials: from design to analysis. Yale J Biol Med. 2013;86(3):343–58. [PMC free article] [PubMed] [Google Scholar]
  • 24.Stekhoven DJ, Bühlmann P. MissForest--non-parametric missing value imputation for mixed-type data. Bioinformatics. 2012;28(1):112–8. doi: 10.1093/bioinformatics/btr597 [DOI] [PubMed] [Google Scholar]
  • 25.Degenhardt F, Seifert S, Szymczak S. Evaluation of variable selection methods for random forests and omics data sets. Brief Bioinform. 2019;20(2):492–503. doi: 10.1093/bib/bbx124 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 26.Chen T, Guestrin C. XGBoost: A Scalable Tree Boosting System. In: Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining; 2016. 785–94. doi: 10.1145/2939672.2939785 [DOI] [Google Scholar]
  • 27.Lundberg SM, Erion G, Chen H, DeGrave A, Prutkin JM, Nair B, et al. From local explanations to global understanding with explainable AI for trees. Nat Mach Intell. 2020;2(1):56–67. doi: 10.1038/s42256-019-0138-9 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 28.Nohara Y, Matsumoto K, Soejima H, Nakashima N. Explanation of machine learning models using shapley additive explanation and application for real data in hospital. Comput Methods Programs Biomed. 2022;214:106584. doi: 10.1016/j.cmpb.2021.106584 [DOI] [PubMed] [Google Scholar]
  • 29.Lou J, Xiang Z, Zhu X, Fan Y, Song J, Cui S, et al. A retrospective study utilized MIMIC-IV database to explore the potential association between triglyceride-glucose index and mortality in critically ill patients with sepsis. Sci Rep. 2024;14(1):24081. doi: 10.1038/s41598-024-75050-8 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 30.Zheng R, Qian S, Shi Y, Lou C, Xu H, Pan J. Association between triglyceride-glucose index and in-hospital mortality in critically ill patients with sepsis: analysis of the MIMIC-IV database. Cardiovasc Diabetol. 2023;22(1):307. doi: 10.1186/s12933-023-02041-w [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 31.Lou J, Xiang Z, Zhu X, Song J, Cui S, Li J, et al. The non-linear association between creatinine-to-albumin ratio and medium-term mortality in patients with sepsis accompanied by acute kidney injury in the intensive care unit: a retrospective study based on the MIMIC database and external validation. Front Cell Infect Microbiol. 2025;15:1602921. doi: 10.3389/fcimb.2025.1602921 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 32.Saito T, Rehmsmeier M. The precision-recall plot is more informative than the ROC plot when evaluating binary classifiers on imbalanced datasets. PLoS One. 2015;10(3):e0118432. doi: 10.1371/journal.pone.0118432 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 33.Akilli NB, Yortanlı M, Mutlu H, Günaydın YK, Koylu R, Akca HS, et al. Prognostic importance of neutrophil-lymphocyte ratio in critically ill patients: short- and long-term outcomes. Am J Emerg Med. 2014;32(12):1476–80. doi: 10.1016/j.ajem.2014.09.001 [DOI] [PubMed] [Google Scholar]
  • 34.Chebl RB, Assaf M, Kattouf N, Haidar S, Khamis M, Abdeldaem K, et al. The association between the neutrophil to lymphocyte ratio and in-hospital mortality among sepsis patients: a prospective study. Medicine (Baltimore). 2022;101(30):e29343. doi: 10.1097/MD.0000000000029343 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 35.Hwang SY, Shin TG, Jo IJ, Jeon K, Suh GY, Lee TR, et al. Neutrophil-to-lymphocyte ratio as a prognostic marker in critically-ill septic patients. Am J Emerg Med. 2017;35(2):234–9. doi: 10.1016/j.ajem.2016.10.055 [DOI] [PubMed] [Google Scholar]
  • 36.Li J-Y, Yao R-Q, Liu S-Q, Zhang Y-F, Yao Y-M, Tian Y-P. Efficiency of monocyte/high-density lipoprotein cholesterol ratio combined with neutrophil/lymphocyte ratio in predicting 28-day mortality in patients with sepsis. Front Med (Lausanne). 2021;8:741015. doi: 10.3389/fmed.2021.741015 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 37.Liu S, Li Y, She F, Zhao X, Yao Y. Predictive value of immune cell counts and neutrophil-to-lymphocyte ratio for 28-day mortality in patients with sepsis caused by intra-abdominal infection. Burns Trauma. 2021;9:tkaa040. doi: 10.1093/burnst/tkaa040 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 38.Liu S, Wang X, She F, Zhang W, Liu H, Zhao X. Effects of neutrophil-to-lymphocyte ratio combined with interleukin-6 in predicting 28-day mortality in patients with sepsis. Front Immunol. 2021;12:639735. doi: 10.3389/fimmu.2021.639735 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 39.Wei W, Huang X, Yang L, Li J, Liu C, Pu Y, et al. Neutrophil-to-lymphocyte ratio as a prognostic marker of mortality and disease severity in septic acute kidney injury patients: a retrospective study. Int Immunopharmacol. 2023;116:109778. doi: 10.1016/j.intimp.2023.109778 [DOI] [PubMed] [Google Scholar]
  • 40.Ye W, Chen X, Huang Y, Li Y, Xu Y, Liang Z, et al. The association between neutrophil-to-lymphocyte count ratio and mortality in septic patients: a retrospective analysis of the MIMIC-III database. J Thorac Dis. 2020;12(5):1843–55. doi: 10.21037/jtd-20-1169 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 41.Denstaedt SJ, Cano J, Wang XQ, Donnelly JP, Seelye S, Prescott HC. Blood count derangements after sepsis and association with post-hospital outcomes. Front Immunol. 2023;14:1133351. doi: 10.3389/fimmu.2023.1133351 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 42.Di Rosa M, Sabbatinelli J, Soraci L, Corsonello A, Bonfigli AR, Cherubini A, et al. Neutrophil-to-lymphocyte ratio (NLR) predicts mortality in hospitalized geriatric patients independent of the admission diagnosis: a multicenter prospective cohort study. J Transl Med. 2023;21(1):835. doi: 10.1186/s12967-023-04717-z [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 43.Chiew CJ, Liu N, Wong TH, Sim YE, Abdullah HR. Utilizing machine learning methods for preoperative prediction of postsurgical mortality and intensive care unit admission. Ann Surg. 2020;272(6):1133–9. doi: 10.1097/SLA.0000000000003297 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 44.Keilwagen J, Grosse I, Grau J. Area under precision-recall curves for weighted and unweighted data. PLoS One. 2014;9(3):e92209. doi: 10.1371/journal.pone.0092209 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 45.Salciccioli JD, Marshall DC, Pimentel MAF, Santos MD, Pollard T, Celi LA, et al. The association between the neutrophil-to-lymphocyte ratio and mortality in critical illness: an observational cohort study. Crit Care. 2015;19(1):13. doi: 10.1186/s13054-014-0731-6 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 46.Lou J, Kong H, Xiang Z, Zhu X, Cui S, Li J, et al. The J-shaped association between the ratio of neutrophil counts to prognostic nutritional index and mortality in ICU patients with sepsis: a retrospective study based on the MIMIC database. Front Cell Infect Microbiol. 2025;15:1603104. doi: 10.3389/fcimb.2025.1603104 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 47.Cai J, Jin Y, Lou J, Xu B, Qi H, Li J. The V-shaped association between the ratio of neutrophil counts to prognostic nutritional index and 30-, 60-, and 90-day mortality in elderly critically ill patients aged 65 and older with sepsis: a retrospective study based on the MIMIC database. Front Nutr. 2025;12:1602016. doi: 10.3389/fnut.2025.1602016 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 48.Karakonstantis S, Kalemaki D, Tzagkarakis E, Lydakis C. Pitfalls in studies of eosinopenia and neutrophil-to-lymphocyte count ratio. Infect Dis (Lond). 2018;50(3):163–74. doi: 10.1080/23744235.2017.1388537 [DOI] [PubMed] [Google Scholar]

Decision Letter 0

Siddharth Gosavi

9 Feb 2026

-->PONE-D-25-63990-->-->The Prognostic Value of the Early Neutrophil-to-Lymphocyte Ratio for 28-Day Mortality in Sepsis Patients: A Machine Learning-Based Investigation of the MIMIC Database-->-->PLOS One

Dear Dr. Lai,

Thank you for submitting your manuscript to PLOS ONE. After careful consideration, we feel that it has merit but does not fully meet PLOS ONE’s publication criteria as it currently stands. Therefore, we invite you to submit a revised version of the manuscript that addresses the points raised during the review process.

Please submit your revised manuscript by Mar 26 2026 11:59PM. If you will need more time than this to complete your revisions, please reply to this message or contact the journal office at plosone@plos.org. When you're ready to submit your revision, log on to https://www.editorialmanager.com/pone/ and select the 'Submissions Needing Revision' folder to locate your manuscript file.

Please include the following items when submitting your revised manuscript:-->

  • A letter that responds to each point raised by the academic editor and reviewer(s). You should upload this letter as a separate file labeled 'Response to Reviewers'.

  • A marked-up copy of your manuscript that highlights changes made to the original version. You should upload this as a separate file labeled 'Revised Manuscript with Track Changes'.

  • An unmarked version of your revised paper without tracked changes. You should upload this as a separate file labeled 'Manuscript'.

-->If you would like to make changes to your financial disclosure, please include your updated statement in your cover letter. Guidelines for resubmitting your figure files are available below the reviewer comments at the end of this letter.

If applicable, we recommend that you deposit your laboratory protocols in protocols.io to enhance the reproducibility of your results. Protocols.io assigns your protocol its own identifier (DOI) so that it can be cited independently in the future. For instructions see: https://journals.plos.org/plosone/s/submission-guidelines#loc-laboratory-protocols. Additionally, PLOS ONE offers an option for publishing peer-reviewed Lab Protocol articles, which describe protocols hosted on protocols.io. Read more information on sharing protocols at https://plos.org/protocols?utm_medium=editorial-email&utm_source=authorletters&utm_campaign=protocols.

We look forward to receiving your revised manuscript.

Kind regards,

Siddharth Gosavi, MBBS, MD Internal Medicine,DNB Internal Medicine

Academic Editor

PLOS One

Journal Requirements:

When submitting your revision, we need you to address these additional requirements.

1.Please ensure that your manuscript meets PLOS ONE's style requirements, including those for file naming. The PLOS ONE style templates can be found at

https://journals.plos.org/plosone/s/file?id=wjVg/PLOSOne_formatting_sample_main_body.pdf and

https://journals.plos.org/plosone/s/file?id=ba62/PLOSOne_formatting_sample_title_authors_affiliations.pdf

2. Please include your tables as part of your main manuscript and remove the individual files. Please note that supplementary tables (should remain/ be uploaded) as separate "supporting information" files.

3. Please note that funding information should not appear in any section or other areas of your manuscript. We will only publish funding information present in the Funding Statement section of the online submission form. Please remove any funding-related text from the manuscript.

4. Thank you for stating the following financial disclosure:

“This project was supported by the grants of High-level Medical Team Project in Baoan (No.202405) and Medical and Health Scientific Research Project of Shenzhen Bao 'an District in 2023 (No.2023JD118).”

Please state what role the funders took in the study.  If the funders had no role, please state: "The funders had no role in study design, data collection and analysis, decision to publish, or preparation of the manuscript."

If this statement is not correct you must amend it as needed.

Please include this amended Role of Funder statement in your cover letter; we will change the online submission form on your behalf.

5. In the online submission form you indicate that your data is not available for proprietary reasons and have provided a contact point for accessing this data. Please note that your current contact point is a co-author on this manuscript. According to our Data Policy, the contact point must not be an author on the manuscript and must be an institutional contact, ideally not an individual. Please revise your data statement to a non-author institutional point of contact, such as a data access or ethics committee, and send this to us via return email. Please also include contact information for the third party organization, and please include the full citation of where the data can be found.

6. Your ethics statement should only appear in the Methods section of your manuscript. If your ethics statement is written in any section besides the Methods, please move it to the Methods section and delete it from any other section. Please ensure that your ethics statement is included in your manuscript, as the ethics statement entered into the online submission form will not be published alongside your manuscript.

7. Please include a separate caption for each figure in your manuscript.

8. Please include captions for your Supporting Information files at the end of your manuscript, and update any in-text citations to match accordingly. Please see our Supporting Information guidelines for more information: http://journals.plos.org/plosone/s/supporting-information.

9. If the reviewer comments include a recommendation to cite specific previously published works, please review and evaluate these publications to determine whether they are relevant and should be cited. There is no requirement to cite these works unless the editor has indicated otherwise.

10. Please review your reference list to ensure that it is complete and correct. If you have cited papers that have been retracted, please include the rationale for doing so in the manuscript text, or remove these references and replace them with relevant current references. Any changes to the reference list should be mentioned in the rebuttal letter that accompanies your revised manuscript. If you need to cite a retracted article, indicate the article’s retracted status in the References list and also include a citation and full reference for the retraction notice.

Additional Editor Comments:

This is a well-conducted retrospective study that leverages a large public database and advanced machine learning techniques to investigate the prognostic value of the neutrophil-to-lymphocyte ratio (NLR) in sepsis. The manuscript is generally well-written, and the methodological approach, including the two-stage feature selection process using Boruta and XGBoost, is a strength. However, several important issues need to be addressed to strengthen the manuscript's claims and contextualize its findings within the existing literature. Firstly, the novelty statement should be moderated. While the combination of methods is robust, claiming this as the "first large-scale retrospective study employing machine learning-based feature selection to evaluate the ability of the NLR" may be too absolute. Recent research has increasingly applied ML techniques to similar prognostic questions in sepsis using the MIMIC database. For instance, a 2024 study in Scientific Reports used the MIMIC-IV database to explore the triglyceride-glucose (TyG) index, another accessible biomarker, and its association with mortality in septic patients, employing similar analytical techniques including restricted cubic splines and subgroup analysis (DOI: 10.1038/s41598-024-75050-8). Reframing the contribution to highlight the specific, rigorous feature selection pipeline and the detailed exploration of NLR's biphasic predictive pattern across clinical subgroups would provide a more accurate and compelling narrative.

Several methodological aspects require further clarification. The decision to use only the first NLR value within six hours of ICU admission is pragmatic but represents a significant limitation, as a single static measurement may not capture the dynamic inflammatory state critical in sepsis. The discussion should more thoroughly acknowledge that this "early" snapshot might miss peak values or trends that carry greater prognostic weight. Furthermore, while the use of the missForest package for imputation is appropriate, the manuscript would benefit from a brief validation statement or supplementary table comparing the distribution of key variables before and after imputation to assure readers of the integrity of the imputed dataset. The model performance is commendable, but the discussion would be enriched by a more direct comparison with other contemporary biomarker-based models developed on similar cohorts. For example, studies have examined composite indices like the creatinine-to-albumin ratio (CAR) for sepsis-associated acute kidney injury, finding non-linear associations with mortality, which echoes your findings on NLR's biphasic pattern (DOI: 10.3389/fcimb.2025.1602921). Situating your XGBoost model's performance relative to these emerging indicators would better define its relative clinical utility.

The most substantial limitation is the lack of external validation, which is rightly mentioned but deserves greater emphasis in the discussion. The reliance solely on the MIMIC database, while common, introduces risks related to selection bias and protocol homogeneity from a single healthcare system. The conclusions about NLR's predictive value, particularly its diminished utility in severe organ failure, must be framed as hypothesis-generating and require validation in independent, prospective cohorts. Additionally, please ensure all figures and tables are correctly numbered and embedded in the main text for review, and standardize the formatting of abbreviations throughout (e.g., SAPS II vs. SAPSII). In the results, when stating survivors had better vital signs, please specify which parameters showed significant differences. The statistical methods section should explicitly state which p-values were adjusted using the Benjamini-Hochberg procedure.

The discussion effectively interprets the biphasic nature of the NLR's predictive value. To further strengthen it, consider integrating the concept that the predictive weight of inflammatory markers like NLR may be context-dependent, overshadowed by organ dysfunction in advanced disease. This resonates with research on other composite markers that combine inflammatory and nutritional/renal dimensions, such as the ratio of neutrophil counts to the prognostic nutritional index (NPNR), which has shown J- or V-shaped associations with mortality in septic patients, highlighting the complex interplay between inflammation and host reserve (DOI: 10.3389/fcimb.2025.1603104; DOI: 10.3389/fnut.2025.1602016). Citing such work would help build a more nuanced theoretical framework for why NLR's utility plateaus in critical illness. Finally, the reference list should be formatted consistently according to the journal's style, and including some of the more recent, relevant studies suggested here would update the scholarly context.

In conclusion, this study provides valuable insights into the conditional prognostic value of the NLR in sepsis. By addressing the points above—particularly tempering the novelty claim, elaborating on methodological limitations, and engaging with contemporary literature on sepsis biomarkers—the manuscript can be significantly enhanced. I recommend a minor revision to allow the authors to refine these aspects and produce a more robust and contextualized final paper.

[Note: HTML markup is below. Please do not edit.]

Reviewers' comments:

Reviewer's Responses to Questions

-->Comments to the Author

1. Is the manuscript technically sound, and do the data support the conclusions?

The manuscript must describe a technically sound piece of scientific research with data that supports the conclusions. Experiments must have been conducted rigorously, with appropriate controls, replication, and sample sizes. The conclusions must be drawn appropriately based on the data presented. -->

Reviewer #1: Partly

**********

-->2. Has the statistical analysis been performed appropriately and rigorously? -->

Reviewer #1: Yes

**********

-->3. Have the authors made all data underlying the findings in their manuscript fully available?

The PLOS Data policy requires authors to make all data underlying the findings described in their manuscript fully available without restriction, with rare exception (please refer to the Data Availability Statement in the manuscript PDF file). The data should be provided as part of the manuscript or its supporting information, or deposited to a public repository. For example, in addition to summary statistics, the data points behind means, medians and variance measures should be available. If there are restrictions on publicly sharing data—e.g. participant privacy or use of data from a third party—those must be specified.-->

Reviewer #1: Yes

**********

-->4. Is the manuscript presented in an intelligible fashion and written in standard English?

PLOS ONE does not copyedit accepted manuscripts, so the language in submitted articles must be clear, correct, and unambiguous. Any typographical or grammatical errors should be corrected at revision, so please note any specific errors here.-->

Reviewer #1: Yes

**********

-->5. Review Comments to the Author

Please use the space provided to explain your answers to the questions above. You may also include additional comments for the author, including concerns about dual publication, research ethics, or publication ethics. (Please upload your review as an attachment if it exceeds 20,000 characters)-->

Reviewer #1: This is a well-conducted retrospective study that leverages a large public database and advanced machine learning techniques to investigate the prognostic value of the neutrophil-to-lymphocyte ratio (NLR) in sepsis. The manuscript is generally well-written, and the methodological approach, including the two-stage feature selection process using Boruta and XGBoost, is a strength. However, several important issues need to be addressed to strengthen the manuscript's claims and contextualize its findings within the existing literature. Firstly, the novelty statement should be moderated. While the combination of methods is robust, claiming this as the "first large-scale retrospective study employing machine learning-based feature selection to evaluate the ability of the NLR" may be too absolute. Recent research has increasingly applied ML techniques to similar prognostic questions in sepsis using the MIMIC database. For instance, a 2024 study in Scientific Reports used the MIMIC-IV database to explore the triglyceride-glucose (TyG) index, another accessible biomarker, and its association with mortality in septic patients, employing similar analytical techniques including restricted cubic splines and subgroup analysis (DOI: 10.1038/s41598-024-75050-8). Reframing the contribution to highlight the specific, rigorous feature selection pipeline and the detailed exploration of NLR's biphasic predictive pattern across clinical subgroups would provide a more accurate and compelling narrative.

Several methodological aspects require further clarification. The decision to use only the first NLR value within six hours of ICU admission is pragmatic but represents a significant limitation, as a single static measurement may not capture the dynamic inflammatory state critical in sepsis. The discussion should more thoroughly acknowledge that this "early" snapshot might miss peak values or trends that carry greater prognostic weight. Furthermore, while the use of the missForest package for imputation is appropriate, the manuscript would benefit from a brief validation statement or supplementary table comparing the distribution of key variables before and after imputation to assure readers of the integrity of the imputed dataset. The model performance is commendable, but the discussion would be enriched by a more direct comparison with other contemporary biomarker-based models developed on similar cohorts. For example, studies have examined composite indices like the creatinine-to-albumin ratio (CAR) for sepsis-associated acute kidney injury, finding non-linear associations with mortality, which echoes your findings on NLR's biphasic pattern (DOI: 10.3389/fcimb.2025.1602921). Situating your XGBoost model's performance relative to these emerging indicators would better define its relative clinical utility.

The most substantial limitation is the lack of external validation, which is rightly mentioned but deserves greater emphasis in the discussion. The reliance solely on the MIMIC database, while common, introduces risks related to selection bias and protocol homogeneity from a single healthcare system. The conclusions about NLR's predictive value, particularly its diminished utility in severe organ failure, must be framed as hypothesis-generating and require validation in independent, prospective cohorts. Additionally, please ensure all figures and tables are correctly numbered and embedded in the main text for review, and standardize the formatting of abbreviations throughout (e.g., SAPS II vs. SAPSII). In the results, when stating survivors had better vital signs, please specify which parameters showed significant differences. The statistical methods section should explicitly state which p-values were adjusted using the Benjamini-Hochberg procedure.

The discussion effectively interprets the biphasic nature of the NLR's predictive value. To further strengthen it, consider integrating the concept that the predictive weight of inflammatory markers like NLR may be context-dependent, overshadowed by organ dysfunction in advanced disease. This resonates with research on other composite markers that combine inflammatory and nutritional/renal dimensions, such as the ratio of neutrophil counts to the prognostic nutritional index (NPNR), which has shown J- or V-shaped associations with mortality in septic patients, highlighting the complex interplay between inflammation and host reserve (DOI: 10.3389/fcimb.2025.1603104; DOI: 10.3389/fnut.2025.1602016). Citing such work would help build a more nuanced theoretical framework for why NLR's utility plateaus in critical illness. Finally, the reference list should be formatted consistently according to the journal's style, and including some of the more recent, relevant studies suggested here would update the scholarly context.

In conclusion, this study provides valuable insights into the conditional prognostic value of the NLR in sepsis. By addressing the points above—particularly tempering the novelty claim, elaborating on methodological limitations, and engaging with contemporary literature on sepsis biomarkers—the manuscript can be significantly enhanced. I recommend a minor revision to allow the authors to refine these aspects and produce a more robust and contextualized final paper.

**********

-->6. PLOS authors have the option to publish the peer review history of their article (what does this mean?). If published, this will include your full peer review and any attached files.

If you choose “no”, your identity will remain anonymous but your review may still be made public.

Do you want your identity to be public for this peer review? For information about this choice, including consent withdrawal, please see our Privacy Policy.-->

Reviewer #1: No

**********

[NOTE: If reviewer comments were submitted as an attachment file, they will be attached to this email and accessible via the submission site. Please log into your account, locate the manuscript record, and check for the action link "View Attachments". If this link does not appear, there are no attachment files.]

To ensure your figures meet our technical requirements, please review our figure guidelines: https://journals.plos.org/plosone/s/figures

You may also use PLOS’s free figure tool, NAAS, to help you prepare publication quality figures: https://journals.plos.org/plosone/s/figures#loc-tools-for-figure-preparation.

NAAS will assess whether your figures meet our technical requirements by comparing each figure against our figure specifications.

-->

PLoS One. 2026 Jun 2;21(6):e0348676. doi: 10.1371/journal.pone.0348676.r002

Author response to Decision Letter 1


18 Feb 2026

Response to Reviewer 1-Comment 1

Reviewer’s Comment:

*This is a well-conducted retrospective study that leverages a large public database and advanced machine learning techniques to investigate the prognostic value of the neutrophil-to-lymphocyte ratio (NLR) in sepsis. The manuscript is generally well-written, and the methodological approach, including the two-stage feature selection process using Boruta and XGBoost, is a strength. However, several important issues need to be addressed to strengthen the manuscript's claims and contextualize its findings within the existing literature. Firstly, the novelty statement should be moderated. While the combination of methods is robust, claiming this as the "first large-scale retrospective study employing machine learning-based feature selection to evaluate the ability of the NLR" may be too absolute. Recent research has increasingly applied ML techniques to similar prognostic questions in sepsis using the MIMIC database. For instance, a 2024 study in Scientific Reports used the MIMIC-IV database to explore the triglyceride-glucose (TyG) index, another accessible biomarker, and its association with mortality in septic patients, employing similar analytical techniques including restricted cubic splines and subgroup analysis (DOI: 10.1038/s41598-024-75050-8). Reframing the contribution to highlight the specific, rigorous feature selection pipeline and the detailed exploration of NLR's biphasic predictive pattern across clinical subgroups would provide a more accurate and compelling narrative.*

________________________________________

Author’s Response:

We sincerely thank the reviewer for the thoughtful and constructive feedback. We fully agree that the original wording of our novelty statement was overly absolute and insufficiently contextualized within the growing body of machine learning (ML) research using the MIMIC database. We are grateful for the reviewer’s suggestion to cite the relevant 2024 Scientific Reports study on the TyG index (DOI: 10.1038/s41598-024-75050-8), which indeed represents a methodologically similar and high-quality contribution. We have now revised the opening paragraph of the Discussion section to:

1.Remove all absolute claims (e.g., “To our knowledge, this is the first…”) and instead position our work within the broader context of recent ML-based prognostic studies in sepsis.

2.Explicitly cite and acknowledge the TyG index study, using it as an example of how ML is increasingly employed to reevaluate accessible biomarkers-thereby clarifying that our contribution lies not in being the first to use such techniques, but in applying them specifically to the NLR with a particularly rigorous feature selection protocol and a novel focus on clinical phenotype stratification.

3.Reframe the core contribution around two specific, defensible innovations:

a) Methodological: a robust two-stage (Boruta + XGBoost) feature selection pipeline combined with SHAP-based marginal contribution analysis.

b) Clinical: the first detailed description of a biphasic, phenotype-specific predictive pattern of the NLR, which distinguishes inflammation-driven mortality risk (evident in patients with preserved organ function) from organ failure-driven mortality (where NLR loses its predictive utility).

________________________________________

We believe this revised framing more accurately reflects the study’s genuine contributions and appropriately acknowledges prior work in the field. We are grateful to the reviewer for guiding us toward a more precise and compelling narrative.

Response to Reviewer 1-Comment 2

Reviewer’s Comment:

Several methodological aspects require further clarification. The decision to use only the first NLR value within six hours of ICU admission is pragmatic but represents a significant limitation, as a single static measurement may not capture the dynamic inflammatory state critical in sepsis. The discussion should more thoroughly acknowledge that this "early" snapshot might miss peak values or trends that carry greater prognostic weight. Furthermore, while the use of the missForest package for imputation is appropriate, the manuscript would benefit from a brief validation statement or supplementary table comparing the distribution of key variables before and after imputation to assure readers of the integrity of the imputed dataset. The model performance is commendable, but the discussion would be enriched by a more direct comparison with other contemporary biomarker-based models developed on similar cohorts. For example, studies have examined composite indices like the creatinine-to-albumin ratio (CAR) for sepsis-associated acute kidney injury, finding non-linear associations with mortality, which echoes your findings on NLR's biphasic pattern (DOI: 10.3389/fcimb.2025.1602921). Situating your XGBoost model's performance relative to these emerging indicators would better define its relative clinical utility.

________________________________________

Author’s Response:

We sincerely thank the reviewer for these thoughtful and constructive methodological suggestions. We have carefully addressed each of the three specific concerns: (1) limitation of single NLR measurement, (2) validation of missForest imputation, and (3) contextualization against other emerging biomarker-based models, as detailed below. All corresponding revisions have been clearly marked in the tracked-change version of the manuscript.

1. Limitation of a Single Early NLR Measurement

Reviewer’s concern:

The use of only the first NLR value within six hours of ICU admission is a pragmatic but potentially significant limitation, as a static measurement may not reflect the dynamic inflammatory trajectory of sepsis.

Our response:

We fully agree with the reviewer. Although we had briefly acknowledged this limitation in the original manuscript, we have now substantially expanded this discussion in the Limitations section. Specifically, we now explicitly state that a single baseline measurement captures only a static snapshot of the inflammatory state and may underestimate the prognostic value of NLR trajectories or peak values, which could better reflect the evolving host response. We also explicitly call for future studies incorporating serial NLR measurements or longitudinal trajectory modeling to further validate and refine our findings. This revision directly aligns with our overall conclusion that the clinical utility of NLR lies in dynamic monitoring rather than isolated static assessment.

Action taken:

a) Manuscript location: Discussion – Limitations (please see track changed)

2.Validation of missForest Imputation

Reviewer’s concern:

While the use of missForest for imputation is appropriate, the manuscript would benefit from a brief validation statement or supplementary table comparing the distribution of key variables before and after imputation to assure readers of the integrity of the imputed dataset.

Our response:

We appreciate this suggestion for enhanced transparency. In response, we performed a systematic comparison of the distributions of key continuous and categorical variables before and after imputation using the missForest algorithm. The results are now presented in S2 Table.

a) For variables with low missingness (<5%), the distributions remained virtually identical.

b) For variables with higher missing rates (approximately 30%), such as SpO₂, MAP, heart rate, and blood gas parameters, we observed a modest convergence in dispersion (SD, IQR)—an expected behavior of model based imputation when recovering missing values from observed patterns.

c) Critically, all post imputation means, medians, and quartiles remained within clinically plausible and physiologically coherent ranges, and no systematic shift would alter the clinical interpretation of these parameters (e.g., median SpO₂: 98.0% → 98.4%; median MAP: 81.0 mmHg → 79.3 mmHg).

We have added a brief validation statement in the Methods section (Data extraction) and direct readers to S2 Table for full details.

Action taken:

a) Manuscript location: Methods – Data extraction (please see track changed).

b) Supplementary material: Added S2 Table (“Comparison of Key Variables Before and After Imputation Using the MissForest Algorithm”).

3. Comparison with Other Emerging Biomarker-Based Models

Reviewer’s concern:

The discussion would be enriched by a more direct comparison with other contemporary biomarker-based models developed on similar cohorts—particularly studies examining the creatinine-to-albumin ratio (CAR) in sepsis-associated acute kidney injury, which reported non linear associations with mortality echoing our biphasic NLR pattern. Situating our XGBoost model’s performance relative to these emerging indicators would better define its relative clinical utility.

Our response:

We thank the reviewer for guiding us to contextualize our findings within the broader landscape of novel prognostic biomarkers. We have now substantially expanded the Discussion to include a dedicated paragraph that directly compares:

• Study populations and sample sizes (our N = 4,376 vs. 1,257–2,712 in comparator studies);

• Model performance (our XGBoost ROC AUC = 0.875, PR AUC = 0.603 vs. CAR based AUCs of 0.68-0.75; TyG studies did not report AUC but reported hazard ratios of 1.4-1.8);

• Non linear patterns (our biphasic NLR mortality curve with inflection points at NLR = 7.3 and 27.3 vs. the J shaped CAR mortality curve with inflection at CAR = 1.2 mg/dL, and the non linear TyG mortality curve reported by Lou et al. [29] with inflection at TyG = 8.9);

• Practical advantages of NLR: zero additional cost, routine availability from complete blood count, and direct biological linkage to innate adaptive immune imbalance-a pathophysiological axis distinct from insulin resistance (TyG) or renal inflammatory crosstalk (CAR).

We emphasize that these biomarkers are not mutually exclusive but rather capture complementary aspects of sepsis pathophysiology, and their combined use may enable more refined phenotype specific risk stratification. Our model’s superior discriminative performance underscores the value of integrating a simple, inexpensive marker like NLR within a robust machine learning framework.

Action taken:

• Manuscript location: Discussion – second paragraph (please see track changed).

________________________________________

We are confident that these revisions have substantially improved the methodological transparency, scholarly depth, and clinical contextualization of our manuscript. We thank the reviewer again for these insightful suggestions, which we believe have greatly strengthened our work.

Response to Reviewer 1-Comment 3

Reviewer’s Comment:

The most substantial limitation is the lack of external validation, which is rightly mentioned but deserves greater emphasis in the discussion. The reliance solely on the MIMIC database, while common, introduces risks related to selection bias and protocol homogeneity from a single healthcare system. The conclusions about NLR's predictive value, particularly its diminished utility in severe organ failure, must be framed as hypothesis-generating and require validation in independent, prospective cohorts. Additionally, please ensure all figures and tables are correctly numbered and embedded in the main text for review, and standardize the formatting of abbreviations throughout (e.g., SAPS II vs. SAPSII). In the results, when stating survivors had better vital signs, please specify which parameters showed significant differences. The statistical methods section should explicitly state which p-values were adjusted using the Benjamini-Hochberg procedure.

________________________________________

Author’s Response:

We thank the reviewer for these thoughtful and constructive comments. We have carefully addressed each point as follows:

1. External validation and hypothesis generating framing.

We fully agree that the absence of external validation is the most substantial limitation of our study. While we had acknowledged this in the original manuscript, the reviewer correctly notes that the emphasis was insufficient. In response, we have substantially strengthened the Limitations section to explicitly state that:

• Our single center MIMIC derived findings are subject to selection bias and protocol homogeneity;

• The proposed biphasic, phenotype dependent predictive pattern of the NLR must be regarded as hypothesis generating, not definitive;

• Independent validation in prospective, multi center cohorts is urgently required before any clinical implementation can be considered;

• Future external validation should specifically test the generalizability of the identified inflection points (NLR = 7.3 and 27.3) and the differential predictive utility across SOFA strata observed in our cohort.

We have also revised the Conclusions section to echo this cautious framing, explicitly stating that our conclusions remain hypothesis generating pending external validation. All changes are clearly marked in the tracked change manuscript.

2. Figures and tables numbering/embedding.

We have verified that all figures and tables are correctly numbered and that each is explicitly cited in the main text.

3. Abbreviation standardization.

We have standardized all abbreviations throughout the manuscript. Specifically, “SAPSII” has been uniformly corrected to “SAPS II” in all instances. Other abbreviations (SOFA, CRRT, VIS, etc.) were already consistent and remain unchanged.

4. Specification of vital signs in Results.

We have revised the Results – Baseline characteristics section to explicitly list which vital signs showed significant differences and which did not.

5. Clarification of Benjamini Hochberg adjustment.

We have now explicitly clarified the application of the Benjamini–Hochberg procedure in the Statistical analysis section. Specifically, we state that:

a) For multiple comparisons across categorical covariates in baseline tables (e.g., comorbidities, organ dysfunction components, treatment categories), p values were adjusted using the Benjamini–Hochberg method to control the false discovery rate.

b) For interaction p values derived from subgroup analyses, the same FDR correction was applied to comprehensively evaluate potential effect modifiers.

Response to Reviewer 1-Comment 4

Reviewer’s Comment:

The discussion effectively interprets the biphasic nature of the NLR's predictive value. To further strengthen it, consider integrating the concept that the predictive weight of inflammatory markers like NLR may be context-dependent, overshadowed by organ dysfunction in advanced disease. This resonates with research on other composite markers that combine inflammatory and nutritional/renal dimensions, such as the ratio of neutrophil counts to the prognostic nutritional index (NPNR), which has shown J- or V-shaped associations with mortality in septic patients, highlighting the complex interplay between inflammation and host reserve (DOI: 10.3389/fcimb.2025.1603104; DOI: 10.3389/fnut.2025.1602016). Citing such work would help build a more nuanced theoretical framework for why NLR's utility plateaus in critical illness. Finally, the reference list should be formatted consistently according to the journal's style, and including some of the more recent, relevant studies suggested here would update the scholarly context.

________________________________________

Author’s Response:

We thank the reviewer for this thoughtful and constructive suggestion. We fully agree that the context dependent nature of inflammatory biomarkers is a critical concept that deserves explicit articulation within our theoretical framework. In response, we have:

1.Expanded the Discussion to incorporate the NPNR analogy.

We have added a new paragraph in the section of Discussion that explicitly links our findings on the NLR’s biphasic pattern to recent research on the neutrophil count to prognostic nutritional index ratio (NPNR). Both NLR and NPNR have been shown to exhibit J or V shaped associations with mortality, suggesting a shared underlying principle: inflammatory markers are most informative when the host still possesses physiological reserve, and their predictive utility wanes once organ failure becomes the dominant

Attachment

Submitted filename: Response to Reviewers.docx

pone.0348676.s013.docx (29.3KB, docx)

Decision Letter 1

Chiara Lazzeri

20 Apr 2026

The Prognostic Value of the Early Neutrophil-to-Lymphocyte Ratio for 28-Day Mortality in Sepsis Patients: A Machine Learning-Based Investigation of the MIMIC Database

PONE-D-25-63990R1

Dear Dr. Lai,

We’re pleased to inform you that your manuscript has been judged scientifically suitable for publication and will be formally accepted for publication once it meets all outstanding technical requirements.

Within one week, you’ll receive an e-mail detailing the required amendments. When these have been addressed, you’ll receive a formal acceptance letter and your manuscript will be scheduled for publication.

An invoice will be generated when your article is formally accepted. Please note, if your institution has a publishing partnership with PLOS and your article meets the relevant criteria, all or part of your publication costs will be covered. Please make sure your user information is up-to-date by logging into Editorial Manager at Editorial Manager® and clicking the ‘Update My Information' link at the top of the page. For questions related to billing, please contact billing support.

If your institution or institutions have a press office, please notify them about your upcoming paper to help maximize its impact. If they’ll be preparing press materials, please inform our press team as soon as possible -- no later than 48 hours after receiving the formal acceptance. Your manuscript will remain under strict press embargo until 2 pm Eastern Time on the date of publication. For more information, please contact onepress@plos.org.

Kind regards,

Chiara Lazzeri

Academic Editor

PLOS One

Additional Editor Comments (optional):

Reviewers' comments:

Reviewer's Responses to Questions

-->Comments to the Author

1. If the authors have adequately addressed your comments raised in a previous round of review and you feel that this manuscript is now acceptable for publication, you may indicate that here to bypass the “Comments to the Author” section, enter your conflict of interest statement in the “Confidential to Editor” section, and submit your "Accept" recommendation.-->

Reviewer #1: All comments have been addressed

Reviewer #2: All comments have been addressed

**********

-->2. Is the manuscript technically sound, and do the data support the conclusions?

The manuscript must describe a technically sound piece of scientific research with data that supports the conclusions. Experiments must have been conducted rigorously, with appropriate controls, replication, and sample sizes. The conclusions must be drawn appropriately based on the data presented. -->

Reviewer #1: Yes

Reviewer #2: Partly

**********

-->3. Has the statistical analysis been performed appropriately and rigorously? -->

Reviewer #1: Yes

Reviewer #2: I Don't Know

**********

-->4. Have the authors made all data underlying the findings in their manuscript fully available?

The PLOS Data policy requires authors to make all data underlying the findings described in their manuscript fully available without restriction, with rare exception (please refer to the Data Availability Statement in the manuscript PDF file). The data should be provided as part of the manuscript or its supporting information, or deposited to a public repository. For example, in addition to summary statistics, the data points behind means, medians and variance measures should be available. If there are restrictions on publicly sharing data—e.g. participant privacy or use of data from a third party—those must be specified.-->

Reviewer #1: Yes

Reviewer #2: Yes

**********

-->5. Is the manuscript presented in an intelligible fashion and written in standard English?

PLOS ONE does not copyedit accepted manuscripts, so the language in submitted articles must be clear, correct, and unambiguous. Any typographical or grammatical errors should be corrected at revision, so please note any specific errors here.-->

Reviewer #1: Yes

Reviewer #2: Yes

**********

-->6. Review Comments to the Author

Please use the space provided to explain your answers to the questions above. You may also include additional comments for the author, including concerns about dual publication, research ethics, or publication ethics. (Please upload your review as an attachment if it exceeds 20,000 characters)-->

Reviewer #1: The authors have thoroughly and satisfactorily addressed all of the concerns raised in the previous review round. The revised manuscript is substantially improved in terms of methodological transparency, contextualization within the existing literature, and appropriate framing of conclusions. The study makes a valuable contribution to the field of sepsis prognostication by applying a rigorous machine learning pipeline to re-evaluate the prognostic role of the neutrophil-to-lymphocyte ratio (NLR).

Minor Suggestions:

Consider briefly mentioning in the Limitations that the optimal NLR cutoffs (7.3 and 27.3) were derived from a single dataset and may require calibration before clinical use in other settings.

The abbreviation "SAPS II" is correct, but in Table 2 and a few places, it appears as "SAPS II" – ensure consistency (looks fine; just a quick double-check).

In the Discussion, the phrase "infection point" appears twice (e.g., "first infection point"). It should be "inflection point". Please correct.

These do not affect the overall quality or validity of the work. The manuscript is well revised and suitable for publication. I recommend acceptance without further revision.

Reviewer #2: Overall Assessment: Minor Revision

Summary: This study uses MIMIC-IV database and machine learning (XGBoost) to evaluate the prognostic value of early NLR for 28-day mortality in sepsis patients (N=4,376). The authors found that NLR was independently associated with mortality (OR=1.16) and identified a biphasic predictive pattern where NLR's utility was pronounced in patients with SOFA≤4 but lost significance in SOFA≥9.

Strengths:

1.Large sample size with rigorous feature selection (Boruta + XGBoost)

2.Novel finding of biphasic, phenotype-specific predictive pattern

3.Thorough discussion of limitations

Weaknesses & Suggestions:

1.Only the first NLR measurement was used; dynamic changes may carry greater prognostic weight. Please discuss.

2.Tables 1 and 2 are overly dense. Consider moving less critical variables to supplementary materials.

3.In the abstract, PO2 value (221.07 ± 115.04 mmHg) appears unusually high for arterial blood. Please verify.

4.While the ML approach is robust, the clinical utility of NLR alone remains limited (SHAP rank #14). This should be more explicitly stated in the conclusion.

Recommendation: Minor revision.

**********

-->7. PLOS authors have the option to publish the peer review history of their article (what does this mean?). If published, this will include your full peer review and any attached files.

If you choose “no”, your identity will remain anonymous but your review may still be made public.

Do you want your identity to be public for this peer review? For information about this choice, including consent withdrawal, please see our Privacy Policy.-->

Reviewer #1: No

Reviewer #2: No

**********

Acceptance letter

Chiara Lazzeri

PONE-D-25-63990R1

PLOS One

Dear Dr. Lai,

I'm pleased to inform you that your manuscript has been deemed suitable for publication in PLOS One. Congratulations! Your manuscript is now being handed over to our production team.

At this stage, our production department will prepare your paper for publication. This includes ensuring the following:

* All references, tables, and figures are properly cited

* All relevant supporting information is included in the manuscript submission,

* There are no issues that prevent the paper from being properly typeset

You will receive further instructions from the production team, including instructions on how to review your proof when it is ready. Please keep in mind that we are working through a large volume of accepted articles, so please give us a few days to review your paper and let you know the next and final steps.

Lastly, if your institution or institutions have a press office, please let them know about your upcoming paper now to help maximize its impact. If they'll be preparing press materials, please inform our press team within the next 48 hours. Your manuscript will remain under strict press embargo until 2 pm Eastern Time on the date of publication. For more information, please contact onepress@plos.org.

You will receive an invoice from PLOS for your publication fee after your manuscript has reached the completed accept phase. If you receive an email requesting payment before acceptance or for any other service, this may be a phishing scheme. Learn how to identify phishing emails and protect your accounts at https://explore.plos.org/phishing.

If we can help with anything else, please email us at customercare@plos.org.

Thank you for submitting your work to PLOS ONE and supporting open access.

Kind regards,

PLOS ONE Editorial Office Staff

on behalf of

Dr. Chiara Lazzeri

Academic Editor

PLOS One

Associated Data

    This section collects any data citations, data availability statements, or supplementary materials included in this article.

    Supplementary Materials

    S1 Fig. Boruta algorithm feature selection results for XGBoost modeling.

    Box plots display the Z scores of each parameter, with the x-axis showing their names and the y-axis showing the Z values. Parameters are color-coded by importance (blue for important, green for tentative, and red for unimportant).

    (TIF)

    pone.0348676.s001.tif (5.2MB, tif)
    S2 Fig. Feature collinearity assessment in XGBoost modeling.

    (TIF)

    pone.0348676.s002.tif (4.8MB, tif)
    S3 Fig. Unadjusted associations between NLR levels and mortality outcomes.

    Panels A-C correspond to 28-day mortality (A), hospital mortality (B), and ICU mortality (C), respectively. Error bars represent 95% CIs, with the intermediate NLR concentration group serving as the reference category.

    (TIF)

    pone.0348676.s003.tif (511.8KB, tif)
    S4 Fig. Restricted cubic spline plots illustrating the nonlinear relationship between the NLR and mortality.

    (A) 28-day mortality, (B) in-hospital mortality, and (C) ICU mortality. The solid line represents the adjusted odds ratio; the pink band indicates the 95% confidence interval. The histogram displays the distribution of the NLR values, with the median (7.6) and the ROC-derived optimal cutoff (8.08) marked. The nonlinear association was statistically significant for all of the endpoints (P for nonlinearity < 0.05).

    (TIFF)

    pone.0348676.s004.tiff (1.4MB, tiff)
    S1 Table. Missing rate for demographics and clinical variables extracted from the database during the observation period.

    (PDF)

    pone.0348676.s005.pdf (101.6KB, pdf)
    S2 Table. Comparison of Key Variables Before and After Imputation Using the misForest Algorithm.

    (PDF)

    pone.0348676.s006.pdf (94.4KB, pdf)
    S3 Table. Selected features and discarded features of collinearity assessment in XGBoost modeling.

    (PDF)

    pone.0348676.s007.pdf (102KB, pdf)
    S4 Table. Performance evaluation of XGBoost model across training and test datasets.

    (PDF)

    pone.0348676.s008.pdf (118.9KB, pdf)
    S5 Table. Feature importances in the XGBoost model.

    (PDF)

    pone.0348676.s009.pdf (105.6KB, pdf)
    S6 Table. Subgroup analysis for the association of the NLR with 28-day mortality, hospital mortality and ICU mortality.

    (PDF)

    pone.0348676.s010.pdf (88.3KB, pdf)
    S7 Table. Abbreviations.

    (PDF)

    pone.0348676.s011.pdf (100.3KB, pdf)
    Attachment

    Submitted filename: Response to Reviewers.docx

    pone.0348676.s013.docx (29.3KB, docx)

    Data Availability Statement

    The complete minimal anonymized dataset necessary to replicate all study findings is publicly available from Zenodo (https://doi.org/10.5281/zenodo.20139761).


    Articles from PLOS One are provided here courtesy of PLOS

    RESOURCES