Skip to main content
iScience logoLink to iScience
. 2026 Jan 29;29(3):114821. doi: 10.1016/j.isci.2026.114821

Stratifying mortality risk in hypotension using outcome-oriented clustering of vital sign trajectories

Zhuolun Liu 1,4, Li Li 2,3,4, Yongmin Zhang 1,5,∗, Shasha Xue 1, Lina Zhang 2,3,∗∗, Zhaoxin Qian 2,3,∗∗∗, Shuai Zhang 1, Feng Lyu 1, Fan Wu 1
PMCID: PMC12969146  PMID: 41809043

Summary

The prognostic significance of early vital sign trajectories in critically ill patients with hypotension remains unclear. This study applied an outcome-oriented two-step clustering model to characterize vital sign trajectories within 12 h of symptom onset, with the optimal window determined by weight learning, to improve mortality risk stratification. Patients (mean arterial pressure ≤65 mmHg or on vasopressors) were from the MIMIC-IV (development), eICU, and Xiangya multi-ICU (validation) databases. Three clusters with progressively increasing mortality were identified. Cluster 3 vs. 1 had the highest adjusted hazard ratio (2.30, 95%CI [2.04, 2.58]). The respiratory rate trajectory within the initial 2 h was the primary mortality discriminator. External validation demonstrated the model’s generalizability, with hazard ratios (eICU, 2.96, 95% CI: 2.70, 3.23) and (Xiangya, 2.28, 95% CI: 1.93, 2.71) for cluster 3 vs. 1, respectively. This study demonstrated that hypotension can be risk-stratified by respiratory rate trajectories, with elevated breaths linked to higher mortality.

Subject areas: Health sciences, Medicine, Medical specialty, Internal medicine, Cardiovascular medicine, Intensive care medicine

Graphical abstract

graphic file with name fx1.jpg

Highlights

  • •

    Early respiratory rates were the strongest predictor of mortality risk in hypotension

  • •

    Outcome-oriented clustering identified 2 h as the optimal analysis window

  • •

    The model, using simple vital signs, is easily integrated for clinical early warning

  • •

    External validation across multiple cohorts confirmed the model’s generalizability


Health sciences; Medicine; Medical specialty; Internal medicine; Cardiovascular medicine; Intensive care medicine

Introduction

Hypotension occurs in 10%–20% of ICU patients and results in 30%–43% mortality rate.1 If not promptly managed, it can lead to a vicious cycle of systemic ischemic-hypoxic damage and multi-organ dysfunction. The early identification of patients at risk for hypotension was achieved and validated by well-designed machine learning models.1,2,3,4 Although these models demonstrate robust performance in external cohort validation,3 their clinical adoption faces two key challenges. First, their reliance on numerous metrics necessitates integration with electronic medical records or real-time monitoring systems, imposing technical and equipment demands. Second, focusing solely on predicting the onset of hemodynamic instability provides limited clinical utility, as it neither quantifies severity nor informs the direction of intervention.

Hypotension arises from diverse, interconnected causes—such as volume depletion, cardiac dysfunction, and vasoplegia. Despite this etiologic complexity, the condition often manifests in a phenotypically similar way through decreased blood pressure and impaired perfusion. This ambiguity obscures prognosis and underscores the critical need for a model-driven, interpretable system to stratify mortality risk in the ICU.5,6

Vital signs are among the most extensively and intensively monitored parameters in the ICU, making predictive models based on vital signs highly valuable, even in resource-limited settings. Group-based trajectory modeling has been widely applied to identify subphenotypes of vital sign trajectories across various critical syndromes, demonstrating strong potential for distinguishing mortality risk and treatment responses. Bhavani7 analyzed vital signs within the first 8 h of hospitalization, identifying four novel sepsis subphenotypes with distinct outcomes and differing therapeutic responses to balanced crystalloids versus saline. Similarly, Li et al.8 demonstrated that longitudinal subphenotypes based on respiratory rate (RR) and heart rate during the first 72 h of ICU stay were predictive of 28-sday mortality in elderly patients with acute respiratory distress syndrome (ARDS). As such, we hypothesize that early vital sign trajectories encode the phenotypic characteristics and mortality risk of hypotension.

However, the identification of clinically meaningful phenotypes from high-dimensional features, exemplified by vital sign time-series data, presents challenges. The indiscriminate use of all variables within the complete available time window for clustering may introduce redundant features, namely noise,9 obscuring the true distances between samples in the high-dimensional feature space that are critical for identifying desired subgroups.10 Consequently, it’s essential to develop a clustering strategy enabling optimization of feature selection and time window. This study aims to: (1) identify outcome-oriented subgroups by clustering vital sign trajectories that stratify mortality risk in patients with hypotension; (2) determine the optimal combination of vital signs and time windows for trajectory analysis to maximize parsimony without compromising performance; and (3) assess the robustness and generalizability through external validation in two transpacific cohorts. The overall framework of this study is shown in Figure 1.

Figure 1.

Figure 1

The overall framework of the clustering strategy and analysis method

Results

Baseline characteristics of cohorts

This study included 13,654, 21,562, and 4,146 patients with hypotension from MIMIC IV, eICU, and Xiangya databases respectively. The flowchart of inclusion and exclusion in the three cohorts are shown in Figures S1–S13. Baseline characteristics of the cohorts are shown in Table 1.

Table 1.

Baseline characteristics of the three datasets used in this study

MIMIC IV (n = 13,654) eICU (n = 21,562) Xiangya (n = 4,146)
Demographic, n (%)

Female 5,195 (38.0) 8,578 (39.8) 1,492 (36.0)
Age mean (SD) (years) 67.9 (13.9) 64.7 (14.8) 57.0 (15.3)

Comorbidities, n (%)

Hypertension 6,595 (48.3) 11,420 (53.0) 1,428 (34.4)
Diabetes 2,178 (16.0) 2,889 (13.4) 772 (18.6)
Heart disease 4,423 (32.4) 1,782 (8.3) 593 (14.3)
kidney disease 2,668 (19.5) 3,186 (14.8) 417 (10.1)
Liver disease 538 (3.9) 220 (1.0) 438 (10.6)
Atrial fibrillation 4,841 (35.4) 2,469 (11.5) 78 (1.9)
Asthma 1,094 (8.0) 1,218 (5.6) 31 (0.7)
Chronic obstructive pulmonary disease 755 (5.5) 3,187 (14.8) 396 (9.6)

Major intervention, n (%)

Vasopressor∗ 10,257 (75.1) 12,423 (57.6) 3,723 (89.8)
Inotropes∗ 854 (6.3) 1,928 (8.9) 307 (7.4)
Invasive ventilation 9,101 (66.7) 16,480 (76.4) 2,333 (56.3)
CRRT 893 (6.5) NA 6 (0.1)
ECMO 46 (0.3) NA 97 (2.3)

Severity score & outcome, median (IQR)

SOFA score at T0 6 (4, 9) 6 (3, 9) 9 (7, 12)
Length of ICU stay (days) 3.1 (1.7, 6.4) 3.8 (2.1, 7.2) 7.9 (3.4, 16.5)
Length of hospital stay (days) 8.8 (5.8, 14.7) 9.4 (5.8, 16.2) 15 (8, 27)
In-hospital mortality, n (%) 1,722 (12.6) 3,735 (17.3) 910 (21.9)

Notes: CRRT, continuous renal replacement treatment; ECMO, extracorporeal membrane oxygenation; IQR, interquartile range; SOFA, sequential organ failure assessment; NA, not applicable. Vasopressors∗ include: norepinephrine, vasopressin, terlipressin, dopamine, phenylephrine, epinephrine, angiotensin II, droxidopa, ephedrine; Inotropes∗ includes: dobutamine, milrinone.

Vital signs selection

The selection of vital signs relied on the full 12-h time window. The overall SHapley Additive exPlanations (SHAP) values for time-series vital signs, normalized across different classifiers, are shown in Figure 2 The standard deviation of mortality rates for different combinations of vital signs at each effective cluster number was calculated, as shown in Figure 3. Compared to variable combinations with more vital signs, clustering based solely on RR exhibited the highest standard deviation in mortality rates. Although three out of four evaluation metrics indicated that two was the optimal number of clusters (Figure S4A), we prioritized the identification of intermediary risk groups by choosing three clusters, a decision that may reduce inter-cluster separation as a trade-off but enabled a further division of the lower-risk subgroup, revealing a gradient of phenotypic and outcome differences. This facilitated subsequent analysis of the relationship between phenotypes and outcomes. Finally, clustering based on time-series RR identified three subgroups with distinct mortality rates.

Figure 2.

Figure 2

Normalized total SHAP values for the time series of vital signs in each model

The variables in the figure are arranged in ascending order of the mean normalized total SHAP values on the five models. Notes: RR, respiratory rate; SpO2, pulse oxygen saturation; HR, heart rate; MAP, mean arterial pressure; SAP, systolic arterial pressure; DAP, diastolic arterial pressure.

Figure 3.

Figure 3

The standard deviation of subgroup mortality for each combination of variables at each effective number of clusters

The x axis represents various variable combinations, accumulated sequentially from RR onward. Clusters with a population size of less than 5% of the total or whose survival curves did not show significant differences (p < 0.05) from those of any other clusters were considered invalid, and the standard deviation of subgroup mortality was not calculated for these clusters.

It is also noteworthy that the Hopkins statistic for the time series of systolic arterial pressure exceeded 0.85, while the statistics for the other time-series vital signs were all above 0.90. This indicates that the distribution of all time-series vital signs was significantly different from a random uniform distribution, suggesting their suitability for clustering.

Vital signs trajectories and mortality of subgroups

The development cohort exhibited three subgroups with stratified mortality risk (Figure 4), presenting distinct trajectories of vital signs from the onset of hypotension to 12 h thereafter, as illustrated in Figure 5A. Cluster 1 had the lowest in-hospital mortality (8.5%, 95%CI [7.7%, 9.3%]) with lowest RR (range of median11,12 breaths/min) across 12h time window; cluster 2 had intermediate mortality (12.3%, 95%CI [11.6%, 13.0%]) and rapider RR (range of median13,14,15 rates/min); while cluster 3 had the highest mortality (32. 7%, 95%CI [29.9%, 35.4%]) with the rapidest RR (range of median [26–28] rates/min). Compared to cluster 1 and 2, cluster 3 had higher heart rate (range of median: cluster 1 vs. 2 vs. 3 [78–80] vs. [80–84] vs. [91–93] rates/min, p < 0.001), and lower SPO2 (range of median: cluster 1 vs. 2 vs. 3 [98%–99%] vs. [97%–99%] vs. [97%–97%], p < 0.001) As shown, there was moderate overlap in interquartile ranges of heart rate and SPO2. The dynamic blood pressure trajectories of the three clusters demonstrate substantial overlap.

Figure 4.

Figure 4

Survival curves for subgroups based on 12-h and 2-h RRs from the three cohorts

(A, D) Survival curves for subgroups identified within the 12-h and 2-h time windows, respectively, from MIMIC IV. (B, E) Survival curves for the 12-h and 2-h windows from eICU. (C, F) Survival curves for the 12-h and 2-h windows from Xiangya. Limited by data availability, in-hospital survival curves were drawn and ended at 90 days. The number at risk at every 10-day interval is annotated on the corresponding coordinates. 90-day survival curves for MIMIC IV were shown in Figure S8. Notes: RR, respiratory rate.

Figure 5.

Figure 5

Trajectories of vital signs in subgroups based on the 12-h time window

(A) Trajectories of vital signs in subgroups from MIMIC IV, (B) Corresponding trajectories from eICU, (C) Corresponding trajectories from Xiangya. Upon testing, none of the vital signs at each time step follow a normal distribution. Therefore, the trajectory is composed of the median vital signs at each time step within each subgroup. The shaded area represents the interquartile range (IQR). Notes: RR, respiratory rate; SpO2, pulse oxygen saturation; HR, heart rate; MAP, mean arterial pressure; SAP, systolic arterial pressure; DAP, diastolic arterial pressure.

The external validation of the model yielded similar results (Figure 4). The eICU cohort was divided into three clusters, with in-hospital mortality for cluster 1 vs. 2 vs. 3 being 14.1% (95%CI [13.5%, 14.7%]), 18.4% (95%CI [17.41%, 19.4%]), and 31.9% (95%CI [30.0%, 33.8%]), respectively. The fluctuations and trends of vital sign trajectories across the three clusters were nearly identical to those observed in the MIMIC IV cohort. The Xiangya cohort also exhibited three clusters with different in-hospital mortality risks (in-hospital mortality, Cluster 1 vs. 2 vs. 3, 16.2% (95%CI [14.4%, 17.9%]) vs. 22.6% (95%CI [20.5%, 24.6%]) vs. 32.8% (95%CI [29. 5%, 36.2%]). Notably, the mortality in Cluster 1 and Cluster 2 of the Xiangya cohort was significantly higher, with heart rate (range of median: cluster 1 vs. 2 vs. 3 [86–89] vs. [96–100] vs. [108–112], p < 0.001) notably exceeding those observed in the MIMIC IV and eICU cohorts. In all cohorts, the mortality risk was positively correlated with the overall level of RR for all three subgroups (in all cohorts, p < 0.001).

Characteristics and mortality predictors of subgroups

Table S1 summarizes the demographic, laboratory, and outcome characteristics of the three subgroups in the MIMIC-IV cohort. Cluster 3 was characterized by the longest ICU and hospital stays and the highest in-hospital mortality, with significant differences observed across all subgroups. Meanwhile, this group received continuous renal replacement therapy and vasopressor support more frequently, despite having a lower prevalence of comorbidities such as hypertension, heart disease, and atrial fibrillation. The mean age was 63.8 ± 16.4 years, significantly younger than those in cluster 1 (67.9 ± 13.9 years) and cluster 2 (67.4 ± 14.4 years).

Moreover, the XGBoost classification model identified key mortality predictors within each subgroup, incorporating laboratory and demographic variables. The AUCs for the XGBoost classifiers targeting cluster 1, cluster 2, and cluster 3 on the test set were 0.85 ± 0.01, 0.87 ± 0.01, and 0.79 ± 0.01, respectively, showing an improvement over the performance of the model trained on the overall population, which achieved AUCs of 0.84 ± 0.01, 0.84 ± 0.01, and 0.75 ± 0.03 (Figure S9). These results support the potential of the subgroups to inform targeted outcome prediction or clinical management. Furthermore, the top five mortality predictors ranked by average absolute SHAP values, and their SHAP value distributions across all samples are presented in Figure 6. Minimum lactate emerged as the primary mortality predictor in the high-risk subgroup, maximum anion gap in the intermediate-risk subgroup, and maximum glucose in the low-risk subgroup.

Figure 6.

Figure 6

Critical mortality influences across subtypes in MIMIC IV based on 12-h RRs

Notes: max, maximum; min, minimum; bun, blood urea nitrogen; pt, prothrombin time; po2, arterial partial pressure of oxygen; abs, absolute; ph, potential of hydrogen; ptt, partial thromboplastin time.

Time window optimization and cox regression analysis

In the MIMIC-IV cohort, time-series RRs were identified as the sole determinant for subgroups with the greatest outcome differences. The learned average weight of the first three time-steps (T0–T2) was 0.754 ± 0.073, 0.198 ± 0.038, 0.048 ± 0.041 respectively (Table 2), contributing majority of the total weight. Mortality prediction based solely on the weighted average of RR at the first three time-steps achieved an AUC of 0.68 on the validation set, comparable to the XGBoost model using the full RR time window (AUC = 0.68).

Table 2.

The weight of the weighted average RRs at each time step

Time steps T0 T1 T2 T3 T4 T5 T6 T7 T8 T9 T10 T11 T12
Mean of the weights 0.754 0.198 0.048 <0.0001 <0.0001 <0.0001 <0.0001 <0.0001 <0.0001 <0.0001 <0.0001 <0.0001 <0.0001
Standard deviation 0.073 0.038 0.041 <0.0001 <0.0001 <0.0001 <0.0001 <0.0001 <0.0001 <0.0001 <0.0001 <0.0001 <0.0001

In the 2-h clustering model, both cluster 2 and cluster 3 were associated with a significantly higher risk of in-hospital mortality compared with cluster 1 across all cohorts (Table 3). In the MIMIC-IV cohort, the hazard ratios (HRs, 95%CI) for mortality were 2.04 (1.87, 2.23) for cluster 2 and 3.73 (3.32, 4.18) for cluster 3. Similar trends were observed in the eICU cohort, with HRs of 1.53 (1.43, 1.65) and 2.96 (2.70, 3.23), respectively, and in the Xiangya cohort, with ratios of 1.52 (1.31, 1.76) and 2.28 (1.93, 2.71). These findings indicate robust and consistent performance of the model across external validation datasets. Results based on the 12-h clustering were comparable, showing elevated mortality risks for cluster 3 versus cluster 1 in the MIMIC-IV (HR 3.48, 95%CI [3.09, 3.93]), eICU 2.56 (2.35, 2.78), and Xiangya 2.18 (1.84, 2.59) cohorts.

Table 3.

Hazard ratios of mortality between clusters within each time window

Clusters MIMIC IV eICU Xiangya
12 h

3 vs.1 3.48 (3.09,3.93) 2.56 (2.35,2.78) 2.18 (1.84,2.59)
2 vs.1 1.31 (1.20,1.44) 1.33 (1.24,1.43) 1.43 (1.22,1.66)

2 h

3 vs.1 3.73 (3.32,4.18) 2.96 (2.70,3.23) 2.28 (1.93,2.71)
2 vs.1 2.04 (1.87,2.23) 1.53 (1.43,1.65) 1.52 (1.31,1.76)

In the MIMIC-IV cohort, multivariable Cox proportional hazards analysis demonstrated that, after adjustment for potential confounders, cluster membership remained a strong independent predictor of in-hospital mortality (Figure 7). For the 2-h clustering model, the adjusted HR (95% CI) were 1.46 (1.35, 1.58) for cluster 2 and 2.30 (2.04, 2.58) for cluster 3, ranking among the top two predictors among all covariates. In the subgroup of patients with septic shock, who accounted for 61% of the MIMIC-IV cohort, similar associations were observed (cluster 2: HR 1.50 [1.37, 1.64]; cluster 3: HR 2.26 [2.00, 2.55]). Results based on the 12-h clustering model were consistent, with cluster 3 remaining the strongest independent predictor of in-hospital mortality (HR 2.34, [2.07, 2.64]), while cluster 2 was also an independent predictor but of lesser magnitude. Furthermore, multiple rounds of clustering based on three randomly selected time steps resulted in subgroup survival curves with clearly lower HRs than the current results (Figure S12). This further validates the rationality of the selected time window.

Figure 7.

Figure 7

Independent prognostic factors associated with survival outcome identified by multivariate Cox regression analysis

The analysis was conducted across all clusters derived from the 2-h RRs, with a specific focus on the septic shock subgroup. The forest plot presents the adjusted hazard ratios and 95% confidence intervals of variables that retained independent statistical significance (p < 0.05).

In different cohorts, the subgroups based on 2-h time window exhibited similar sample distributions and clustering structures, confirming the generalizability of this strategy (Figure S10). The median ranges of RR across 2 h in cluster 1, 2, and 3 were as follows: MIMIC IV (16–16), (20–22), and (28–28); eICU (15–15), (21–21), and (29–29); and Xiangya (16–16), (23–24), and (30–31) (p < 0.001). For heart rate, the corresponding medians were as follows: MIMIC IV (79–80), (84–85), and (93–95); eICU (81–81), (87–88), and (96–96); and Xiangya (89–90), (103–105), and (117–128) (p < 0.001).

Discussion

This study developed an outcome-oriented clustering framework that identified three prognostically distinct subgroups of hypotensive patients based on early vital sign trajectories. The 2-h RR pattern showed the strongest mortality discrimination, validated externally, underscoring its potential utility in diverse clinical settings.

Within the first 12 h after hypotension onset, RR trajectories diverged markedly across subgroups. High-risk patients (cluster 3) showed persistently rapid RR (28–30 breaths/min), intermediate-risk patients (cluster 2) had moderate rates (20–24 breaths/min), and low-risk patients (cluster 1) maintained stable, slower RR (15–16 breaths/min). Heart rate exhibited a similar pattern, though less distinct, while arterial pressure showed minimal separation. After adjustment for confounders, the cluster variable remained the strongest independent predictor, including in septic shock. This finding was in line with prior evidence from sepsis9,14 and ward populations,15,16 identifying RR as the most sensitive predictor of mortality. Mechanistically, early RR dynamics may capture key aspects of disease progression and host response, likely representing a compensatory response to hypoxia and acidosis and integrating circulatory, metabolic, and neurogenic signals. In contrast, arterial pressure remained deceptively stable under hemodynamic treatment. Consistent with this interpretation, RR exhibited smaller fluctuations over 0–12 h, whereas heart rate and blood pressure increased during the first 2 h before stabilizing.

Using an outcome-oriented two-step clustering framework that applied time-dependent weighting based on outcome relevance, this study demonstrated that a 2-h observation window achieved mortality discrimination comparable to the full 12-h period. This suggests that early RR dynamics capture critical information about the magnitude and direction of the host response—whether compensatory or declining—providing similar prognostic value to longer-term trajectories. The short time window enhanced model efficiency and improved clinical feasibility. Furthermore, consistent performance across independent cohorts from the United States and China underscores the model’s robustness, generalizability, and potential for broad clinical application. Although baseline characteristics may introduce subtle differences in the RR ranges used for risk stratification across populations, this does not substantially limit the generalizability of our model. By inputting readily available retrospective RR data into our model, population-specific risk stratification ranges can be rapidly calibrated.

An interesting observation was that in the model development cohort, patients in cluster 3 were younger and had lower rates of hypertension and atrial fibrillation, with only a slightly higher prevalence of chronic obstructive pulmonary disease. Paradoxically, this subgroup exhibited the highest mortality risk. Multivariable analysis showed that only advanced age remained an independent risk factor, suggesting that the prognostic effect of early respiratory patterns outweighed that of baseline comorbidities.

Clinically, our model could be integrated into bedside decision-support tools or early-warning systems to identify high-risk patients and prompt timely investigation (e.g., acid-base assessment) and intervention (e.g., ventilatory optimization, modulation of respiratory drive).17,18 In managing circulatory instability, RR remains the most overlooked vital sign—clinicians often prioritize blood pressure, heart rate, or oxygen saturation over tachypnea. This underrecognition may reflect limited awareness of its prognostic value and the complexity of adjusting RR, as sedation or analgesia may necessitate invasive ventilation. Identifying patient phenotypes characterized by early tachypnea, therefore, carries therapeutic significance. As shown in phenotype-guided approaches in ARDS19 and sepsis,20 integrating such RR-based phenotypes into predictive and decision-support systems could enable more individualized and timely interventions.

In conclusion, this study stratified hypotension patients into three subgroups with distinct mortality risks based on RR trajectories within the first 2 h of onset. After adjustment for confounders, cluster variable was the strongest predictor. These findings have significant therapeutic implications, as the model demonstrated consistent and robust performance across diverse regions. Further research on targeted treatments for each subgroup may enhance individualized and precise management.

Limitations of the study

This study has several limitations. As a retrospective analysis, the findings are inherently vulnerable to biases related to data quality and confounders. Although rigorous preprocessing procedures were applied, missing or inaccurate data may still introduce bias into the identified Vital Sign trajectory patterns of the subgroups. Compared with heart rate and blood pressure, recordings of RR are generally less precise and often require manual verification by nursing staff, introducing potential subjectivity. Nevertheless, hourly RR recordings by well-trained ICU nurses are considered relatively reliable,21 as evidenced by the clear separation of RR trajectory quartiles across the three subgroups. Consciousness level and temperature trajectories were not included as predictors due to limited data availability and low variability. The temporal resolution of the 12-hourly time series may not fully represent physiological dynamics that could affect outcomes. Extending the observation window or increasing sampling granularity may provide deeper insights into trajectory clustering and patient stratification.

Additionally, our analysis did not account for specific treatments administered during the study period, such as mechanical ventilation or vasopressor use, which may have influenced outcomes. Future research will aim to explore personalized treatment strategies tailored to distinct subgroups with varying health statuses, with an emphasis on mutual validation and interpretation in alignment with the findings of this study.

Resource availability

Lead contact

Further information and requests for resources and reagents should be directed to and will be fulfilled by the lead contact, Yongmin Zhang (zhangyongmin@csu.edu.cn).

Materials availability

This study did not generate new unique reagents.

Data and code availability

  • •

    The dataset of Xiangya used and/or analyzed during this study are not publicly available due to the confidentiality policy of the National Health Commission of China but are available from the corresponding author L.Z. upon reasonable request.

  • •

    This study does not report original code.

  • •

    Any additional information required to reanalyze the data reported in this article is available from the lead contact upon request.

Acknowledgments

This study was funded by the Key Research and Development Program of Hunan Province, China (2023SK2020, 2022SK2039), the National Key Research and Development Program of China (2022YFC2009800), and Central South University Research Program of Advanced Interdisciplinary Studies (grant no. 2023QYJC022).

Author contributions

Z.Q., L.Z., L.L., Y.Z., and Z.L. conceived the study, and Z.L. performed, designed, and built the models. L.L., Z.L., S.X., and S.Z. extracted data from the databases and performed the analysis. L.L. and Z.L. interpreted the results and drafted the manuscript. Z.Q., L.Z., Y.Z., F.L., and F.W. revised the manuscript for important intellectual content. All the authors read and approved the final manuscript. The corresponding author had full access to all data in the study and had final responsibility for the decision to submit for publication.

Declaration of interests

The authors declare that they have no competing interests.

STAR★Methods

Key resources table

REAGENT or RESOURCE SOURCE IDENTIFIER
Deposited data

MIMIC IV database Johnson et al.22 Johnson et al.22
eICU database Pollard et al.23 Pollard et al.23
the Multi-ICU database of Xiangya Hospital, Central South University This paper The dataset of Xiangya used and/or analyzed during this study are not publicly available due to the confidentiality policy of the National Health Commission of China but are available from the corresponding author Lina Zhang upon reasonable request.

Software and algorithms

Python version 3.9.0 Python Software Foundation https://www.python.org/downloads/release/python-390/
fastdtw = = 0.3.4 PyPI https://pypi.org/project/fastdtw/
shap = = 0.45.1 PyPI https://pypi.org/project/shap/0.45.1/
xgboost = = 2.0.3 PyPI https://pypi.org/project/xgboost/
scikit-learn = = 1.5.2 PyPI https://pypi.org/project/scikit-learn/
scipy = = 1.11.2 PyPI https://pypi.org/project/scipy/
lightgbm = = 4.3.0 PyPI https://pypi.org/project/lightgbm/
lifelines = = 0.28.0 PyPI https://pypi.org/project/lifelines/
scikit-survival = = 0.23.1 PyPI https://pypi.org/project/scikit-survival/
statsmodels = = 0.14.1 PyPI https://pypi.org/project/statsmodels/

Experimental model and study participant details

This study utilized data from MIMIC IV: https://doi.org/10.1038/s41597-022-01899-x,22 eICU: https://doi.org/10.1038/sdata.2018.178,23 and the Multi-ICU database of Xiangya Hospital, Central South University. MIMIC-IV and eICU are publicly available datasets of critically ill patients, accessible to researchers upon completion of a credentialing process. MIMIC IV consists of anonymized health-related data from 73,181 patients admitted to the intensive care units of Beth Israel Deaconess Medical Center between 2008 and 2019. eICU consists of 200,859 patient unit encounters for 139,367 unique patients admitted between 2014 and 2015. These patients were admitted to one of 335 units at 208 hospitals in the US. The Xiangya database is a proprietary database built on multi-ICU data including patients admitted to seven adult ICUs of Xiangya Hospital from September 2011 and November 2023. Xiangya Hospital is a tertiary teaching hospital in central southern China. All these databases have undergone ethical review to ensure patient privacy protection and have waived the requirement for informed consent, as the data are de-identified and do not contain personally identifiable information. The authors have signed the Data Use Agreements for both MIMIC IV and eICU and have completed the CITI Program course on “Data or Specimens Only Research” (record ID 65878900). The study was approved by the institution’s Internal Review Board (Medical Ethics Committee of Xiangya Hospital Central South University, registry number 202309806).

Method details

Patient cohort

Patient inclusion criteria for all cohorts were (1) Patients with mean arterial pressure (MAP) ≤ 65 mmHg, or receiving vasopressors or inotropes2; (2) Patients with age ≥18 years; and (3) Patients with ICU length of stay ≥24 h.

Data pre-processing

Based on the considerations regarding the accessibility and prognostic value of vital signs discussed above, six time-series vital signs were collected as potential variables for clustering, including respiratory rate (RR), heart rate, pulse oxygen saturation (SpO2), mean arterial pressure (MAP), systolic arterial pressure (SAP), and diastolic arterial pressure (DAP). We utilized the raw measurements of invasive arterial blood pressure, without application of any calibrations. The onset of hypotension (T0) was defined as the earliest time of either the occurrence of MAP ≤65 mmHg or the initial administration of vasopressors. A complete temporal vital sign consisted of the value at T0 and the mean values recorded hourly during the 12 h following T0, resulting in a total of 13 values within the 12-h time window. Extreme outliers were set to missing based on predefined ranges.2 Missing data were imputed using forward filling for terminal gaps and linear interpolation for internal gaps, with a maximum of four consecutive time steps filled. Samples with any remaining missing values were excluded, and all variables were subsequently standardized using Z score normalization before modeling.

To understand the clinical characteristic differences among subgroups and identify key mortality factors, a total of 68 variables, including demographics, respiratory, circulation, metabolism, coagulation, enzymes, electrolytes, liver and renal function on the first day admitted to ICU, were collected. Maximum and minimum value of each variable were extracted. Features with more than 25% missing data were excluded. The remaining missing data were imputed using the K-Nearest Neighbor algorithm,24,25 and these variables were also Z score normalized.

Vital signs selection

As previously mentioned, the expected subgroups may only be driven by a subset of vital signs, and using all the vital signs for clustering may introduce redundant features, namely noise. The noise could obscure the true sample distances in the feature space, impairing the clustering model’s ability to identify expected subgroups. Therefore, an outcome-oriented trajectory modeling approach is required to screen key features and eliminate irrelevant noise. In the development cohort, SHapley Additive exPlanations (SHAP) values11 were computed for five commonly used classifiers—Support Vector Machine, Logistic Regression, XGBoost, LightGBM, and Random Forest—to quantify the contribution of each time-series vital sign to patient mortality. The Hopkins statistic was then employed to evaluate the spatial randomness of feature distributions and assess their suitability for clustering.12

The average absolute SHAP values for each time step of the six vital signs were computed across the five classifiers, and the overall SHAP values for each vital sign were aggregated across all time steps. The normalized overall SHAP values from the five classifiers were used to rank the variables. Then, variables ranked higher were added to the clustering model one at a time to form various combinations. For each combination, the standard deviation of in-hospital mortality within clusters and the log rank test on survival curves were employed to assess differences in outcomes. To ensure the stability and generalizability of subgroups, clusters with a sample size smaller than 5% of the total population were not considered. The final aim of variable selection was to identify the optimal combination of the vital signs to drive a subgroup structure with the most significant differences in outcomes.

Trajectory clustering & survival analysis

The trajectories of vital signs were clustered using Hierarchical Agglomerative Clustering (HAC)26 based on Dynamic Time Warping (DTW).13 DTW aligns two time series non-linearly by minimizing their cumulative Euclidean distance to capture their underlying similarities. The DTW similarity matrix of the selected time-series vital signs was computed, followed by applying HAC to this matrix to partition the patients into subgroups. Four internal validation metrics (Sum of Squared Errors, Silhouette Coefficient, Davies-Bouldin Index, and Calinski-Harabasz Index) were used to determine the optimal number of clusters, ranging from 2 to 11.

Cox proportional hazards models were fitted to assess the association between cluster membership and mortality. Univariate analyses were first performed for each cluster, followed by multivariable models adjusted for potential confounders, including demographic characteristics, pre-existing comorbidities, and etiological factors contributing to hypotension.

Characteristics analysis of subgroups

To understand the clinical phenotypes of the subgroups better, relevant analyses were performed in the MIMIC IV cohort. Firstly, Student’s t tests, Kruskal-Wallis tests and chi-square tests were conducted in the demographic, laboratory, and outcome variables to compare the statistical differences between subgroups. Furthermore, three binary classification models based on XGBoost were constructed and interpreted using the SHAP values. The models were trained with in-hospital mortality as the label and samples from within each subgroup to identify the crucial predictors of mortality specific to each cluster. To further evaluate the clinical relevance of the identified clusters in guiding differentiated intervention or outcome prediction, we compared the performance of the best global model, trained on the overall population, against the three subgroup-specific models. All models were evaluated on the test sets corresponding to each cluster.

Time window optimization

The determination of an optimal observation time window for dynamic trajectories involves a fundamental trade-off: excessively long windows risk introducing noise and data gaps, whereas overly short windows may fail to capture essential patterns, thereby compromising predictive accuracy. To address the optimal observation time window, we utilized the weighted average of vital signs across sequential time steps. The weight values were learned from MIMIC IV cohort to identify the time steps most predictive of mortality risk. Specifically, the metric is defined as μw = ∑t=1Twtxt, where xt denotes the vital sign value at time step t, and wtis the corresponding weight. The weights satisfy the constraints ∑t=1Twt = 1 and 0≤wt ≤ 1. These constraints ensure stable convergence of the model within a reasonable range. The loss function is defined as Loss = μalive/μdead, where μalive and μdead represent the mean of the weighted averages of 12-h RRs within the surviving and deceased patient groups, respectively. The weights are iteratively adjusted through backpropagation and gradient updates to optimize the loss. By minimizing the loss, the distinct in the weighted average metrics between the deceased and surviving patient groups is maximized, thereby enhancing the association between the metric and patient outcomes. Furthermore, the weight at each time step captures its contribution to the relationship with patient outcomes.

A 5-fold cross-validation strategy was employed for internal validation, and the resulting weights were averaged across folds to obtain robust estimates. Hazard ratios of mortality between subgroups before and after optimization was compared to assess the effect of optimization.

External validation

The generalizability of the clustering strategy was validated in eICU and Xiangya cohorts. Patients were clustered using the same method based on the selected time-series vital signs, and the optimal number of clusters was determined. The progression patterns of vital signs, mortality, sample distribution, and clustering structure were visually compared among the subgroups from the three cohorts to verify the generalizability of the clustering strategy.

Quantification and statistical analysisb

Epidemiologic, clinical, and routine laboratory statistical analysis

Statistical analyses were performed using Python (version 3.9.0). Normal distribution and variance homogeneity of data were assessed using the Shapiro–Wilk test and Levene’s tests, respectively. Normally distributed variables were expressed as mean and standard deviation and compared using Student’s t tests. Non-normally distributed variables were expressed as median and 25th and 75th percentiles and compared using Kruskal-Wallis tests. Categorical data were compared using the chi-squared test. A p-value <0.05 was considered statistically significant.

Survival analysis

Cox proportional hazards models were fitted to assess the association between cluster membership and mortality. Univariate analyses were first performed for each cluster, followed by multivariable models adjusted for potential confounders, including demographic characteristics, pre-existing comorbidities, and etiological factors contributing to hypotension.

Published: January 29, 2026

Footnotes

Supplemental information can be found online at https://doi.org/10.1016/j.isci.2026.114821.

Contributor Information

Yongmin Zhang, Email: zhangyongmin@csu.edu.cn.

Lina Zhang, Email: zln7095@csu.edu.cn.

Zhaoxin Qian, Email: xyqzx@csu.edu.cn.

Supplemental information

Document S1. Figures S1–S13 and Tables S1–S5
mmc1.pdf (11.8MB, pdf)

References

  • 1.Chiang D.H., Jiang Z., Tian C., Wang C.Y. Development and validation of a dynamic early warning system with time-varying machine learning models for predicting hemodynamic instability in critical care: a multicohort study. Crit. Care. 2025;29:318. doi: 10.1186/s13054-025-05553-x. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 2.Hyland S.L., Faltys M., Hüser M., Lyu X., Gumbsch T., Esteban C., Bock C., Horn M., Moor M., Rieck B., et al. Early prediction of circulatory failure in the intensive care unit using machine learning. Nat. Med. 2020;26:364–373. doi: 10.1038/s41591-020-0789-4. [DOI] [PubMed] [Google Scholar]
  • 3.Rahman A., Chang Y., Dong J., Conroy B., Natarajan A., Kinoshita T., Vicario F., Frassica J., Xu-Wilson M. Early prediction of hemodynamic interventions in the intensive care unit using machine learning. Crit. Care. 2021;25:388. doi: 10.1186/s13054-021-03808-x. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 4.Dung-Hung C., Cong T., Zeyu J., Yu-Shan O.Y., Yung-Yan L. External validation of a machine learning model to predict hemodynamic instability in intensive care unit. Crit. Care. 2022;26:215. doi: 10.1186/s13054-022-04088-9. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 5.Møller J.E., Hassager C., Proudfoot A., De Backer D., Morrow D.A., Ravn H.B., Krychtiuk K.A., Zeymer U., Thiele H. Cardiogenic shock: diagnosis, phenotyping and management. Intensive Care Med. 2025;51:1651–1663. doi: 10.1007/s00134-025-08049-y. [DOI] [PubMed] [Google Scholar]
  • 6.Marcucci M., Painter T.W., Conen D., Lomivorotov V., Sessler D.I., Chan M.T.V., Borges F.K., Leslie K., Duceppe E., Martínez-Zapata M.J., et al. Hypotension-Avoidance Versus Hypertension-Avoidance Strategies in Noncardiac Surgery : An International Randomized Controlled Trial. Ann. Intern. Med. 2023;176:605–614. doi: 10.7326/M22-3157. [DOI] [PubMed] [Google Scholar]
  • 7.Bhavani S.V., Semler M., Qian E.T., Verhoef P.A., Robichaux C., Churpek M.M., Coopersmith C.M. Development and validation of novel sepsis subphenotypes using trajectories of vital signs. Intensive Care Med. 2022;48:1582–1592. doi: 10.1007/s00134-022-06890-z. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 8.Li M., Liu F., Yang Y., Lao J., Yin C., Wu Y., Yuan Z., Wei Y., Tang F. Identifying vital sign trajectories to predict 28-day mortality of critically ill elderly patients with acute respiratory distress syndrome. Respir. Res. 2024;25:8. doi: 10.1186/s12931-023-02643-8. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 9.Meng L., Avram D., Tseng G., Huo Z. Outcome-Guided Sparse K-Means for Disease Subtype Discovery via Integrating Phenotypic Data with High-Dimensional Transcriptomic Data. J. Royal Stat. Soc C: Appl Stat. 2022;71:352–375. doi: 10.1111/rssc.12536. [DOI] [Google Scholar]
  • 10.Wu H., Chen Y., Zhu W., Cai Z., Heidari A.A., Chen H. Feature selection in high-dimensional data: an enhanced RIME optimization with information entropy pruning and DBSCAN clustering. Int. J. Mach. Learn. Cybern. 2024;15:4211–4254. doi: 10.1007/s13042-024-02143-1. [DOI] [Google Scholar]
  • 11.Lundberg S.M., Lee S.-I. Proceedings of the 31st International Conference on Neural Information Processing Systems (NIPS’17) Curran Associates Inc., Red Hook; NY, USA: 2017. A unified approach to interpreting model predictions; pp. 4768–4777. [Google Scholar]
  • 12.Lugner M., Gudbjörnsdottir S., Sattar N., Svensson A.M., Miftaraj M., Eeg-Olofsson K., Eliasson B., Franzén S. Comparison between data-driven clusters and models based on clinical features to predict outcomes in type 2 diabetes: nationwide observational study. Diabetologia. 2021;64:1973–1981. doi: 10.1007/s00125-021-05485-5. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 13.Bhavani S.V., Xiong L., Pius A., Semler M., Qian E.T., Verhoef P.A., Robichaux C., Coopersmith C.M., Churpek M.M. Comparison of time series clustering methods for identifying novel subphenotypes of patients with infection. J. Am. Med. Inform. Assoc. 2023;30:1158–1166. doi: 10.1093/jamia/ocad063. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 14.Shi T., Ye X., Sakurai T. IEEE International Conference on Bioinformatics and Biomedicine. BIBM; 2023. Multi-omics Clustering Based on Interpretable and Discriminative Features for Cancer Subtyping. [Google Scholar]
  • 15.Sánchez-Díaz J.S., Peniche-Moguel K.G., Reyes-Ruiz J.M., Del Carpio-Orantes L., Escarramán-Martínez D., Zamarrón-López É.I., Pérez-Nieto O.R., Calyeca-Sánchez M.V. Assessment of mortality risk by respiratory rate-specific quartiles in mechanically ventilated patients with acute respiratory distress syndrome in Mexico. Intern. Emerg. Med. 2025 doi: 10.1007/s11739-025-04112-0. [DOI] [PubMed] [Google Scholar]
  • 16.Candel B.G.J., de Groot B., Nissen S.K., Thijssen W.A.M.H., Lameijer H., Kellett J. The prediction of 24-h mortality by the respiratory rate and oxygenation index compared with National Early Warning Score in emergency department patients: an observational study. Eur. J. Emerg. Med. 2023;30:110–116. doi: 10.1097/MEJ.0000000000000989. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 17.Legouis D., Monard C., Ourahmoune A., Sgardello S., Quintard H., Criton G., Sangla F., Schneider A. Differential effects of thiamine and ascorbic acid in clusters of septic patients identified by latent variable analysis. Crit. Care. 2024;28:396. doi: 10.1186/s13054-024-05188-4. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 18.Choudhary T., Upadhyaya P., Davis C.M., Yang P., Tallowin S., Lisboa F.A., Schobel S.A., Coopersmith C.M., Elster E.A., Buchman T.G., et al. Derivation and validation of generalized sepsis-induced acute respiratory failure phenotypes among critically ill patients: a retrospective study. Crit. Care. 2024;28:321. doi: 10.1186/s13054-024-05061-4. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19.Pensier J., Fosset M., Paschold B.S., von Wedel D., Redaelli S., Braeuer B.L.P., Novack V., Balzer F., Jung B., Amato M.B.P., et al. Temporal stability of phenotypes of acute respiratory distress syndrome: clinical implications for early corticosteroid therapy and mortality. Intensive Care Med. 2025;51:1784–1796. doi: 10.1007/s00134-025-08089-4. [DOI] [PubMed] [Google Scholar]
  • 20.Kiernan E., Zelnick L.R., Khader A., Coston T.D., Bailey Z.A., Speckmaier S., Lo J.J., Siew E.D., Sathe N.A., Kestenbaum B.R., et al. Molecular Phenotyping of Sepsis and Differential Response to Fluid Resuscitation. Am. J. Respir. Crit. Care Med. 2025;211:1681–1688. doi: 10.1164/rccm.202412-2377OC. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 21.Churpek M.M., Snyder A., Twu N.M., Edelson D.P. Accuracy Comparisons between Manual and Automated Respiratory Rate for Detecting Clinical Deterioration in Ward Patients. J. Hosp. Med. 2018;13:486–487. doi: 10.12788/jhm.2914. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22.Johnson A.E.W., Bulgarelli L., Shen L., Gayles A., Shammout A., Horng S., Pollard T.J., Hao S., Moody B., Gow B., et al. MIMIC-IV, a freely accessible electronic health record dataset. Sci. Data. 2023;10:1. doi: 10.1038/s41597-022-01899-x. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 23.Pollard T.J., Johnson A.E.W., Raffa J.D., Celi L.A., Mark R.G., Badawi O. The eICU Collaborative Research Database, a freely available multi-center database for critical care research. Sci. Data. 2018;5 doi: 10.1038/sdata.2018.178. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 24.Huang M.W., Tsai C.F., Tsui S.C., Lin W.C. Combining data discretization and missing value imputation for incomplete medical datasets. PLoS One. 2023;18 doi: 10.1371/journal.pone.0295032. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 25.Bomrah S., Uddin M., Upadhyay U., Komorowski M., Priya J., Dhar E., Hsu S.C., Syed-Abdul S. A scoping review of machine learning for sepsis prediction- feature engineering strategies and model performance: a step towards explainability. Crit. Care. 2024;28:180. doi: 10.1186/s13054-024-04948-6. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 26.Pan W., Hathi D., Xu Z., Zhang Q., Li Y., Wang F. Identification of predictive subphenotypes for clinical outcomes using real world data and machine learning. Nat. Commun. 2025;16:3797. doi: 10.1038/s41467-025-59092-8. [DOI] [PMC free article] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Document S1. Figures S1–S13 and Tables S1–S5
mmc1.pdf (11.8MB, pdf)

Data Availability Statement

  • •

    The dataset of Xiangya used and/or analyzed during this study are not publicly available due to the confidentiality policy of the National Health Commission of China but are available from the corresponding author L.Z. upon reasonable request.

  • •

    This study does not report original code.

  • •

    Any additional information required to reanalyze the data reported in this article is available from the lead contact upon request.


Articles from iScience are provided here courtesy of Elsevier

RESOURCES