Skip to main content
Chest logoLink to Chest
. 2024 Jul 2;166(5):1035–1045. doi: 10.1016/j.chest.2024.06.3769

Development and Validation of the Hospital Medicine Safety Sepsis Initiative Mortality Model

Hallie C Prescott a,b,, Megan Heath a, Elizabeth S Munroe a, John Blamoun c, Paul Bozyk d, Rachel K Hechtman a, Jennifer K Horowitz a, Namita Jayaprakash e, Keith E Kocher b,f, Mariam Younas g, Stephanie P Taylor a, Patricia J Posa a, Elizabeth McLaughlin a, Scott A Flanders a
PMCID: PMC11638544  PMID: 38964673

Abstract

Background

When comparing outcomes after sepsis, it is essential to account for patient case mix to make fair comparisons. We developed a model to assess risk-adjusted 30-day mortality in the Michigan Hospital Medicine Safety sepsis initiative (HMS-Sepsis).

Research Question

Can HMS-Sepsis registry data adequately predict risk of 30-day mortality? Do performance assessments using adjusted vs unadjusted data differ?

Study Design and Methods

Retrospective cohort of community-onset sepsis hospitalizations in the HMS-Sepsis registry (April 2022-September 2023), with split derivation (70%) and validation (30%) cohorts. We fit a risk-adjustment model (HMS-Sepsis mortality model) incorporating acute physiologic, demographic, and baseline health data and assessed model performance using concordance (C) statistics, Brier scores, and comparisons of predicted vs observed mortality by deciles of risk. We compared hospital performance (first quintile, middle quintiles, fifth quintile) using observed vs adjusted mortality to understand the extent to which risk adjustment impacted hospital performance assessment.

Results

Among 17,514 hospitalizations from 66 hospitals during the study period, 12,260 hospitalizations (70%) were used for model derivation and 5,254 hospitalizations (30%) were used for model validation. Thirty-day mortality for the total cohort was 19.4%. The final model included 13 physiologic variables, two physiologic interactions, and 16 demographic and chronic health variables. The most significant variables were age, metastatic solid tumor, temperature, altered mental status, and platelet count. The model C statistic was 0.82 for the derivation cohort, 0.81 for the validation cohort, and ≥ 0.78 for all subgroups assessed. Overall calibration error was 0.0%, and mean calibration error across deciles of risk was 1.5%. Standardized mortality ratios yielded different assessments than observed mortality for 33.9% of hospitals.

Interpretation

The HMS-Sepsis mortality model showed strong discrimination and adequate calibration and reclassified one-third of hospitals to a different performance category from unadjusted mortality. Based on its strong performance, the HMS-Sepsis mortality model may aid in fair hospital benchmarking, assessment of temporal changes, and observational causal inference analysis.

Key Words: benchmarking, health care quality indicator, hospitalization, risk adjustment


Take-home Points.

Study Question: Can the Hospital Medicine Safety sepsis initiative (HMS-Sepsis) registry data adequately predict risk of 30-day mortality after sepsis, and does risk adjustment impact hospital performance assessments?

Results: The final HMS-Sepsis mortality model included 13 physiologic variables, two physiologic interactions, and 16 demographic and chronic health variables. The most significant variables were age, metastatic solid tumor, temperature, altered mental status, and platelet count. The model concordance statistic was 0.82 for the derivation cohort, 0.81 for the validation cohort, and ≥ 0.78 for all subgroups assessed. Adjusted assessments using standardized mortality ratios yielded different quintile assessments than observed mortality for 33.9% of hospitals.

Interpretation: Based on its strong performance, the HMS-Sepsis mortality model may aid in fair hospital in this study benchmarking, assessment of temporal changes, and observational causal inference analysis.

Sepsis, a life-threatening complication of infection, is a leading cause of hospitalization, hospital mortality, and health care costs.1 Recent estimates indicate that sepsis contributes to 1.7 million hospitalizations,1 350,000 deaths,1 and $38 billion2 in direct hospitalization costs annually in the United States. Furthermore, sepsis contributes to one-third to one-half of all hospital deaths,3,4 and as such, is a major driver of overall hospital mortality and quality.

The state of Michigan has a long-standing history of data-driven collaborative quality initiatives (CQIs), funded by Blue Cross Blue Shield of Michigan (BCBSM) since the early 2000s. The CQIs aim to improve outcomes and reduce costs of health care throughout the state.5,6 They have been successful at both goals,7 yielding a large return on investment.5 Recognizing the burden of sepsis, BCBSM began funding the CQI the Hospital Medicine Safety sepsis initiative (HMS-Sepsis) in 2020. Participating hospitals submit data to the HMS-Sepsis registry, receive real-time performance data, implement local quality improvement initiatives facilitated by the HMS-Sepsis coordinating center, and share successes and challenges at collaborative-wide meetings to promote improvement across all hospitals. Additionally, hospitals are scored on an annual performance index that impacts BCBSM hospital reimbursement under a pay-for-performance model. Sepsis process measures were added to the HMS performance index in 2024.

Because of the wide heterogeneity of sepsis, risk adjustment is necessary to yield fair comparisons of outcomes across hospitals and over time. We sought to develop a risk-adjustment model (herein referred to as the HMS-Sepsis mortality model) to account for differences in patient case mix and illness severity. The goals of the model were to facilitate cross-sectional benchmarking of hospitals, to measure temporal changes in outcomes, and to support causal inference analyses using HMS-Sepsis registry data. Rather than use an existing mortality model, we developed a new model to leverage best the unique human-abstracted HMS-Sepsis registry data, which differ from administrative and electronic health record data. Additionally, we wanted to ensure that the model was optimally tuned to sepsis (as opposed to all-cause admissions) and to build it using data from the first 6 hours of presentation to avoid adjusting away illness severity resulting from delayed initial management of sepsis. Herein, we describe the development and internal validation of the HMS-Sepsis mortality model.

Study Design and Methods

Setting and Data Source

The Michigan Hospital Medicine Safety (HMS) CQI was founded in 2010 to improve the quality of care for hospitalized medical patients. Participating hospitals range from larger urban academic medical centers to smaller rural hospitals. More information on the HMS is provided in e-Appendix 1.

For the HMS-Sepsis, a random sample of adult hospitalizations for community-onset sepsis are selected for inclusion into the HMS-Sepsis registry at each hospital. Hospitalizations are identified via a two-step process. First, hospitalizations with principal diagnostic coding for sepsis or infection are identified, then professional abstractors review the electronic health record to confirm presence of infection and acute organ dysfunction during the first 2 calendar days of the encounter. Specifically, abstractors assess for lactate elevation, acute respiratory dysfunction, acutely altered mental status, acute renal dysfunction, acute hematologic dysfunction, acute liver dysfunction, and vasopressor treatment, similar to the Centers for Disease Control and Prevention’s Adult Sepsis Event criteria.8 This two-step process was designed to have a consistent threshold for inclusion, despite variable use of sepsis diagnostic codes over time and across hospitals.9 More details on HMS-Sepsis inclusion criteria, exclusion criteria, and sampling strategy are provided in e-Appendix 2.

Data on eligible hospitalizations are entered into the centralized HMS-Sepsis registry by professional abstractors at each hospital using structured abstraction forms and a standardized data dictionary. Vital status at 90 days after discharge is determined by chart review, search of public obituaries, and a telephone call to the patient. To ensure robust data collection, abstractors undergo training by the HMS-Sepsis coordinating center and abstractions are audited for accuracy. For this analysis, we included 18 months of hospitalizations (April 2022 through September 2023). Hospitalizations before April 2022 were excluded, given the high prevalence of COVID-19 unlikely to be representative of future case-mix and outcomes.

HMS-Sepsis Mortality Model Approach

The model for predicting 30-day mortality among hospitalizations for sepsis was developed and validated according to a prespecified statistical analysis plan developed by two of the authors (H. C. P. and M. H.). The analysis plan followed the approach used to develop the United Kingdom’s Intensive Care National Audit and Research Centre mortality model10 and adhered to principles laid out in the Centers for Medicare and Medicaid Services’ Risk Adjustment in Quality Measurement.11 The study cohort was divided into a derivation cohort (70%) for model development and a validation cohort (30%) for model assessment. We split hospitalizations at random, rather than by time or by hospital, because we had fewer than 2 calendar years of data (making it infeasible to split by year) and wanted data from every hospital to be eligible for both the development and validation cohorts (making it not feasible to split by hospitals).

We developed the HMS-Sepsis mortality model based on three guiding principles. First, the primary goal of the model is to track HMS-Sepsis quality improvement efforts. To ensure that the model was suited optimally to this goal, we developed and internally validated a novel risk-adjustment model, rather than using or adapting an off-the-shelf model. This allowed us to leverage the unique human-abstracted data elements available in the HMS-Sepsis registry (eg, baseline functional limitations, acute alteration of mental status) and to ensure the model was calibrated optimally to sepsis and our hospitals. In general, purpose-built models perform better than off-the-shelf models.12,13 Second, we aimed to create an interpretable model. Following the approach taken by Intensive Care National Audit and Research Centre, we considered only variables known to be associated with mortality based on prior literature or clinical experience and retained only variables that meaningfully improved model performance.10 Third, we aimed to generate a model accounting for illness severity on presentation and not for subsequent illness severity that might have resulted from delayed recognition or treatment of sepsis.11 Thus, the model primarily used data acquired within 6 hours of presentation to the health care system. However, because some data elements are collected daily and not hourly, data beyond 6 hours were permitted for some variables, as detailed in e-Appendix 3.

HMS-Sepsis Mortality Model Development

The HMS-Sepsis mortality model was built in a multistep process using the derivation cohort. Before model development, the optimal functional form for each candidate variable was determined using a SAS macro.14 Fit statistics for linear, categorical, squared, spline, and log forms were compared for each continuous variable. If the optimal functional form was a spline term, three to five nodes were considered. As the first step, a parsimonious model using acute physiologic variables was fit. Second, demographic and chronic health variables were added, and variables that worsened or did not improve model fit significantly were removed via backward selection. In this process, we determined the change in concordance (C) statistic, Akaike information criterion, Bayesian information criterion, and Brier score that would occur with the removal of each individual variable, and we removed variables that resulted in improved Akaike information criterion with removal. Third, we added select a priori physiology-physiology and chronic health-physiology interaction terms (eg, invasive mechanical ventilation × systolic BP, chronic kidney disease × creatinine) deemed potentially important based on clinical experience and prior literature10 and again removed variables that worsened or did not improve model fit via backward selection. Fourth, we used bootstrap resampling to evaluate Wald P values for each predictor and removed variables that were not significant consistently across bootstrap estimations. This step ensured that rare characteristics were not included in the model. The full list of candidate variables considered is presented in e-Appendix 3. Missing values were imputed as normal under the assumption that the data were not collected because they were assumed to be normal, as is conventional in risk adjustment.10,15,16

HMS-Sepsis Mortality Model Validation

After the final model was fit in the derivation cohort, we evaluated model performance in the validation cohort. We assessed model discrimination using C statistics; overall model performance using Brier scores17; and model calibration using plots of observed vs predicted mortality by decile of predicted mortality, as well as maximum and average calibration error across deciles of risk.18 We considered model discrimination to be strong when the C statistic was > 0.8, consistent with standard practice.16,19 We considered the model to be useful when Brier score was < 0.25, consistent with prior work.20 Although no standard threshold exists for grading model calibration, we considered overall and mean calibration errors of < 1% to reflect excellent model calibration.16,21,22

We completed several additional analyses to examine model performance, some of which were added post hoc. First, we assessed model performance in a priori subgroups defined by baseline health status,23 age, sex, and positive COVID-19 testing status and post hoc subgroups defined by site of infection to assess for differential performance across subgroups. Second, to understand the impact of risk adjustment on hospital performance assessment, we compared hospital performance categories (quintile 1, quintiles 2-4, and quintile 5) defined by standardized mortality ratio (SMR; calculated as observed divided by predicted mortality) vs unadjusted mortality, similar to prior work.24 This comparison was limited to hospitals with > 25 hospitalizations. Third, we assessed the performance of a parsimonious version of the HMS-Sepsis mortality model that excluded data elements not readily extracted from structured electronic health record data (post hoc). Fourth, to understand the impact of each variable included in the model, we examined model performance in the derivation cohort using progressively simpler models created by serial removal of the least predictive variable (post hoc).

Analyses were completed in SAS version 9.4 software (SAS Institute). We considered P < .05 to indicate statistical significance. We followed the Transparent Reporting of a Multivariable Prediction Model for Individual Prognosis or Diagnosis guidance.25 Because HMS-Sepsis is a quality improvement initiative, this study was deemed not regulated by the University of Michigan Institutional Review Board (Identifier: HUM-00179611).

Results

During the study period, 17,514 hospitalizations were in the HMS-Sepsis registry from 66 hospitals, of which 12,260 hospitalizations (70%) were allocated for model derivation and 5,254 hospitalizations (30%) were allocated for model validation. Cohort characteristics are presented in Table 1 and e-Table 1. The overall cohort was a median of 71 years of age (interquartile range, 61-81 years) and 50.0% male; 39.8% had baseline functional impairment, 19.5% had baseline cognitive impairment, 7.1% had metastatic solid malignancy, 33.2% had been hospitalized in the prior 90 days, and 14.8% were admitted from a facility; and 8.4% had positive results for COVID-19. Further, 10.6% were treated with vasopressors and 6.6% were treated with invasive mechanical ventilation within 6 hours of hospital presentation. Thirty-day mortality was 19.4% (3,403/17,514), of which 60.2% of deaths (2,049/3,403) occurred in the hospital and 39.8% of deaths (1,353/3,402) occurred after discharge. Characteristics were similar between the derivation and validation cohorts, with standardized mean differences of < 0.05 for all characteristics.

Table 1.

Characteristics of the Derivation, Validation, and Total Cohorts

Characteristic Derivation (n = 12,260) Validation (n = 5,254) Total (n = 17,514) Standardized Mean Difference
Age, y 71 (61-80) 71 (61-81) 71 (61-81) –0.003
Sex
 Female 6,066 (49.5) 2,596 (49.4) 8,662 (49.5) 0.001
 Male 6,133 (50.0) 2,629 (50.0) 8,762 (50.0) 0.000
 Unknown 61 (0.5%) 29 (0.6) 90 (0.5) –0.007
BMI, kg/m2 27.7 (23.4-33.3) 27.7 (23.0-33.1) 27.7 (23.2-33.3) 0.014
Hospitalized in prior 90 d 4,030 (32.9) 1,776 (33.8) 5,806 (33.2) –0.020
Admitted from SNF, SAR, or LTAC 1,814 (14.8) 778 (14.8) 2,592 (14.8) 0.000
Chronic health conditions
 Baseline cognitive impairment 2,353 (19.2) 1,054 (20.1) 3,407 (19.5) –0.022
 Baseline functional impairment 4,905 (40.0) 2,074 (39.5) 6,979 (39.8) 0.011
 Cardiovascular disease 5,060 (41.3) 2,219 (42.2) 7,279 (41.6) –0.019
 Cerebrovascular disease 1,955 (15.9) 881 (16.8) 2,836 (16.2) –0.022
 Chronic pulmonary disease 3,894 (31.8) 1,632 (31.1) 5,526 (31.6) 0.015
 Dementia 1,654 (13.5) 722 (13.7) 2,376 (13.6) –0.007
 Diabetes 4,842 (39.5) 2,098 (39.9) 6,940 (39.6) –0.009
 Leukemia or lymphoma 369 (3.0) 152 (2.9) 521 (3.0) 0.007
 Moderate or severe kidney disease 4,588 (37.4) 1,969 (37.5) 6,557 (37.4) –0.001
 Solid malignancy, no metastasis 1,826 (14.9) 773 (14.7) 2,599 (14.8) 0.005
 Metastatic solid tumor 873 (7.1) 365 (6.9) 1,238 (7.1) 0.007
 Moderate or severe liver disease 365 (3.0) 159 (3.0) 524 (3.0) –0.003
Admission physiologic features and treatments
 Maximum creatinine, mg/dL 1.3 (0.9-1.9) 1.3 (0.9-2.0) 1.3 (0.9-2.0) –0.006
 Maximum lactate, mg/dL 2.1 (1.2-3.2) 2.2 (1.3-3.3) 2.1(1.2-3.2) –0.006
 Minimum Pao2 to Fio2 ratio, mm Hg 310 (203-476) 310 (203-476) 310 (203-476) –0.001
 Acutely altered mental status 5,461 (44.5) 2,300 (43.8) 7,761 (44.3) 0.015
 Positive COVID-19 testing 1,049 (8.6) 420 (8.0) 1,469 (8.4) 0.021
 Vasopressors within 6 h 1,341 (10.9) 516 (9.8) 1,857 (10.6) 0.037
 Mechanical ventilation within 6 h 814 (6.6) 346 (6.6) 1,160 (6.6) 0.002

Data are presented as No. (%) or median (interquartile range) unless otherwise indicated. Baseline functional impairment was defined as dependency on others for completion of any of the following activities or instrumental activities of daily living: bathing, dressing, toileting, transferring between bed and chair, walking, and managing medications. The Hospital Medicine Safety sepsis initiative registry includes daily minimum and maximum laboratory values and hourly minimum and maximum vital signs for hours 1 through 3. We used laboratory values from encounter day 1 in the model. However, if a patient sought treatment at the hospital after 6:00 pm and had no laboratory values measured on encounter day 1, then values from day 2 were used. Otherwise, if no values were available during the time frame, then normal values were imputed. LTAC = long-term acute care facility; SAR = subacute rehabilitation; SNF = skilled nursing facility.

HMS-Sepsis Mortality Model Development

All physiologic variables considered contributed to model fit and were retained. Data completeness for physiologic variables is presented in e-Table 2. Missingness was greatest for bilirubin and lactate, which were imputed in 14.2% and 13.6% of hospitalizations, respectively. Demographics and chronic health variables were added subsequently, with the following removed via the backward selection process because they did not improve model fit: sex, cerebrovascular disease, connective tissue disorders, transplant status, HIV and AIDS, chronic pulmonary disease, thromboembolic disease, dementia, peptic ulcer disease, and smoking status. Physiology-physiology and chronic health-physiology interactions were considered subsequently. Two interaction terms improved model fit and were retained: COVID-19 positivity × invasive mechanical ventilation and chronic kidney disease × creatinine. Five additional variables were removed in the final step of model assessment via bootstrap resampling because they were not consistently significant across bootstrap estimations: asthma, inflammatory bowel disease, solid malignancy without metastasis, diabetes, and hemiplegia or paraplegia. The variables included in the final model are presented in Table 2. The five most significant variables by Wald χ2 test were age, metastatic solid cancer, temperature, altered mental status, and platelet count.

Table 2.

Variables Included in Final HMS-Sepsis Mortality Model

Variable Measurement Time Frame Functional Form df Wald χ2 P Value
Physiologic characteristics
 Maximum temperature First 3 h Categorical 6 114.81 < .0001
 Altered mental status Day 1 Dichotomous 1 106.87 < .0001
 Minimum platelet Day 1 Spline (5 knots) 4 106.80 < .0001
 Minimum Pao2 to Fio2 ratio First 3 h Spline (4 knots) 3 82.01 < .0001
 Maximum creatinine Day 1 Spline (5 knots) 4 60.05 < .0001
 Maximum lactate Day 1 Spline (5 knots) 4 59.25 < .0001
 Mechanical ventilation First 6 h Dichotomous 1 44.32 < .0001
 Maximum respiratory rate First 3 h Categorical 4 23.57 < .0001
 Positive COVID-19 findings Days 1-2 Dichotomous 1 16.87 < .0001
 Vasopressor treatment First 6 h Dichotomous 1 15.38 < .0001
 Maximum heart rate First 3 h Categorical 4 13.18 .0104
 Maximum bilirubin Day 1 Squared 1 11.28 .0008
 Minimum SBP First 3 h Categorical 2 10.28 .0058
Demographics and chronic health characteristics
 Age Day 1 Spline (4 knots) 3 272.03 < .0001
 Metastatic solid tumor Day 1 Dichotomous 1 246.36 < .0001
 BMI Day 1 Spline (5 knots) 4 71.39 < .0001
 Hospitalization in prior 90 d Day 1 Dichotomous 1 44.72 < .0001
 Admission from SNF, SAR, or LTAC Day 1 Dichotomous 1 34.46 < .0001
 Baseline functional limitations Day 1 Categorical 6 24.06 .0005
 Moderate or severe liver disease Day 1 Dichotomous 1 26.15 < .0001
 Cognitive impairment Day 1 Dichotomous 1 15.54 < .0001
 Hematologic malignancy Day 1 Dichotomous 1 10.48 .0012
 Atrial fibrillation Day 1 Dichotomous 1 10.24 .0014
 Hypertension Day 1 Dichotomous 1 10.17 .0014
 Congestive heart failure Day 1 Dichotomous 1 5.34 .0208
 Cardiovascular disease Day 1 Dichotomous 1 6.45 .0111
 Peripheral vascular disorders Day 1 Dichotomous 1 4.83 .0279
 Mild liver disease Day 1 Dichotomous 1 4.75 .0292
 Moderate or severe kidney disease Day 1 Dichotomous 1 4.26 .0391
Physiologic interactions
 Kidney disease × creatinine N/A Dichotomous 4 16.43 .0025
 Positive COVID-19 findings × IMV N/A Dichotomous 1 4.30 .0381

Baseline functional impairment was defined as dependency on others for completion of any of the following activities or instrumental activities of daily living: bathing, dressing, toileting, transferring between bed and chair, walking, and managing medications. The HMS-Sepsis registry includes daily minimum and maximum laboratory values and hourly minimum and maximum vital signs for hours 1 through 3. We used laboratory values from encounter day 1 in the model. However, if a patient sought treatment at the hospital after 6:00 pm and had no laboratory values measured on encounter day 1, then values from day 2 were used. Otherwise, if no values were available during the time frame, then normal values were imputed. Variable are presented by category (physiologic, demographic and chronic health, and interaction terms), ordered by Wald χ2 value from most to least significant. df = degrees of freedom; HMS-Sepsis = Hospital Medicine Safety sepsis initiative; IMV = invasive mechanical ventilation; LTAC = long-term acute care facility; N/A = not applicable; SAR = subacute rehabilitation; SNF = skilled nursing facility.

HMS-Sepsis Mortality Model Validation

Predicted mortality for the overall cohort was a median of 12.4% (interquartile range, 5.4%-26.7%) and a mean of 19.3%. The distribution of predicted mortality is presented in e-Figure 1. Model performance is presented in Table 3. The C statistic was 0.82 for the derivation cohort and 0.81 for the validation cohort, indicating strong model discrimination. The C statistic was ≥ 0.78 for all subgroups assessed. The Brier score was 0.121 for the derivation cohort, 0.124 for the validation cohort, and ≤ 0.153 for all subgroups assessed. Observed vs predicted mortality by decile of risk is presented in e-Figure 2 and e-Table 3. Overall calibration error was 0.0%, mean calibration error was 1.5%, and maximum calibration error was 3.0%.

Table 3.

Model Performance in the Derivation Cohort, Validation Cohort, and Subgroups

Cohort No. 30-d Mortality C Statistic Brier Score
All observations
 Derivation 12,260 2,387 (19.5) 0.821 0.121
 Validation 5,254 1,015 (19.3) 0.807 0.124
Previously healthy
 Derivation 1,067 51 (4.8) 0.906 0.034
 Validation 451 19 (4.2) 0.839 0.038
Comorbid
 Derivation 11,193 2,336 (20.9) 0.807 0.129
 Validation 4,803 996 (20.7) 0.796 0.132
Age, y
 ≤ 70
 Derivation 5,808 805 (13.9) 0.843 0.091
 Validation 2,503 322 (12.9) 0.817 0.093
 ≥ 71
 Derivation 6,452 1,582 (24.5) 0.787 0.148
 Validation 2,751 693 (25.2) 0.778 0.153
Sex
 Male
 Derivation 6,133 1,199 (19.5) 0.808 0.126
 Validation 2,629 501 (19.1) 0.802 0.125
 Female
 Derivation 6,066 1,178 (19.4) 0.832 0.116
 Validation 2,596 506 (19.5) 0.813 0.123
COVID-19 status
 Positive
 Derivation 1,049 260 (24.8) 0.819 0.140
 Validation 420 95 (22.6) 0.783 0.142
 Negative
 Derivation 11,211 2,127 (19.0) 0.820 0.119
 Validation 4,834 920 (19.0) 0.809 0.123
Respiratory infection
 Derivation 6,083 1,271 (20.9) 0.814 0.127
 Validation 2,574 543 (21.1) 0.802 0.132
Genitourinary infection
 Derivation 3,797 670 (17.6) 0.814 0.115
 Validation 1,669 279 (16.7) 0.791 0.120
Skin and soft tissue
 Derivation 1,621 261 (16.1) 0.823 0.107
 Validation 700 106 (15.1) 0.810 0.102
GI infection
 Derivation 1,413 275 (19.5) 0.797 0.128
 Validation 553 106 (19.2) 0.790 0.134

Data are presented as No. (%) unless otherwise indicated. Patients were defined as previously healthy (vs comorbid) based on having no major comorbidities and few minor comorbidities or other evidence of health impairment.23 Patients were classified as having respiratory, genitourinary, skin or soft tissue, and gastrointestinal (GI) infection based on hospital discharge codes mapped to Agency for Healthcare Research and Quality’s Clinical Classification Software Revised categories.26

When comparing performance categories defined by HMS-Sepsis mortality model SMR vs observed mortality, 39 of 59 hospitals (66.1%) with > 25 hospitalizations were classified concordantly as quintile 1, quintiles 2 through 4, or quintile 5 by both metrics, whereas 20 of 59 hospitals (33.9%) showed discrepant classifications (Table 4, Fig 1). Ten hospitals were classified as having lower mortality by SMR (including five hospitals in the lowest quintile of SMR), whereas 10 hospitals were classified as having higher mortality by SMR (including five hospitals in the highest quintile of SMR).

Table 4.

Comparison of Hospital Performance by Quintiles of Unadjusted Mortality vs SMR

Variable Quintiles Defined by SMR
Quintile 1: Lowest SMR Quintiles 2-4 Quintile 5: Highest SMR
Quintile 1: lowest unadjusted mortality 6 (10.2)a 5 (8.5)b 0 (0.0)
Quintiles 2-4 5 (8.5)c 26 (44.1)a 5 (8.5)b
Quintile 5: highest unadjusted mortality 0 (0.0) 5 (8.5)c 7 (11.9)a

Data are presented as No. (%). This table shows that mortality quintiles defined by unadjusted mortality (row) vs SMR (column) yielded concordant assessments for 39 hospitals (66.1%) and discrepant assignments for 20 hospitals (33.9%). Five hospitals in the highest quintile of observed mortality had an SMR in quintiles 2 through 4, whereas five hospitals in quintiles 2 through 4 of observed mortality were in the highest quintile of SMR. SMR = standardized mortality ratio.

a

Consistent performance assessments.

b

Discrepant assessments, SMR performance worse.

c

Discrepant assessment, SMR performance better.

Figure 1.

Figure 1

Graph showing a comparison of unadjusted 30-day mortality vs standardized mortality ratio (SMR), with the SMR (dot) and 95% CI (bar) for each of 59 hospitals with > 25 hospitalizations. Hospitals are ordered from lowest SMR to highest SMR, with quintiles of SMR depicted by the shaded columns. Hospitals in the lowest quintile of unadjusted mortality are shown in gray, whereas hospitals with the highest quintile of unadjusted mortality are shown in red.

The parsimonious model (omitting baseline functional limitation, acute alteration of mental status, and minimum Pao2 to Fio2 ratio) showed a C statistic of 0.80 in the validation cohort. Across subgroups, the C statistic was 0.007 to 0.044 higher for the full HMS-Sepsis mortality model vs the parsimonious version (e-Table 4). When comparing performance categories defined by the HMS-Sepsis mortality model SMR vs the parsimonious model SMR, 47 hospitals (79.7%) had concordant assessments and 12 hospitals (20.3%) had discordant assessments (e-Table 5). Furthermore, examination of progressively simpler models demonstrated degraded performance of simplified models (e-Table 6).

Discussion

In this study of > 17,000 adult hospitalizations for community-onset sepsis at 66 diverse hospitals in Michigan, we developed and internally validated the HMS-Sepsis mortality model. Where possible, the model uses data from just the first 6 hours of presentation to predict 30-day mortality. The model showed strong discrimination, with a C statistic of 0.81 in the validation cohort. It also demonstrated good calibration, with an overall calibration error of 0.0% and an average calibration error of 1.5% across deciles of predicted risk. Importantly, for one-third of hospitals, the performance assessment by the HMS-Sepsis SMR differed from the unadjusted assessment of hospital mortality, showing the necessity of risk adjustment when comparing mortality across hospitals, even among condition-specific cohorts. The strong performance of the HMS-Sepsis mortality model suggests that it can be used to support fair hospital benchmarking, unbiased assessment of temporal changes in sepsis mortality, and robust observational causal inference analyses.

The performance of the HMS-Sepsis mortality model should be considered in the context of specific design features. Many risk scores, such as the Acute Physiology and Chronic Health Evaluation score,27 Intensive Care National Audit and Research Centre risk models,10,13 and Kaiser Permanente Northern California risk models,15,28 demonstrate stronger discrimination, with C statistics of > 0.85. However, these models evaluate risk in more heterogenous populations of all-cause hospital admissions or ICU admissions. Each of these models include reason for admission within the model, which contributes significantly to model performance. By contrast, the HMS-Sepsis mortality model focuses exclusively on hospitalizations for community-onset sepsis, thereby eliminating an important source of variability in patient risk. The HMS-Sepsis mortality model performs as well or better than most prior sepsis-specific risk models.29, 30, 31, 32 Second, most risk models use a 24-h period to assess patient illness severity. This longer observation window allows for greater data availability (ie, less missingness) and inclusion of data elements at later periods (ie, closer to the outcome). For this reason, models with longer periods display better performance, but have the potential drawback of adjusting away illness severity that may have developed as a result of suboptimal initial management.11 The first 6 hours of sepsis management are critical and are a focus of the HMS-Sepsis initiative, so where possible, we limited the model to data elements within the first 6 h, although this may decrease measured model performance.

A second major finding of the study is the relative importance of the various components of the HMS-Sepsis mortality model, some of which are not included routinely in other models. Altered mental status, for example, is captured in the HMS-Sepsis registry through abstractor review of clinical documentation on presentation to the ED and hospital. However, this information is available rarely in structured data, so is not included in the Centers for Disease Control and Prevention’s adult sepsis event definition criteria or most risk-adjustment models.8 However, altered mental status has been associated with poor outcomes in other studies of patients with infection, sepsis, or both where data are available. For example, altered mental status is one of just three items prioritized for inclusion in the qSOFA (quick sequential [sepsis-associated] organ failure assessment score) risk-stratification tool for patients with infection33 and is the acute organ dysfunction most strongly associated with mortality after hospitalization in sepsis.34 Baseline functional limitations likewise are captured in the HMS-Sepsis registry from review of clinical documentation and contribute significantly to the model. Finally, clinical documentation of cognitive impairment outperformed diagnosis of dementia, which was removed from the model in the backward selection process.

However, although human-abstracted registry data such as altered mental status and baseline functional limitations are informative, they also are labor intensive and costly to collect. Thus, despite the value of such data, they are not feasible to collect in many performance assessment scenarios. To understand the necessity of human-abstracted data to performance of the HMS-Sepsis mortality model, we fit a parsimonious version of the HMS-Sepsis model omitting baseline functional limitation, acute alteration of mental status, and minimum Pao2 to Fio2 ratio. This parsimonious model performed well, but was not as strong as the full HMS-Sepsis mortality model and yielded discrepant performance assessments from the full model for 20% of hospitals.

Our study should be interpreted in the context of several limitations. First, the HMS-Sepsis mortality model uses data collected in the HMS-Sepsis registry, which may not be available in other settings. To address this limitation, we assessed a parsimonious model that may be implemented more readily outside HMS. Second, additional factors influence patient risk that were not included in the HMS-Sepsis model because they were unavailable or not captured within 6 h, such as site of infection.35, 36, 37 However, model performance consistently was strong across subgroups defined by site of infection coded at discharge. Third, we split our data into 70% derivation and 30% validation cohorts, but cross-validation may yield a more stable assessment of model performance with smaller datasets.38 We mitigated this issue by using bootstrapping in the model development process and assessing model performance across multiple subgroups. Fourth, although the model performed well in internal validation, we did not perform external validation of the model, which would be warranted for use of the HMS-Sepsis model in other settings. Fifth, we ascertained vital status by chart review, public obituary data, and telephone follow-up, rather than linkage to national death index data. Our approach allows for real-time ascertainment of vital status to support timely benchmarking and has been used successfully in prior research.39 Sixth, we are unable to assess how the HMS-Sepsis mortality model compares with other existing models.

Interpretation

In this study of > 17,000 recent adult hospitalizations for community-onset sepsis in the state of Michigan, we developed and validated the HMS-Sepsis mortality model. The HMS-Sepsis mortality model demonstrated strong discrimination and adequate calibration and reclassified one-third of hospitals to a different performance category from unadjusted mortality. Based on its strong performance, the HMS-Sepsis mortality model can aid in fair hospital benchmarking, assessment of temporal changes in sepsis mortality, and observational causal inference analysis.

Funding/Support

This work was supported by Blue Cross Blue Shield of Michigan as part of their Value Partnerships program, as well as with resources and use of facilities at the Ann Arbor VA Medical Center. E. S. M. was supported by the National Heart, Lung, and Blood Institute, National Institutes of Health [Grants T32HL007749and F32HL172463]. R. K. H. was supported by National Heart, Lung, and Blood Institute, National Institutes of Health [Grant T32HL007749].

Financial/Nonfinancial Disclosures

The authors have reported to CHEST the following: H. C. P., M. H., J. K. H., P. J. P., E. M., and S. A. F.) receive salary support from BCBSM for work on HMS. K. K. has salary support from BSBSM for work on the Michigan Emergency Department Improvement Collaborative, an ED-based quality network. The authors report grant funding and salary support unrelated to this study from the National Institutes of Health (H. C. P., N. J., S. P. T.), the Agency for Healthcare Research and Quality (H. C. P., S. P. T.), Department of Veterans Affairs (H. C. P.), and Centers for Disease Control and Prevention (H. C. P., J. K. H.). H. C. P. and J. N. serve on the advisory board to Sepsis Alliance (unpaid). H. C. P. serves on the Surviving Sepsis Campaign guidelines. None declared (J. B., P. B., M. Y.).

Acknowledgments

Author contributions: H. C. P. and M. H. designed the analysis. M. H. completed the analysis. H. C. P. drafted the manuscript. All authors interpreted the data, reviewed the manuscript critically for important intellectual content, approved the final version to be published, and are accountable for all aspects of the work.

Role ofsponsors: The sponsor had no role in the design of the study, the collection and analysis of the data, or the preparation of the manuscript.

Disclaimers: Although Blue Cross Blue Shield of Michigan and HMS work collaboratively, the opinions, beliefs, and viewpoints expressed by the authors do not necessarily reflect the opinions, beliefs, and viewpoints of BCBSM or any of its employees. This manuscript does not necessarily reflect the position of the Department of Veterans Affairs or the US government.

Other contributions: The authors thank the Michigan Hospital Medicine Safety (HMS) Consortium Data, Design, and Publications Committee and the HMS Critical Care steering committee for providing feedback on this study and all HMS hospitals for contributing to data collection.

Additional information: The e-Appendixes, e-Figures, and e-Tables are available online under “Supplementary Data.”

Supplementary Data

e-Online Data
mmc1.docx (505.1KB, docx)

References

  • 1.Rhee C., Dantes R., Epstein L., et al. Incidence and trends of sepsis in US hospitals using clinical vs claims data, 2009-2014. JAMA. 2017;318(13):1241–1249. doi: 10.1001/jama.2017.13836. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 2.Liang L (AHRQ), Moore B (IBM Watson Health), Soni A (AHRQ). National inpatient hospital costs: the most expensive conditions by payer, 2017. HCUP Statistical Brief no. 261. Month 2020. Agency for Healthcare Research and Quality website. Accessed February 21, 2024. www.hcup-us.ahrq.gov/reports/statbriefs/sb261-Most-Expensive-Hospital-Conditions-2017.pdf [PubMed]
  • 3.Rhee C., Jones T.M., Hamad Y., et al. Prevalence, underlying causes, and preventability of sepsis-associated mortality in US acute care hospitals. JAMA Netw Open. 2019;2(2) doi: 10.1001/jamanetworkopen.2018.7571. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 4.Liu V., Escobar G.J., Greene J.D., et al. Hospital deaths in patients with sepsis from 2 independent cohorts. JAMA. 2014;312(1):90–92. doi: 10.1001/jama.2014.5804. [DOI] [PubMed] [Google Scholar]
  • 5.Blue Cross Blue Shield of Michigan Blue Cross Blue Shield of Michigan’s collaborative quality initiatives. Blue Cross Blue Shield of Michigan website. https://www.valuepartnerships.com/programs/collaborative-quality-initiatives/
  • 6.McGowan J.G., Martin G.P., Krapohl G.L., et al. What are the features of high-performing quality improvement collaboratives? A qualitative case study of a state-wide collaboratives programme. BMJ Open. 2023;13(12) doi: 10.1136/bmjopen-2023-076648. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 7.Share D.A., Campbell D.A., Birkmeyer N., et al. How a regional collaborative of hospitals and physicians in Michigan cut costs and improved the quality of care. Health Aff. 2011;30(4):636–645. doi: 10.1377/hlthaff.2010.0526. [DOI] [PubMed] [Google Scholar]
  • 8.Centers for Disease Control and Prevention Hospital toolkit for adult sepsis surveillance. Centers for Disease Control and Prevention website. https://www.cdc.gov/sepsis/pdfs/Sepsis-Surveillance-Toolkit-Mar-2018_508.pdf
  • 9.Lindenauer P.K., Lagu T., Shieh M.S., Pekow P.S., Rothberg M.B. Association of diagnostic coding with trends in hospitalizations and mortality of patients with pneumonia, 2003-2009. JAMA. 2012;307(13):1405–1413. doi: 10.1001/jama.2012.384. [DOI] [PubMed] [Google Scholar]
  • 10.Harrison D.A., Ferrando-Vivas P., Shahin J., Rowan K.M. NIHR Journals Library; Southampton, UK: 2015. Ensuring Comparisons of Health-care Providers Are Fair: Development and Validation of Risk Prediction Models for Critically Ill Patients. [PubMed] [Google Scholar]
  • 11.Centers for Medicare and Medicaid Services Risk adjustment in quality measurement. Centers for Medicare and Medicaid Services website. https://mmshub.cms.gov/sites/default/files/Risk-Adjustment-in-Quality-Measurement.pdf
  • 12.Soares M., Dongelmans D.A. Why should we not use APACHE II for performance measurement and benchmarking? Rev Bras Ter Intensiva. 2017;29(3):268–270. doi: 10.5935/0103-507X.20170043. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 13.Harrison D.A., Parry G.J., Carpenter J.R., Short A., Rowan K. A new risk prediction model for critical care: the Intensive Care National Audit & Research Centre (ICNARC) model. Crit Care Med. 2007;35(4):1091–1098. doi: 10.1097/01.CCM.0000259468.24532.44. [DOI] [PubMed] [Google Scholar]
  • 14.Liu S., Stedman M.R. A SAS® macro for covariate specification in linear, logistic and survival regression [paper 1223-2017]. Proceedings of the SAS Global 2017 Conference, Orlando FL. SAS Institute, Inc., website. http://support.sas.com/resources/papers/proceedings17/1223-2017.pdf
  • 15.Escobar G.J., Greene J.D., Scheirer P., Gardner M.N., Draper D., Kipnis P. Risk-adjusting hospital inpatient mortality using automated inpatient, outpatient, and laboratory databases. Med Care. 2008;46(3):232–239. doi: 10.1097/MLR.0b013e3181589bb6. [DOI] [PubMed] [Google Scholar]
  • 16.Prescott H.C., Kadel R.P., Eyman J.R., et al. Risk-adjusting mortality in the nationwide Veterans Affairs healthcare system. J Gen Intern Med. 2022;37(15):3877–3884. doi: 10.1007/s11606-021-07377-1. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 17.Brier G.W. Verification of forecasts expressed in terms of probability. Monthly Weather Review. 1950;75:1–3. [Google Scholar]
  • 18.Huang Y., Li W., Macheret F., Gabriel R.A., Ohno-Machado L. A tutorial on calibration measurements and calibration models for clinical prediction models. J Am Med Inform Assoc. 2020;27(4):621–633. doi: 10.1093/jamia/ocz228. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19.Hosmer D.W., Lemeshow S. Wiley; 1989. Applied Logistic Regression. [Google Scholar]
  • 20.Gerds T.A., Cai T., Schumacher M. The performance of risk prediction models. Biom J. 2008;50(4):457–479. doi: 10.1002/bimj.200810443. [DOI] [PubMed] [Google Scholar]
  • 21.Altman D.G., Vergouwe Y., Royston P., Moons K.G. Prognosis and prognostic research: validating a prognostic model. BMJ. 2009;338 doi: 10.1136/bmj.b605. [DOI] [PubMed] [Google Scholar]
  • 22.Royston P., Moons K.G., Altman D.G., Vergouwe Y. Prognosis and prognostic research: developing a prognostic model. BMJ. 2009;338 doi: 10.1136/bmj.b604. [DOI] [PubMed] [Google Scholar]
  • 23.Munroe E., Basu T., O’Malley M.E., et al. Identfying potentially preventable death from sepsis. Am J Respir Crit Care Med. 2023;207 [Google Scholar]
  • 24.Vincent B.M., Molling D., Escobar G.J., et al. Hospital-specific template matching for benchmarking performance in a diverse multihospital system. Med Care. 2021;59(12):1090–1098. doi: 10.1097/MLR.0000000000001645. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 25.Collins G.S., Moons K.G.M., Dhiman P., et al. TRIPOD+AI statement: updated guidance for reporting clinical prediction models that use regression or machine learning methods. BMJ. 2024;385 doi: 10.1136/bmj-2023-078378. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 26.Agency for Healthcare Research and Quality, Clinical Classifications Software Refined (CCSR) for ICD-10-CM Diagnoses . Agency for Healthcare Research and Quality website; 2020. Healthcare Cost and Utilization Project (HCUP)https://www.hcup-us.ahrq.gov/toolssoftware/ccsr/ccs_refined.jsp [Google Scholar]
  • 27.Zimmerman J.E., Kramer A.A., McNair D.S., Malila F.M. Acute Physiology and Chronic Health Evaluation (APACHE) IV: hospital mortality assessment for today's critically ill patients. Crit Care Med. 2006;34(5):1297–1310. doi: 10.1097/01.CCM.0000215112.84523.F0. [DOI] [PubMed] [Google Scholar]
  • 28.Escobar G.J., Gardner M.N., Greene J.D., Draper D., Kipnis P. Risk-adjusting hospital mortality using a comprehensive electronic record in an integrated health care delivery system. Med Care. 2013;51:446–453. doi: 10.1097/MLR.0b013e3182881c8e. [DOI] [PubMed] [Google Scholar]
  • 29.Phillips G.S., Osborn T.M., Terry K.M., Gesten F., Levy M.M., Lemeshow S. The New York Sepsis Severity Score: development of a risk-adjusted severity model for sepsis. Crit Care Med. 2018;46(5):674–683. doi: 10.1097/CCM.0000000000002824. [DOI] [PubMed] [Google Scholar]
  • 30.Bisarya R., Song X., Salle J., Liu M., Patel A., Simpson S.Q. Antibiotic timing and progression to septic shock among patients in the ED with suspected infection. Chest. 2022;161(1):112–120. doi: 10.1016/j.chest.2021.06.029. [DOI] [PubMed] [Google Scholar]
  • 31.Peltan I.D., Brown S.M., Bledsoe J.R., et al. ED door-to-antibiotic time and long-term mortality in sepsis. Chest. 2019;155(5):938–946. doi: 10.1016/j.chest.2019.02.008. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 32.Rhee C., Wang R., Song Y., et al. Risk adjustment for sepsis mortality to facilitate hospital comparisons using Centers for Disease Control and Prevention’s adult sepsis event criteria and routine electronic clinical data. Crit Care Explor. 2019;1(10) doi: 10.1097/CCE.0000000000000049. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 33.Seymour C.W., Liu V.X., Iwashyna T.J., et al. Assessment of clinical criteria for sepsis: for the Third International Consensus Definitions for Sepsis and Septic Shock (Sepsis-3) JAMA. 2016;315(8):762–774. doi: 10.1001/jama.2016.0288. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 34.Schuler A., Wulf D.A., Lu Y., et al. The impact of acute organ dysfunction on long-term survival in sepsis. Crit Care Med. 2018;46(6):843–849. doi: 10.1097/CCM.0000000000003023. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 35.Ranzani O.T., Shankar-Hari M., Harrison D.A., et al. A comparison of mortality from sepsis in Brazil and England: the impact of heterogeneity in general and sepsis-specific patient characteristics. Crit Care Med. 2019;47(1):76–84. doi: 10.1097/CCM.0000000000003438. [DOI] [PubMed] [Google Scholar]
  • 36.Leligdowicz A., Dodek P.M., Norena M., Wong H., Kumar A., Kumar A. Association between source of infection and hospital mortality in patients who have septic shock. Am J Respir Crit Care Med. 2014;189:1204–1213. doi: 10.1164/rccm.201310-1875OC. [DOI] [PubMed] [Google Scholar]
  • 37.Prescott H.C., Harrison D.A., Rowan K.M., Shankar-Hari M., Wunsch H. Temporal trends in mortality of critically ill patients with sepsis in the United Kingdom, 1988-2019. Am J Respir Crit Care Med. 2024;209(5):507–516. doi: 10.1164/rccm.202309-1636OC. [DOI] [PubMed] [Google Scholar]
  • 38.Jin Y., Kattan M.W. Methodologic issues specific to prediction model development and evaluation. Chest. 2023;164(5):1281–1289. doi: 10.1016/j.chest.2023.06.038. [DOI] [PubMed] [Google Scholar]
  • 39.Wilson M.W., Labaki W.W., Choi P.J. Mortality and healthcare use of patients with compensated hypercapnia. Ann Am Thorac Soc. 2021;18(12):2027–2032. doi: 10.1513/AnnalsATS.202009-1197OC. [DOI] [PMC free article] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

e-Online Data
mmc1.docx (505.1KB, docx)

Articles from Chest are provided here courtesy of American College of Chest Physicians

RESOURCES