Skip to main content
NIHPA Author Manuscripts logoLink to NIHPA Author Manuscripts
. Author manuscript; available in PMC: 2026 Jul 24.
Published before final editing as: Ann Emerg Med. 2026 Jun 18:S0196-0644(26)00311-2. doi: 10.1016/j.annemergmed.2026.05.006

Improving End-of-Life Screening in the Emergency Department with Collaborative Artificial Intelligence

Adrian D Haimovich 1,*, Gabriel Erion-Barner 1,*, Larry A Nathanson 1, Caroline Cohen 1, Roger Orcutt 1, Smit Desai 2, David Rubins 3, Ula Hwang 4,5, R Andrew Taylor 6, Nathan I Shapiro 1, Kei Ouchi 7,**, Mara A Schonberg 3,**
PMCID: PMC13394201  NIHMSID: NIHMS2189575  PMID: 42313042

Abstract

Study Objectives:

To compare end-of-life predictions as measured by the physician-answered Surprise Question (SQ, “would you be surprised if this patient died in the next 6 months?”), the Geriatric End-of-Life Screening Tool (GEST) artificial intelligence (AI) model, and a new collaborative GEST+SQ model for predicting 6-month mortality in older emergency department (ED) patients.

Methods:

This was a single-site prospective cohort study (Nov 2022 - June 2023) at a tertiary academic ED of patients aged ≥65 years. Answers to the SQ were collected within the electronic health record (EHR) at ED disposition and GEST scores were calculated from available records using lab, vital sign, demographic and historical data. Six-month mortality was adjudicated via EHR and state records. SQ and GEST were compared using sensitivity and specificity. A new logistic regression model was developed combining SQ and GEST (GEST+SQ) and compared to GEST alone using area under receiver-operating characteristic curves (ROC-AUC) for discrimination and expected calibration error (ECE) for calibration. We modeled a sequential screening pathway where low- and high-risk patients received only GEST screening while intermediate risk patients received both GEST and SQ, reporting the proportion of patients for whom adding the SQ to GEST would change a theoretical referral to intervention.

Results:

From 9,256 eligible patients, 3,479 had SQ responses (37.6%), with 13.3% 6-month mortality. When matching GEST sensitivity to SQ (83.8%), GEST had greater specificity than the SQ (61.5% [56.7–67.1] vs. 50.8% [49.1–52.6]). At matching specificity (50.8%), GEST sensitivity (90.0% [87.0–92.7]) exceeded the SQ (83.8% [80.3–87.0]). GEST had an ROC-AUC of 0.79 (0.77–0.81), while the GEST+SQ model had ROC-AUC of 0.80 (0.78–0.82). The GEST+SQ model had significantly improved ECE of 0.01 (0.01–0.02) for GEST+SQ vs. 0.042 (0.03–0.05) for GEST alone. In a sequential screening pathway, as few as 5% of patients required SQ screening following GEST risk scoring.

Conclusions:

GEST modestly outperformed the SQ for predicting 6-month mortality. A GEST+SQ collaborative model did not improve discrimination (ROC-AUC) over GEST alone, but improved calibration. Sequential screening using GEST and then the SQ for intermediate risk patients could decrease physician screening burden by 95% relative to manual, SQ-only screening. Collaborative approaches integrating automated tools with targeted physician input may enhance ED mortality risk assessment while reducing clinician effort.

Introduction:

Mortality risk assessment is implicit in much of emergency medical care. Near-term (i.e., ≤ 30 day) mortality risk factors heavily into workup intensity, disposition, and level-of-care decision-making. For older adults, 75% of whom are seen in the emergency department (ED) within the last 6 months of life,1,2 assessment of longer-term mortality risk (≥ 6 months) is a cornerstone of evidence-based tools for shared decision-making3, ED-based serious illness communication4, palliative care referral5, comprehensive geriatric assessments6, and transitions of care7, among other diverse processes. Similarly, mortality risk is tightly linked to patient complexity and multimorbidity8, which connects to critical themes in ED quality including diagnostic quality9 and confidence10, workup intensity11, and operational throughput.12,13 Taken together, mortality risk is a latent variable highly relevant to the care of older adults in the ED. Many of the implications of mortality risk relate to broader goals of care. in an ideal setting this would be comprehensively discussed by the patient’s outpatient team before an ED visit, but in practice, basic goals of care documentation such as advance directives are available less than 50% of the time in an ED encounter, even in patients with chronic comorbidities, nursing facility residents, or older patients.14 Thus, ED physicians often must start from a blank slate when addressing issues related to mortality and goals of care.

Given the implicit importance of longer-term mortality risk among older adults in the ED, several tools have been developed to measure ED physician gestalt for mortality risk. We consider the best studied of these tools – the surprise question (SQ). The SQ, “would you be surprised if this patient died in the next six months” is validated in the ED, among other settings, where it predicts both mortality risk and resource utilization15, but, like other manual screening tools, requires clinician participation at each encounter for each eligible patient.1621 Nevertheless, health systems are now systematically incorporating this screening tool into ED operations.19,22

To reduce screening burden on clinicians while improving mortality risk assessment for older adults in the ED, we recently developed and externally validated1an automatable electronic health record (EHR) machine learning algorithm called the Geriatric End-of-life Screening Tool (GEST). GEST, an artificial intelligence (AI) logistic regression algorithm, uses commonly available EHR data including age, vital signs, and select labs to predict mortality and does not require clinician input.23,24 While predictive models, including AI models, are increasingly prevalent in emergency medicine, comparison to usual care or clinician impression are often lacking.21,2528 To address this gap, the objective of this study is to compare end-of-life predictions by the SQ, GEST, and by a new combined model incorporating SQ and GEST (GEST+SQ), both in head-to-head performance analysis as well as in a hypothetical sequential screening application where GEST is used to pre-screen for SQ focusing on prognostic performance and potential reductions in clinician bedside screening effort.

Methods:

Study Design:

This is a single site prospective cohort study conducted between 11/1/2022 and 06/30/2023 at Beth Israel Deaconess Medical Center, a tertiary care academic ED in Boston, MA. During the study period, the SQ question was incorporated into the ED disposition activity for all patients ≥65 years of age. Due to institutional policies, the SQ was implemented as a “soft-stop” at the time of patient admission or discharge and could be skipped without interrupting the clinical workflow.29 The SQ was only recorded only once per ED encounter and was not linked with further interventions. Patients with SQ responses were included in this study. Patients with a home address outside of Massachusetts were excluded due to unreliable mortality data. Only the index visit (first encounter in the study period) for each patient was included. Patient race, ethnicity, and sex were sourced from the EHR as collected during the registration process. Patients who died within 24 hours of ED arrival were excluded in alignment with the GEST derivation and validation studies. We used the TRIPOD+AI reporting guideline.30 The study was approved by the BIDMC Institutional Review Board (2022P000795).

GEST scoring:

GEST scores were calculated using EHR data four hours into the patient encounter using parameters from the previously published method including age, complete blood count and basic metabolic panel data, need for supplemental oxygen in the ED, and data about ED diagnosis, prior diagnoses, and admissions.23,24 Data were centered and missing data mean-imputed according to the prespecified GEST model. Mortality was assessed using EHR and Massachusetts death records which were matched using previously described methods.24 Last names, birth dates, and sex required exact matches, while first names were matched probabilistically. The GEST score is an estimate of the patient’s six-month mortality risk (0–100%).

Methods of comparison between GEST and SQ:

Summary statistics of the patient cohort as well as the SQ-positive and SQ-negative subgroups were calculated. We report demographics and mortality of groups who had the SQ answered and not answered, as well as the SQ responses.3133 Factors associated with the presence or absence of SQ response were analyzed using logistic regression.

We calculated sensitivity and specificity for six-month mortality for SQ responses and for GEST. GEST sensitivity was reported at the operating point where its specificity was equal to the SQ, and vice versa for specificity. Sensitivity, sensitivity and F1 score (the harmonic mean of precision and recall) were calculated for each user. Factors associated with improved F1 score were analyzed using logistic regression. Calibration curves were plotted using a sigmoid curve fit. Population-weighted net reclassification improvement (NRI) was also used to measure the number of cases correctly reclassified by GEST relative to SQ.34 Confidence intervals were calculated using 10,000 bootstraps. A sensitivity analysis was performed by responding provider type (attending or resident). SQ inter-rater reliability (IRR) was assessed separately using Krippendorff’s alpha35 in a cohort of patients with repeat encounters containing SQ responses, allowing within-patient comparison across encounters. As described above, repeat encounters were not included in any other analyses. IRR bootstraps used 100 replicates.

Combined model training and evaluation:

A logistic regression model was trained using ten-fold cross-validation to predict six-month mortality in our cohort using either the SQ alone, GEST alone, or GEST+SQ. Receiver-operating characteristic curves (ROC) and calibration curves were calculated for each model and summarized with area under the ROC curve (ROC-AUC), measuring discrimination and expected calibration error (ECE) respectively. ECE, a standard calibration metric in machine learning, measures how well predicted probabilities match true event rates.36 We used ECE rather than the Brier score because the Brier score is a composite of discrimination and calibration, rather than isolated calibration.37 The final model used in sequential screening modeling resulted from averaging coefficients across all folds. Code was written in Python version (v3.12.3), models were implemented in pure Python and scikit-learn (v1.6.1)38, and statistical analysis was performed with the Scipy (v1.15.2)39 and Statsmodels (v0.14.4)40 packages.

Sequential GEST+SQ screening pathway modeling:

We modeled the integration of GEST into an ED-based screening workflow using the cohort of patients with SQ responses.24 In the baseline approach, all patients aged ≥65 years are screened using the SQ, as is current practice at some hospitals.15,19 In contrast, our sequential screening strategy involves identifying all patients above a pre-specified mortality risk threshold - we chose 20%, 30%, and 40% thresholds, in keeping with prior work.23,41,42 In this analysis, we used the GEST+SQ model to assign each patient into a low, intermediate, or high-risk group. A patient was low risk if their GEST score was sufficiently low such that a SQ response of “no” would not reassign them above the mortality risk threshold, while for the high-risk group, the GEST score was sufficiently high such that the SQ response of “yes” would not reassign them to the low-risk group. A patient was intermediate risk if crossing the pre-specified mortality risk threshold depended on the SQ response. In a sequential screening scenario, only intermediate-risk patients required clinician SQ screening.

Results:

During the study period, there were a total of 12,939 encounters by 9,256 patients aged ≥65 years; 3,479 patients (37.6%) had SQ responses after excluding patients who expired within 24 hours (Supplementary Figure 1). The mean age of this SQ cohort was 77.8 (SD 8.3) years, 54.5% of patients were female, 7.0% were Hispanic, and 66.9% were white (Table 1). The proportion of admissions was 62.2%, the average GEST score across the entire cohort was 11.4%, and 462 (13.3%) patients died within 6-months. Comparison with the GEST derivation demographics is shown in Supplementary Table S1.

Table 1:

Study demographics. Statistics are shown for the subset of the population with “Yes” and “No” answers to the surprise question. For each categorical variable, n and N indicate the number of patients in a given category and the total number of patients with data available for that variable, respectively.

Surprise “No” Surprise “Yes”

Total n 1870 1609

Sex, n/N (%) Female 994/1870 (53.2) 903/1609 (56.1)
Male 876/1870 (46.8) 706/1609 (43.9)

Hispanic ethnicity , n/N (%) 114/1795 (6.4) 131/1572 (8.3)

Race, n/N (%) AI/AN 1/1809 (0.1) 2/1592 (0.1)
Asian 111/1809 (6.1) 76/1592 (4.8)
Black 331/1809 (18.3) 395/1592 (24.8)
Native Hawaiian or other Pacific Islander 1/1809 (0.1) 0/1592 (0.0)
Other 74/1809 (4.1) 82/1592 (5.2)
Unknown 61/1809 (3.4) 17/1592 (1.1)
White 1291/1809 (71.3) 1037/1592 (65.1)

Admitted, n/N (%) 1433/1870 (76.6) 730/1609 (45.4)

Age, mean (SD) 79.6 (8.7) 75.7 (7.3)

GEST Score, mean (SD) 15.6 (16.6) 6.6 (8.8)

Six month mortality, n/N (%) 387/1870 (20.7) 75/1609 (4.7)

“GEST Score” is the output of the model studied in this paper. It is an estimate of the probability of patient mortality in the next 6 months, ranging from 0–100%, with greater values indicating higher risk of mortality. Prior literature has used estimated mortality of roughly 30% as a clinically meaningful threshold for palliative referral.40

Among included encounters, 387 of the 1,870 patients with a SQ response of “No” (would not be surprised) experienced 6-month mortality (20.7%), as compared to 75 of the 1,609 patients with a response of “Yes” (4.7%). Patients for whom the SQ was answered were older, less likely to be Hispanic or white, and had higher GEST scores and mortality than those without the SQ (Supplementary Table S2). Of SQ responses, 91.3% were from resident physicians (Supplementary Table S3). Patients who had attending responses were younger, had lower GEST scores, lower admission rates, and lower mortality rates than those seen by residents (Supplementary Table S3). The median number of SQ responses per physician was 4 (IQR: 2–12.5). The treating physician’s SQ response rate was more strongly associated with SQ response than patient age, sex, race, ethnicity, disposition, GEST score, or whether the physician was a resident or attending (Supplementary Figure S2).

The SQ had a sensitivity of 83.8% (80.3–87.0%) and specificity of 50.8% (49.1–52.6). This performance is similar to other studies of the SQ.16 At a matching sensitivity, GEST had a specificity of 61.5% (56.7–67.1) while at matching specificity, GEST had sensitivity of 90.0% (87.0–92.7) (Figure 1). GEST outperformed resident physician SQ responses (GEST sensitivity 0.90 [0.87–0.93] and specificity 0.61 [0.56–0.67] vs resident SQ sensitivity 0.84 [0.80–0.87] and specificity 0.50 [0.48–0.52]) but confidence intervals overlapped with attendings’ performance in predicting mortality (GEST sensitivity 0.83 [0.67–0.96] and specificity 0.58 [0.06–0.79] vs attending SQ sensitivity 0.87 [0.71–1.0] and specificity 0.63 [0.57–0.68]) (Supplementary Figure S3). NRI for GEST versus the SQ ranged from 0.28 (0.25–0.30) to 0.32 (0.29–0.34), favoring GEST (Supplementary Figure S4). The strongest predictor of physician accuracy was average GEST score of their patients (log odds-ratio 0.51 [−0.029–1.05]), but the confidence interval included zero, as it did for resident vs attending status, rate of SQ response, and patient age, disposition, race, ethnicity, and sex (Supplementary Table S4). Inter-rater reliability was assessed separately using the subset of 400 patients with repeat SQ evaluations during the study period – these repeat encounters were not used in the other analyses. Krippendorff’s alpha was 0.190 (0.113–0.292).

Figure 1:

Figure 1:

ROC-AUC of GEST vs clinician gestalt using the Surprise Question

ROC curve for GEST model achieves higher sensitivity and specificity than clinicians answering the Surprise Question (SQ). GEST is shown as an ROC curve while clinicians, with binary answers, are shown as points. GEST sensitivity and specificity at operating points that match the SQ are shown with horizontal (sensitivity at matching specificity) and vertical (specificity at matching sensitivity) lines representing a 95% confidence interval on the corresponding value. The SQ is marked with 95% confidence intervals as well. There is considerable variability in the sensitivity and specificity of individual clinicians’ answers to the SQ.

The combined GEST+SQ model (Supplementary Figure S5) had an ROC-AUC of 0.80 (0.78–0.82) compared to the cross-validated ROC-AUC of 0.79 (0.77–0.81) for GEST. The combined model had an expected calibration error of 0.01 (0.01–0.02) compared to the GEST error of 0.04 (0.03–0.05) (Figure 2, Supplementary Figure S6).

Figure 2:

Figure 2:

Calibration of GEST and the combined GEST+SQ model

A model combining the GEST score with clinician gestalt represented by the Surprise Question (SQ) achieves better calibration, in that the predicted p(mortality) from the model more closely estimates the observed mortality rate. A calibration curve matching the dashed line would have perfect performance, i.e., its predicted and empiric mortality rates are equal. In this figure, observed mortality rates are smoothed as described in the methods. Binned probabilities are shown in Supplemental Figure S5.

Using mortality risk cutoffs of 20%, 30%, and 40%, sequential screening models indicated 81.2%, 89.6%, and 93.7% of patients had sufficiently low GEST scores to not require SQ screening because even a positive response would not cross the mortality risk cutoff, while 4.7%, 2.3%, and 1.3% of patients had scores that were sufficiently high that even a negative SQ screen would not cross the mortality risk threshold (Figure 3). This left 14.1%, 8.1%, or 5.0% of patients assigned intermediate-risk and requiring SQ screening.

Figure 3:

Figure 3:

Modeling sequential palliative needs screening using GEST to automatically rule-in and out high- and low-risk patients, respectively and then SQ for intermediate-risk patients.

At proposed palliative referral thresholds of 20%, 30%, and 40% mortality risk, the combined-model referral strategy involves prompting clinicians for input (i.e., provide a SQ response) on 14.1%, 8.1%, and 5.0% of patients, respectively. The intermediate-risk range is highlighted in the purple and indicates patients for whom, if the SQ answer is “no”, the GEST+SQ model would indicate mortality risk above the referral threshold (horizontal gray bar), while if it is ”yes”, the model would indicate risk below the threshold. For all other patients, the GEST+SQ predicted risk would be either below (left, light blue shading) or above (right, light orange shading) the referral threshold regardless of the SQ answer. A histogram showing the percentage of patients with each given GEST score is shown for reference.

Discussion:

In this single-site, prospective cohort study, we found that GEST, an automated EHR-integrated algorithm, had slightly superior prognostic performance to treating clinicians in identifying older adults with 6-month mortality risk. When combining GEST with SQ, model calibration, the reliability of the assigned risk values, improved, but discrimination, the accurate ranking of predictions, did not. We note that in clinical decision support, calibration may be more important than discrimination because calibration focuses on predicting accurate risk for the current patient rather than ranking them relative to the whole population.43

These data suggest that SQ, GEST, or the collaborative clinician-AI GEST+SQ model are reasonable approaches for mortality prediction in the ED. Moreover, we performed a modeling analysis where a sequential screening approach using AI (i.e., GEST) pre-screened patients, ruling in and out high risk and low risk patients, respectively, saving clinician input (i.e., SQ) for the intermediate cases. We found that as many as 95% of patients would not require clinician input with this approach. Given the observed performance equipoise between GEST and SQ, the clinician effort savings of GEST+SQ may be preferable for ED teams implementing mortality risk assessment tools for older adults.

Automated mortality risk screening with GEST has several potential applications. First, mortality risk assessment has operational implications in anticipating resource use and screening for care transition needs.7 Second, GEST may be a useful quality assurance tool, both for screening for high-risk cases for adverse outcomes and for risk-adjustment in analyses of clinician- and site-level care variation. Finally, GEST can be used to trigger palliative care referral or ED team-driven serious illness communication, depending on department resources. 1,4,5,44

There are numerous ED screening tools for the older adult ED population, many of which have partially overlapping target outcomes (e.g., mortality, frailty, functional decline, ED revisit), but systematic implementation of these tools is limited, even among accredited geriatric emergency departments.45 A critical step to filling this gap is transitioning from clinician-driven screening (i.e., physicians, nurses, social workers), to automated systems tightly integrated into electronic health record (EHR) workflows.46 While these automated tools have proliferated, few have been compared to physician judgement, and, therefore, their clinical applicability and adequacy is unknown.27,28 While GEST was not trained to predict other adverse ED outcomes for older adults, the strong associations between mortality risk and frailty, functional decline, and revisits, introduce the possibility that GEST may be able to augment or replace multiple ED screening tools, further streamlining ED processes.14

While not explicitly analyzed in this work, the advent of large language models (LLMs) raises the potential for future iterations of tools like GEST with significantly improved performance. LLMs alone typically do not show clear benefit over classical machine learning methods for the type of probabilistic risk stratification discussed here, but they are able to leverage a new source of data in the form of unstructured text.47 We anticipate that LLM-based feature extraction will provide new, highly informative features for models like GEST resulting in even higher predictive performance.48

Limitations of this work include data from a single academic medical center and the exclusion of patients with primary residence outside of MA due to restrictions on mortality capture. The high admission rate (62.2%) limits generalizability. A large proportion of responses were by resident physicians; patients with SQ responses by resident physicians were older and higher-mortality than those with attending responses, both of which limit head-to-head SQ and GEST comparison. An important limitation is missingness in responses to the SQ, which was primarily driven by provider response rate, rather than patient attributes. This aligns with prior literature on clinician-level variation in EHR use patterns and while this missingness would be addressed by mandatory responses, also called hard-stops, prior work has suggested that these can lead to misleading responses.29,49 The low response rate highlights the difficulty of operationalizing mortality screening in the ED but supports the goal of using models like GEST+SQ to reduce the physician burden of screening tools.

Taken together, we argue that the approach of using AI to pre-screen for clearly high- or low-risk patients while relying on clinicians for intermediate cases has potential broad applicability to ED processes. EDs increasingly support broad individual and population health screening objectives including substance use, suicide, housing instability, elder safety, and others.5052 While, individually, these may be feasible, the cumulative effects of screening programs can burden ED teams. We propose that this sequential, collaborative AI-clinician approach may be applied to many other ED screening activities, saving both clinician effort and enabling a wider array of interventions.

Supplementary Material

1

Grant:

ADH was supported by K12TR004381, UH was supported by R33AG058926, MAS was supported by K24AG071906, KO was supported by K76AG064434.

Footnotes

Conflicts of Interest:

Authors report no conflicts of interest.

Meetings: Preliminary data from this study were presented at the Society for Academic Emergency Medicine Annual Meeting in May, 2025.

IRB/Ethics Information:

IRB/Ethics information: approved by BIDMC IRB 2022P000795.

Study Protocol/Registration:

The study protocol was not separately prepared and the study was not pre-registered.

Patient/Public Involvement:

There was no patient or public involvement in the design, execution, or reporting of this study.

Publisher's Disclaimer: This is a PDF file of an unedited manuscript that has been accepted for publication. As a service to our customers we are providing this early version of the manuscript. The manuscript will undergo copyediting, typesetting, and review of the resulting proof before it is published in its final form. Please note that during the production process errors may be discovered which could affect the content, and all legal disclaimers that apply to the journal pertain.

Data Sharing Agreement:

Data are not available for sharing due to PHI.

Code Sharing:

Code is available upon request.

References

  • 1.Ouchi K, George N, Schuur JD, Aaronson EL, Lindvall C, Bernstein E, Sudore RL, Schonberg MA, Block SD, Tulsky JA. Goals-of-care conversations for older adults with serious illness in the emergency department: Challenges and opportunities. Ann Emerg Med. Elsevier BV; 2019. Aug;74(2):276–284. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 2.Smith AK, McCarthy E, Weber E, Cenzer IS, Boscardin J, Fisher J, Covinsky K. Half of older Americans seen in emergency department in last month of life; most admitted to hospital, and many die there. Health Aff (Millwood). Health Affairs (Project Hope); 2012. June;31(6):1277–1285. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 3.Haimovich AD, Chary A, Burke L, Janke AT, Rodman A, Landon B, Shapiro NI, Naik AD, Schoenfeld E, Ouchi K, Schonberg MA. Marginal dispositions and shared decision-making among older adults in the ED: A prospective cohort study. Acad Emerg Med [Internet]. Wiley; 2025. Dec 19;(acem.70211). Available from: 10.1111/acem.70211 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 4.Ouchi K, Block SD, Rentz DM, Berry DL, Oelschlager H, Shiozawa Y, Rossmassler S, Berger AL, Hasdianda MA, Wang W, Boyer E, Sudore RL, Tulsky JA, Schonberg MA. Serious illness conversations in the emergency department for older adults with advanced illnesses: A randomized clinical trial. JAMA Netw Open. American Medical Association; 2025. June 2;8(6):e2516582. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 5.Grudzen CR, Siman N, Cuthel AM, Adeyemi O, Yamarik RL, Goldfeld KS, PRIM-ER Investigators, Abella BS, Bellolio F, Bourenane S, Brody AA, Cameron-Comasco L, Chodosh J, Cooper JJ, Deutsch AL, Elie MC, Elsayem A, Fernandez R, Fleischer-Black J, Gang M, Genes N, Goett R, Heaton H, Hill J, Horwitz L, Isaacs E, Jubanyik K, Lamba S, Lawrence K, Lin M, Loprinzi-Brauer C, Madsen T, Miller J, Modrek A, Otero R, Ouchi K, Richardson C, Richardson LD, Ryan M, Schoenfeld E, Shaw M, Shreves A, Southerland LT, Tan A, Uspal J, Venkat A, Walker L, Wittman I, Zimny E. Palliative care initiated in the emergency department: A cluster randomized clinical trial: A cluster randomized clinical trial. JAMA: the journal of the American Medical Association. American Medical Association (AMA); 2025. p. 599–608. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6.Galvin R, Gilleit Y, Wallace E, Cousins G, Bolmer M, Rainer T, Smith SM, Fahey T. Adverse outcomes in older adults attending emergency departments: a systematic review and meta-analysis of the Identification of Seniors At Risk (ISAR) screening tool. Age Ageing. Oxford University Press (OUP); 2017. Mar 1;46(2):179–186. [DOI] [PubMed] [Google Scholar]
  • 7.Jacobsohn GC, Jones CMC, Green RK, Cochran AL, Caprio TV, Cushman JT, Kind AJH, Lohmeier M, Mi R, Shah MN. Effectiveness of a care transitions intervention for older adults discharged home from the emergency department: A randomized controlled trial. Acad Emerg Med. Wiley; 2022. Jan;29(1):51–63. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 8.Schneider C, Aubert CE, Del Giovane C, Donzé JD, Gastens V, Bauer DC, Blum MR, Dalleur O, Henrard S, Knol W, O’Mahony D, Curtin D, Lee SJ, Aujesky D, Rodondi N, Feller M. Comparison of 6 mortality risk scores for prediction of 1-year mortality risk in older adults with multimorbidity. JAMA Netw Open. American Medical Association (AMA); 2022. July 1;5(7):e2223911. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 9.Jawad BN, Pedersen KZ, Andersen O, Meier N. Minimizing the risk of diagnostic errors in acute care for older adults: An interdisciplinary patient safety challenge. Healthcare (Basel). MDPI AG; 2024. Sept 13;12(18):1842. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 10.Haimovich AD, Janke AT, Kocher KE, Mangus CW, Parsons AS, McCoy L, Taylor RA, Rodman A, Pusic M. Managing clinical uncertainty: Formalizing management reasoning in emergency care delivery. Ann Emerg Med [Internet]. Elsevier BV; 2025. Oct 10; Available from: 10.1016/j.annemergmed.2025.09.007 [DOI] [PubMed] [Google Scholar]
  • 11.Herring AA, Johnson B, Ginde AA, Camargo CA, Feng L, Alter HJ, Hsia R. High-intensity emergency department visits increased in California, 2002–09. Health Aff (Millwood). Health Affairs (Project Hope); 2013. Oct;32(10):1811–1819. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12.Kanzaria HK, Probst MA, Ponce NA, Hsia RY. The association between advanced diagnostic imaging and ED length of stay. Am J Emerg Med. Elsevier BV; 2014. Oct;32(10):1253–1258. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 13.Perotte R, Lewin GO, Tambe U, Galorenzo JB, Vawdrey DK, Akala OO, Makkar JS, Lin DJ, Mainieri L, Chang BC. Improving emergency department flow: Reducing turnaround time for emergent CT scans. AMIA Annu Symp Proc. 2018. Dec 5;2018:897–906. [PMC free article] [PubMed] [Google Scholar]
  • 14.Oulton J, Rhodes SM, Howe C, Fain MJ, Mohler MJ. Advance directives for older adults in the emergency department: a systematic review. J Palliat Med. SAGE Publications; 2015. June;18(6):500–505. [DOI] [PubMed] [Google Scholar]
  • 15.Haydar SA, Strout TD, Bond AG, Han PK. Prognostic value of a modified surprise question designed for use in the emergency department setting. Clin Exp Emerg Med. The Korean Society of Emergency Medicine; 2019. Mar;6(1):70–76. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16.Ouchi K, Jambaulikar G, George NR, Xu W, Obermeyer Z, Aaronson EL, Schuur JD, Schonberg MA, Tulsky JA, Block SD. The “surprise question” asked of emergency physicians may predict 12-month mortality among older emergency department patients. J Palliat Med. 2018. Feb;21(2):236–240. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 17.van Lummel EV, Ietswaard L, Zuithoff NP, Tjan DH, van Delden JJ. The utility of the surprise question: A useful tool for identifying patients nearing the last phase of life? A systematic review and meta-analysis. Palliat Med. SAGE Publications; 2022. July;36(7):1023–1046. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 18.Downar J, Goldman R, Pinto R, Englesakis M, Adhikari NKJ. The “surprise question” for predicting death in seriously ill patients: a systematic review and meta-analysis. CMAJ. CMA Joule Inc; 2017. Apr 3;189(13):E484–E493. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19.Ouchi K, Strout T, Haydar S, Baker O, Wang W, Bernacki R, Sudore R, Schuur JD, Schonberg MA, Block SD, Tulsky JA. Association of emergency clinicians’ assessment of mortality risk with actual 1-month mortality among older adults admitted to the hospital. JAMA Netw Open. American Medical Association (AMA); 2019. Sept 4;2(9):e1911139. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 20.White N, Kupeli N, Vickerstaff V, Stone P. How accurate is the ‘Surprise Question’ at identifying patients at the end of life? A systematic review and meta-analysis. BMC Med [Internet]. Springer Science and Business Media LLC; 2017. Dec;15(1). Available from: 10.1186/s12916-017-0907-4 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 21.Knack SKS, Scott N, Driver BE, Prekker ME, Black LP, Hopson C, Maruggi E, Kaus O, Tordsen W, Puskarich MA. Early physician gestalt versus usual screening tools for the prediction of sepsis in critically ill emergency patients. Ann Emerg Med. Elsevier BV; 2024. Sept;84(3):246–258. [DOI] [PubMed] [Google Scholar]
  • 22.Dilip M, Van Tonder R, Jubanyik K, Venkatesh A, Rhodes D, Sangal R, Kim N. 509 impact of the mortality surprise question on emergency department clinician behavior. Ann Emerg Med. Elsevier BV; 2025. Sept;86(3):S217. [Google Scholar]
  • 23.Haimovich AD, Xu W, Wei A, Schonberg MA, Hwang U, Taylor RA. Automatable end-of-life screening for older adults in the emergency department using electronic health records. J Am Geriatr Soc. Wiley; 2023. June;71(6):1829–1839. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 24.Haimovich AD, Burke RC, Nathanson LA, Rubins D, Taylor RA, Kross EK, Ouchi K, Shapiro NI, Schonberg MA. Geriatric End-of-Life Screening Tool prediction of 6-month mortality in older patients. JAMA Netw Open. American Medical Association (AMA); 2024. May 1;7(5):e2414213. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 25.Kareemi H, Vaillancourt C, Rosenberg H, Fournier K, Yadav K. Machine learning versus usual care for diagnostic and prognostic prediction in the emergency department: A systematic review. Acad Emerg Med. Wiley; 2021. Feb;28(2):184–196. [DOI] [PubMed] [Google Scholar]
  • 26.Schriger DL, Elder JW, Cooper RJ. Structured clinical decision aids are seldom compared with subjective physician judgment, and are seldom superior. Ann Emerg Med. 2017. Sept;70(3):338–344.e3. [DOI] [PubMed] [Google Scholar]
  • 27.Levin S, Toerper M, Hamrock E, Hinson JS, Barnes S, Gardner H, Dugas A, Linton B, Kirsch T, Kelen G. Machine-learning-based electronic triage more accurately differentiates patients with respect to clinical outcomes compared with the Emergency Severity Index. Ann Emerg Med. 2018. May;71(5):565–574.e2. [DOI] [PubMed] [Google Scholar]
  • 28.Hinson JS, Taylor RA, Venkatesh A, Steinhart BD, Chmura C, Sangal RB, Levin SR. Accelerated chest pain treatment with artificial intelligence-informed, risk-driven triage. JAMA Intern Med. American Medical Association (AMA); 2024. Sept 1;184(9):1125–1127. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 29.Powers EM, Shiffman RN, Melnick ER, Hickner A, Sharifi M. Efficacy and unintended consequences of hard-stop alerts in electronic health record systems: a systematic review. J Am Med Inform Assoc. Oxford University Press (OUP); 2018. Nov 1;25(11):1556–1566. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 30.Collins GS, Moons KGM, Dhiman P, Riley RD, Beam AL, Van Calster B, Ghassemi M, Liu X, Reitsma JB, van Smeden M, Boulesteix AL, Camaradou JC, Celi LA, Denaxas S, Denniston AK, Glocker B, Golub RM, Harvey H, Heinze G, Hoffman MM, Kengne AP, Lam E, Lee N, Loder EW, Maier-Hein L, Mateen BA, McCradden MD, Oakden-Rayner L, Ordish J, Parnell R, Rose S, Singh K, Wynants L, Logullo P. TRIPOD+AI statement: updated guidance for reporting clinical prediction models that use regression or machine learning methods. BMJ. BMJ; 2024. Apr 16;385:e078378. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 31.Johnson KS. Racial and ethnic disparities in palliative care. J Palliat Med. Mary Ann Liebert Inc; 2013. Nov;16(11):1329–1334. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 32.Bazargan M, Bazargan-Hejazi S. Disparities in palliative and hospice care and completion of advance care planning and directives among non-Hispanic Blacks: A scoping review of recent literature. Am J Hosp Palliat Care. SAGE Publications; 2021. June;38(6):688–718. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 33.Ornstein KA, Roth DL, Huang J, Levitan EB, Rhodes JD, Fabius CD, Safford MM, Sheehan OC. Evaluation of racial disparities in hospice use and end-of-life treatment intensity in the REGARDS cohort. JAMA Netw Open. American Medical Association (AMA); 2020. Aug 3;3(8):e2014639. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 34.Kerr KF, Wang Z, Janes H, McClelland RL, Psaty BM, Pepe MS. Net reclassification indices for evaluating risk prediction instruments: a critical review. Epidemiology. Ovid Technologies (Wolters Kluwer Health); 2014. Jan;25(1):114–121. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 35.Castro S Fast Krippendorff: Fast computation of Krippendorff’s alpha agreement measure [Internet]. 2025. Available from: https://github.com/pln-fing-udelar/fast-krippendorff [Google Scholar]
  • 36.Naeini MP, Cooper GF, Hauskrecht M. Obtaining well calibrated probabilities using Bayesian Binning. Proc Conf AAAI Artif Intell. 2015. Jan;2015:2901–2907. [PMC free article] [PubMed] [Google Scholar]
  • 37.Murphy AH. A new vector partition of the probability score. J Appl Meteorol. American Meteorological Society; 1973. June;12(4):595–600. [Google Scholar]
  • 38.Pedregosa F, Varoquaux G, Gramfort A, Michel V, Thirion B, Grisel O, Blondel M, Prettenhofer P, Weiss R, Dubourg V, Vanderplas J, Passos A, Cournapeau D, Brucher M, Perrot M, Duchesnay E. Scikit-learn: Machine Learning in Python. Journal of Machine Learning Research. 2011;12:2825–2830. [Google Scholar]
  • 39.Virtanen P, Gommers R, Oliphant TE, Haberland M, Reddy T, Cournapeau D, Burovski E, Peterson P, Weckesser W, Bright J, van der Walt SJ, Brett M, Wilson J, Millman KJ, Mayorov N, Nelson ARJ, Jones E, Kern R, Larson E, Carey CJ, Polat İ, Feng Y, Moore EW, VanderPlas J, Laxalde D, Perktold J, Cimrman R, Henriksen I, Quintero EA, Harris CR, Archibald AM, Ribeiro AH, Pedregosa F, van Mulbregt P, SciPy 1.0 Contributors. SciPy 1.0: fundamental algorithms for scientific computing in Python. Nat Methods. Springer Science and Business Media LLC; 2020. Mar;17(3):261–272. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 40.Seabold S, Perktold J. statsmodels: Econometric and statistical modeling with python. 9th Python in Science Conference. 2010. [Google Scholar]
  • 41.Courtright KR, Chivers C, Becker M, Regli SH, Pepper LC, Draugelis ME, O’Connor NR. Electronic health record mortality prediction model for targeted palliative care among hospitalized medical patients: A pilot quasi-experimental study. J Gen Intern Med. Springer Science and Business Media LLC; 2019. Sept;34(9):1841–1847. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 42.Parikh RB, Manz C, Chivers C, Regli SH, Braun J, Draugelis ME, Schuchter LM, Shulman LN, Navathe AS, Patel MS, O’Connor NR. Machine learning approaches to predict 6-month mortality among patients with cancer. JAMA Netw Open. American Medical Association (AMA); 2019. Oct 2;2(10):e1915997. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 43.Van Calster B, McLernon DJ, van Smeden M, Wynants L, Steyerberg EW, Topic Group ‘Evaluating diagnostic tests and prediction models’ of the STRATOS initiative. Calibration: the Achilles heel of predictive analytics. BMC Med. Springer Science and Business Media LLC; 2019. Dec 16;17(1):230. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 44.Lupu D, American Academy of Hospice and Palliative Medicine Workforce Task Force. Estimate of current hospice and palliative medicine physician workforce shortage. J Pain Symptom Manage. Elsevier BV; 2010. Dec;40(6):899–911. [DOI] [PubMed] [Google Scholar]
  • 45.Kennedy M, Lesser A, Israni J, Liu SW, Santangelo I, Tidwell N, Southerland LT, Carpenter CR, Biese K, Ahmad S, Hwang U. Reach and adoption of a Geriatric Emergency Department Accreditation program in the United States. Ann Emerg Med. Elsevier BV; 2022. Apr;79(4):367–373. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 46.Haimovich AD, Shah MN, Southerland LT, Hwang U, Patterson BW. Automating risk stratification for geriatric syndromes in the emergency department. J Am Geriatr Soc. Wiley; 2024. Jan;72(1):258–267. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 47.Gu B, Desai RJ, Lin KJ, Yang J. Probabilistic medical predictions of large language models. NPJ Digit Med. 2024. Dec 19;7(1):367. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 48.Wright DS, Socrates V, Huang T, Safranek CW, Sangal RB, Dilip M, Boivin Z, Srica N, Wright CX, Feher A, Miller EJ, Chartash D, Taylor RA. Automated computation of the HEART score with the GPT-4 large language model. Am J Emerg Med. Elsevier BV; 2025. July;93:120–125. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 49.Beeler PE, Orav EJ, Seger DL, Dykes PC, Bates DW. Provider variation in responses to warnings: do the same providers run stop signs repeatedly? J Am Med Inform Assoc. Oxford University Press (OUP); 2016. Apr;23(e1):e93–8. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 50.Patterson BW, Jacobsohn GC, Maru AP, Venkatesh AK, Smith MA, Shah MN, Mendonça EA. RESEARCHComparing strategies for identifying falls in older adult emergency department visits using EHR data. J Am Geriatr Soc. Wiley; 2020. Dec;68(12):2965–2967. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 51.D’Onofrio G, Degutis LC. Preventive care in the emergency department: screening and brief intervention for alcohol problems in the emergency department: a systematic review. Acad Emerg Med. Wiley; 2002. June;9(6):627–638. [DOI] [PubMed] [Google Scholar]
  • 52.Boudreaux ED, Camargo CA Jr, Arias SA, Sullivan AF, Allen MH, Goldstein AB, Manton AP, Espinola JA, Miller IW. Improving suicide risk screening and detection in the emergency department. Am J Prev Med. Elsevier BV; 2016. Apr;50(4):445–453. [DOI] [PMC free article] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

1

Data Availability Statement

Data are not available for sharing due to PHI.

RESOURCES