Skip to main content
JAMIA Open logoLink to JAMIA Open
. 2026 Jul 16;9(4):ooag123. doi: 10.1093/jamiaopen/ooag123

DETERIO-LLM: enhancing traditional deterioration risk scores with clinical context and advanced reasoning

Ghodsieh Ghanbari 1,✉, Joseph C Ahn 2, Eileen Kim 3, Ruth Laverde 4, Zach Pope 5, Shamim Nemati 6
PMCID: PMC13375265  PMID: 42466429

Abstract

Objectives

Accurate and timely prediction of physiological deterioration in hospitalized patients is essential, but existing deep learning models, which rely solely on structured electronic health record (EHR) data, generate high false-positive rates by missing crucial contextual information in clinical notes. We present DETERIO-LLM, a hybrid prediction model that selectively integrates large language models (LLMs) with a deep learning model (DETERIO) to reclassify borderline-risk alerts using narrative clinical notes.

Materials and Methods

This retrospective study used a cohort of 1000 inpatients (4.6% deterioration prevalence). DETERIO-LLM selectively analyzed narrative notes for alerts in the uncertainty range (scores 2.5-3.5). Performance was evaluated using sensitivity, positive predictive value (PPV), and F1 score, benchmarked against comparators including the best-established ML model, eCART.

Results

DETERIO-LLM significantly improved predictive precision over all comparators. At its optimal threshold, the model achieved a PPV of 30.6% and an F1 score of 37.3%, outperforming eCART. This selective integration reduced false-positive alerts by 46.5% within the uncertainty range while maintaining comparable sensitivity.

Discussion

This study demonstrates a practical, interpretable method for integrating LLM-derived contextual knowledge into existing structured EHR prediction workflows. By focusing LLM analysis on high-uncertainty alerts, DETERIO-LLM substantially improves clinical utility through a marked reduction in false-positive cases, which is critical for mitigating alert fatigue.

Conclusion

Selective LLM integration with structured EHR models offers a viable, interpretable path to improving early warning systems in clinical care by leveraging narrative context to reduce alert fatigue while enhancing predictive performance.

Keywords: clinical notes, large language models, clinical decision support systems

Background

Clinical deterioration, characterized by an acute worsening of a patient’s physiological condition, substantially increases the risk of severe morbidity and mortality. Timely identification and intervention in deteriorating patients can significantly improve clinical outcomes, reduce preventable harm, and lower healthcare costs.1–8 To address this need, various predictive frameworks have been developed in order to improve timely identification. Traditional rule-based tools, such as the Modified Early Warning Score (MEWS), National Early Warning Score (NEWS), and its updated version NEWS2, rely on fixed thresholds of vital signs and mental status to generate risk alerts. These systems are validated and widely utilized. However, their low specificity often leads to false-positive alerts and clinician alert fatigue, thereby compromising patient safety.9,10

In recent years, machine learning (ML) and deep learning (DL) models have emerged as more sophisticated alternatives to EWS tools. These models integrate a wide array of structured data elements from electronic health records (EHRs), including laboratory values, vital signs, medications, and demographic characteristics, to learn complex temporal and nonlinear patterns predictive of clinical deterioration.11–17 Models such as eCART, a gradient boosted machine learning system validated across multiple institutions,13 and the Rothman Index (RI),17 have demonstrated improvements in predictive accuracy and lead time over traditional tools. However, many of these models define deterioration based on endpoints such as cardiac arrest, ICU transfer, or in-hospital mortality. These event definitions are heavily influenced by local hospital policies, care practices, and clinical decision-making, which ultimately reduce their generalizability across different institutions.

A recently developed deep learning model, Deep Learning Enhanced Triage and Emergency Response for Inpatient Optimization (DETERIO), addresses these limitations by incorporating the Adult Inpatient Decompensation Event (AIDE) criteria, a consensus-driven, multi-event definition of physiological deterioration.18–19 Unlike traditional models that frame deterioration as a binary outcome, DETERIO employs a temporal difference (TD) learning approach to estimate a patient’s evolving state value over time, enabling dynamic assessment of clinical trajectories throughout hospitalization.20 While DETERIO improves upon traditional methods, its reliance on structured EHR data restricts its capacity to fully leverage nuanced observations, clinician insights, subjective assessments, and contextual information documented within unstructured clinical notes.

Recent advances in large language models (LLMs) such as Claude and GPT have shown promise in addressing this limitation. LLMs are capable of extracting and reasoning through complex narrative text, enabling more nuanced interpretation of patient status from clinical notes. A study comparing different machine learning models for sepsis prediction demonstrated that approaches integrating both unstructured and structured data achieve earlier and more accurate predictions compared to those using structured data alone.21 Similarly, multimodal deep learning models that combined structured and textual data outperformed models reliant on only structured data for predicting heart failure and other conditions.22–23 A recent study on sepsis prediction showed that including data extracted from clinical notes with LLMs in combination with traditional ML models can significantly improve prediction accuracy, particularly for patients with borderline risk scores.24

In this study, we introduce a new framework called DETERIO-LLM, which builds upon the existing DETERIO model by incorporating unstructured clinical notes to improve prediction accuracy in high-uncertainty cases. DETERIO-LLM uses an LLM architecture that processes longitudinal clinical documentation up to the time of an alert and classifies each case as high risk, borderline, or stable, based on a prompt aligned with the AIDE criteria. Our approach selectively targets cases with DETERIO scores in the borderline range (eg, 2.5 to 3.5), where the base model exhibits relatively low positive predictive values (PPVs). This optimizes computational efficiency and reduces false-positive alerts in ambiguous cases, thus improving clinical relevance.

To contextualize the performance of DETERIO-LLM, we benchmarked it against several established deterioration prediction tools. These comparators include the baseline DETERIO model, proprietary and ML-based systems like the Epic Deterioration Index (EDI), RI, and eCART, and traditional scoring methods such as the NEWS, NEWS2, and MEWS. The selection of these comparators was guided by a recent large-scale, multicenter study by Edelson et al.25 That study found that among these systems, eCART achieved the highest overall predictive performance for identifying clinical deterioration, reinforcing its value as the primary machine learning benchmark in our analysis.

Following the methodology established by Edelson et al.,25 2 standardized risk thresholds were applied for comparison: moderate-risk and high-risk. The moderate-risk trigger was defined with a sensitivity closest to that of a NEWS score of 5, and the high-risk trigger had a specificity closest to that of a NEWS score of 7. By comparing DETERIO-LLM against these well-established methods and standardized thresholds, we aim to demonstrate the utility of augmenting an existing ML model with LLM-based reasoning over clinical narratives for significantly improved and more nuanced deterioration detection.

Materials and methods

Study cohort

The baseline deep learning-based DETERIO model was developed and validated in a retrospective cohort study using de-identified EHR data from adult patients (≥18 years) admitted to inpatient services or the emergency departments (EDs) at 2 hospitals within the University of California San Diego Health System from January 1, 2016 to October 31, 2022. To assess the impact of incorporating unstructured clinical notes using LLMs, we constructed a computationally feasible validation subset of 1000 patients from the original DETERIO validation cohort. Given the high token and computational requirements of processing longitudinal narrative notes with an LLM, full-cohort evaluation was not feasible. Because DETERIO-LLM used a pretrained large language model (Claude 3.5 Sonnet) applied through zero-shot prompting, no additional model training or parameter optimization was performed on the study cohort. The 1000-patient subset therefore served exclusively as an independent validation cohort for evaluating LLM-enhanced prediction performance.

We used a stratified random sampling strategy to preserve representativeness while fixing the event prevalence at 4.6%, matching the deterioration rate reported by Edelson et al. Specifically, all patients meeting the deterioration criteria were identified, and a random sample of non-deteriorating patients was selected to achieve the target prevalence. This approach enabled meaningful benchmarking of prevalence-sensitive metrics such as positive predictive value (PPV) and F1 score while maintaining unbiased estimates of sensitivity and specificity within the sampled cohort.

Patients were excluded if their care unit length of stay was less than 2 h or if physiological deterioration occurred before hour 2 of care unit admission. Patients receiving comfort measures only and patients in procedure suites (eg, catheterization lab, gastrointestinal suite) and obstetrics units were additionally excluded. Patients were excluded if they had no available ED provider, history and physical, or progress note prior to their first DETERIO alert. Predictions began 2 h after admission and were updated hourly until deterioration or unit transfer. Table 1 presents the baseline characteristics of the study cohort.

Table 1.

Patient characteristics.

Number of encounters 1000
Age (in years), median [IQR] 60.0 (45.4-72.1)
Male gender (%) 53.2%
White (%) 48.8%
African American (%) 9.8%
Asian (%) 6.2%
Length of stay (in h), median [IQR] 81.9 (41.2-157.1)
SOFA, median [IQR] 1 (1-2)
AIDE cardiovascular—pressors/inotropes (%) 2.1%
AIDE respiratory (%) 3.2%
AIDE hemorrhage (%) 0.1%
In-hospital mortality (%) 0.8%
Transfer to ICU (%) 2.9%

Abbreviations: AIDE=Adult Inpatient Decompensation Event, IQR=interquartile range.

ED provider notes capture the patient’s initial presentation, evaluation, treatment, and discharge or admission plan. History and physical (H&P) notes document the patient’s assessment and plan once admitted and include relevant medical history, family history, social history, allergies, and medications. Progress notes document daily or periodic updates, tracking changes in condition and responses to treatment. These note types typically include physical examination findings as part of their standard clinical structure, although the presence of physical examination documentation was not used as a formal inclusion criterion.

Baseline DETERIO model

The baseline DETERIO model, a 3-layer feedforward neural network (with layer sizes of 100, 80, and 64), was trained using a temporal difference (TD) learning approach. Its objective was to predict a patient’s future state value from the second hour of admission until either physiological deterioration (T0), transfer to another care unit, or up to 14 days in control patients. The dataset comprised 50 variables representing vital signs and laboratory measurements, along with 6 demographic, 11 medication, and 62 comorbidity features. Physiological deterioration was defined by the Adult Inpatient Decompensation Event (AIDE) consensus criteria,18 which established a composite deterioration score. This score aggregated points from 4 distinct adverse events: the need for vasopressors or inotropes (3 points), severe hypoxemia/invasive mechanical ventilation (3 points), ICU transfer (1 point), and in-hospital mortality (1 point). Each event contributed to a binary score, allowing the overall composite deterioration score (or “patient state value”) to range from 0 to 8. For evaluation, a score of ≥3 was considered a positive (deteriorated) class, triggered by the fulfillment of at least 1 AIDE criterion, while a score of <3 constituted the control class. Model predictions were assessed against a threshold calibrated to achieve 50% sensitivity at the encounter level. A predicted risk score exceeding this threshold indicated anticipated physiological deterioration within a 12-h prediction window prior to T0, whereas scores below the threshold suggested no such deterioration.19

Label refinement based on clinical chart review

Prior to final model evaluation, we conducted a targeted chart review within the 1000-patient validation cohort to ensure that the ground-truth deterioration label adequately captured clinically meaningful physiological instability. We focused on a small subset of discordant cases—patients classified as high risk by both the baseline DETERIO model and DETERIO-LLM but labeled as false-positives under the original AIDE criteria (ie, total score <3). This review revealed that several patients exhibited significant clinical instability, most notably severe hypotension or active hemorrhage, but had not yet triggered any of the specific events (eg, vasopressor initiation or ICU transfer) required to meet the AIDE definition. Based on these findings, and in consultation with clinical collaborators, we refined the labeling schema of the baseline DETERIO model cohort to incorporate 2 additional indicators of physiological instability: (1) hypotension, incorporated into the vasopressor/inotrope logic, and (2) hemorrhage, added as a standalone criterion worth 3 points.

These changes expanded the possible deterioration score range from 0-8 to 0-11. The label adjustment led to a modest increase in the number of positive cases. Following finalization of the revised labeling framework, the entire patient cohort was relabeled accordingly, and the baseline DETERIO model was retrained using the updated outcome definition. Final performance metrics for both DETERIO and DETERIO-LLM reflect evaluation under this refined labeling schema. To ensure comparability across models and avoid artificial inflation of prevalence-sensitive metrics (eg, PPV), we maintained a constant event prevalence of 4.6% through adjusted control sampling. This prevalence was used consistently across all model comparisons, including DETERIO, DETERIO-LLM, and external benchmark models reported in Table 4.

Table 4.

Comparison with existing early warning scores.

Risk level Model Threshold Sensitivity Specificity PPV F1 score
Moderate DETERIO-LLM 2.5 52.3 92.9 26.1 34.8
eCART 94 52.2 97.2 17.3 25.9
NEWS 5 51.5 94.5 9.5 16.0
NEWS2 6 50.1 94.2 8.7 14.8
MEWS 3 43.7 94.6 8.2 13.8
RI 41 51.2 92.3 6.9 12.2
EDI 41 50.7 91.5 6.3 11.2
High DETERIO-LLM 3 47.8 94.8 30.6 37.3
eCART 97 42.5 98.4 23.3 30.1
NEWS 7 31.5 98.5 19.1 23.8
NEWS2 8 30.5 98.2 15.8 20.8
MEWS 4 27.6 98.5 16.9 21.0
RI 24 27.9 98.4 16.6 20.8
EDI 54 24.2 98.4 14.5 18.1

DETERIO-LLM model

To enhance the predictive capabilities of the baseline DETERIO model, particularly for cases with high uncertainty, we integrated LLMs with the existing deep learning framework, termed DETERIO-LLM. While the original DETERIO model was developed with a primary decision threshold of 3 for predicting deterioration within a 12-h window, in our implementation we used 2 thresholds to support performance benchmarking: a threshold of 2.5 as the moderate-risk trigger and 3.0 as the high-risk trigger for DETERIO model outputs. This is in line with those used by Edelson et al.25 These thresholds helped us compare DETERIO-LLM directly with established benchmarks while also keeping the model clinically actionable.

For both thresholds, DETERIO-LLM was applied only to cases with baseline model outputs in the borderline range, scores between 2.5 and 3.5 for the 2.5 threshold, and scores between 3.0 and 3.5 for the 3.0 threshold. These cases represent high-uncertainty predictions where the baseline model exhibits lower PPV. For these borderline cases, the LLM was provided with the full text of clinical notes from the time of admission up to the prediction time. The clinical documents included “ED provider note,” “ED note,” “ED floor report,” “Interdisciplinary,” “H&P (History & Physical) note,” “Diagnostic Image report,” and “Progress note.” Progress notes included documentation from physicians, residents, and advanced practice providers. Nursing flowsheets and nursing-only notes were excluded in this iteration to prioritize note types containing higher-level clinical synthesis, assessment, and treatment planning. The statement that these notes include physical examination findings reflects their typical structure in our EHR rather than a formal selection criterion.

We used Claude 3.5 Sonnet, a 175-billion-parameter LLM with a 200 000-token context window for this task. The entire pipeline was implemented in a HIPAA-compliant Amazon Web Services (AWS) environment, utilizing an EC2 instance configured with secure authentication and permissions. This setup allowed for real-time communication with Claude 3.5 Sonnet, accessed through the Amazon Bedrock Converse API, which facilitated structured JSON output generation. All predictions were generated with a model temperature of 0.0 to minimize hallucinations and ensure deterministic responses. Figure 1 illustrates the schematic diagram of the DETERIO-LLM pipeline.

Figure 1.

Flow diagram illustrating the DETERIO-LLM pipeline for refining clinical deterioration alerts. Structured EHR data including vital signs, laboratory results, medications, comorbidities, and demographics are input into the DETERIO prediction model. Cases with a risk score above a higher threshold (θ₂) generate an immediate alarm, while cases with risk scores between lower and upper thresholds (θ₁ ≤ risk score ≤ θ₂) are classified as borderline cases. For borderline cases, clinical notes available up to the alert time are analyzed by a large language model (Claude 3.5 Sonnet) deployed through AWS Bedrock. The LLM produces a risk assessment (high risk, borderline, or stable), along with a rationale and justification. Cases assessed as high risk generate an alarm, whereas cases assessed as borderline or stable do not generate an alarm. The figure demonstrates how the LLM is used to refine alerts for borderline predictions from the DETERIO model.

Overview of the DETERIO-LLM pipeline.

DETERIO-LLM risk classification approach

The DETERIO-LLM framework represents a direct application of an LLM to classify patient deterioration risk based solely on unstructured clinical notes. For each eligible alert within the borderline-risk range, Claude 3.5 Sonnet reviewed all available notes up to the alert timestamp. The model was prompted to determine whether the patient would meet criteria for physiological deterioration within the next 12 h, following the Adult Inpatient Decompensation Event (AIDE) composite score framework. Prompt development and refinement were conducted using a separate set of pilot cases drawn from the original DETERIO dataset but excluded from the 1000-patient validation cohort. No patients from the validation cohort were used during prompt engineering to prevent data leakage or prompt overfitting.

To standardize LLM behavior, a Chain-of-Thought (CoT) prompting strategy was employed. The prompt guided the model to consider recent changes in vital signs, lab abnormalities, medication trends, treatment responses, and physician impressions. Special emphasis was placed on synthesizing both acute and subacute trends over the most recent 24-48 h. The prompt also explicitly listed the AIDE criteria and instructed the model to make predictions aligned with the DETERIO outcome definition (ie, a composite score of ≥3). The LLM returned structured JSON outputs of the form:

{“risk_assessment”: “<High Risk|Borderline|Stable>,” “rationale”: “<concise justification>”}

For evaluation purposes, High Risk classifications were treated as positive predictions, indicating a projected deterioration event, while Borderline and Stable classifications were considered negative. This LLM-based reclassification provided a mechanism for refining ambiguous predictions made by the DETERIO model.

Results

Patient-level performance

At a threshold of 2.5, the baseline DETERIO model identified 25 true positives (TPs) and 127 false-positives (FPs), resulting in a sensitivity of 54.3%, specificity of 86.7%, and a PPV of 16.4%. When enhanced with DETERIO-LLM, performance improved substantially; FPs were reduced (−46.5%; 68 vs 127 FPs) while identifying all but 1 TP case (−4%; 24 vs 25 TPs). This resulted in improved specificity from 86.7% to 92.9% and PPV of 16.4% to 26.1%, with only a small reduction in sensitivity from 54.3% to 52.3%.

At a threshold of 3.0, the baseline DETERIO model identified 23 TPs and 79 FPs, with a sensitivity of 50.0%, specificity of 91.7%, and PPV of 22.5%. Following LLM reevaluation, 29 of 79 FPs (36.7%) were removed, reducing unnecessary alerts. One TP (4.3%) was incorrectly reclassified as negative. Post-LLM specificity an increased to 94.8% and PPV to 30.6%, with a small decrease in sensitivity to 47.8% (Tables 2 and 3).

Table 2.

Patient-level evaluation metrics.

Sensitivity % Specificity % PPV % F1 score %
Threshold = 2.5
DETERIO 54.3 86.7 16.4 25.3
DETERIO-LLM 52.3 92.9 26.1 34.8
Threshold = 3
DETERIO 50.0 91.7 22.5 31.1
DETERIO-LLM 47.8 94.8 30.6 37.3

Table 3.

DETERIO-LLM impact on FP and TP patients reductions by risk score range.

Model FP reduction (%) TP reduction (%)
Threshold = 2.5-3.5 46.5% 4%
Threshold = 3-3.5 36.7% 4.3%

Comparison with existing early warning scores

Table 4 presents the performance of DETERIO-LLM alongside published performance metrics for established early warning systems, including eCART, across 4 key metrics: sensitivity, specificity, PPV, and F1 score. These comparisons are intended to provide contextual benchmarking rather than direct head-to-head evaluation, as the referenced models were assessed in different patient populations and under potentially different outcome definitions.

At the moderate-risk threshold, DETERIO-LLM achieved the highest F1 score (34.8%), driven by a strong balance between PPV (26.1%) and sensitivity (52.3%), outperforming all other models in both predictive precision and clinical relevance.

At the high-risk threshold, DETERIO-LLM again outperformed all benchmarks, achieving the highest F1 score (37.3%) and PPV (30.6%). In contrast, traditional scoring systems demonstrated limited sensitivity and lower overall predictive balance.

Discussion

In this study, we demonstrate that integrating unstructured clinical notes using an LLM can significantly improve the predictive performance of an existing deep learning–based early warning system for inpatient clinical deterioration. Our model, DETERIO-LLM, builds upon the baseline DETERIO framework by selectively reevaluating borderline-risk predictions with an LLM that interprets a wide range of clinical notes. This targeted approach enables more nuanced, context-aware assessments of patient stability, improving PPV and overall F1 score without compromising sensitivity. These findings are consistent with prior research showing that LLMs can extract valuable clinical signals from narrative documentation to enhance risk prediction in critical care settings.21–24

A key strength of our approach is its hybrid architecture, in which the LLM is applied only to high-uncertainty predictions. By restricting LLM inference to borderline DETERIO scores, the system efficiently focuses computational resources where they are most clinically and operationally impactful. This design mitigates the inefficiency, increased latency, and potential overfitting associated with universal LLM use. Our findings support prior studies suggesting that integrating structured model outputs with LLM-based contextual reasoning yields superior performance compared with using LLMs alone.24,26,27,28

When evaluated on a cohort with a 4.6% deterioration prevalence, selected to mirror the population in Edelson et al.,25 DETERIO-LLM consistently outperformed traditional early warning scores (NEWS and MEWS) as well as ML-based models such as eCART, RI, and EDI on key metrics such as PPV and F1 score. At a moderate-risk threshold of 2.5, DETERIO-LLM achieved a PPV of 26.1% and an F1 score of 34.8%, markedly higher than those of all comparator methods. At the high-risk threshold of 3, it maintained a PPV of 30.6% and an F1 score of 37.3%, again exceeding the performance of established EWS models, including eCART, which was the top-performing model in the original Edelson study.

In addition to computational performance gains, DETERIO-LLM yielded a substantial reduction in FPs among borderline-risk predictions. At the 2.5-3.5 score range, the model reduced FP cases by 46.5% while removing only 4% of true positives, striking a favorable balance between sensitivity and specificity. This reduction is particularly important in real-world deployments, where excessive false alarms can lead to alert fatigue, desensitization, and missed opportunities for timely intervention.9–10

In interpreting these results, it is important to recognize the study’s limitations. The quality of LLM output is inherently tied to the completeness and recency of clinical documentation. Notes that are sparse, outdated, or lack meaningful narrative structure may hinder the LLM’s ability to draw reliable inferences. However, the integration of ambient scribes is transforming the semantic richness and temporal fidelity of clinical documentation.

Additionally, although Table 4 provides contextual benchmarking against previously published early warning systems, the absence of a direct head-to-head comparison within the same patient cohort using identical outcome definitions limits definitive comparative conclusions. Differences in reported performance may reflect variations in patient populations, threshold selection, and target label definitions across studies. Accordingly, these cross-study comparisons should be interpreted cautiously. Prospective validation and direct cohort-level comparisons will be important next steps in establishing relative performance.

Finally, while we used fixed score thresholds to determine when to integrate LLM evaluation, adaptive strategies based on uncertainty assessment and risk factors such as age and patient frailty could offer more dynamic and efficient augmentation.

In conclusion, DETERIO-LLM represents a scalable, interpretable, and effective enhancement to early warning systems by incorporating narrative clinical context into risk prediction. By focusing on ambiguous cases and delivering structured, justifiable outputs, this hybrid method addresses key limitations of both traditional scoring systems and deep learning models. Our findings support further investigation of LLM-augmented clinical decision support systems, particularly in settings where timely recognition of deterioration is essential to patient safety and resource optimization.

Conclusion

This study demonstrates that integrating LLMs with a validated DL-based deterioration prediction system can significantly enhance predictive accuracy and clinical utility, particularly in high-uncertainty cases. Our hybrid model, DETERIO-LLM, improves upon the base DETERIO model by selectively applying LLM reassessment to borderline-risk predictions, leveraging the rich contextual information found in unstructured clinical notes. By focusing LLM integration on cases with intermediate DETERIO scores, the model achieved substantial reductions in FPs while maintaining high sensitivity, offering a favorable trade-off between specificity and early detection. Compared to traditional early warning systems and ML-based scores such as NEWS, MEWS, RI, eCART, and EDI, DETERIO-LLM demonstrated superior performance across multiple evaluation metrics, including positive predictive value and F1 score. Future work will focus on prospective validation in real-time clinical settings, further optimization of LLM prompts and thresholds, and the exploration of adaptive triage mechanisms to maximize impact while ensuring efficiency, interpretability, and clinician trust.

Contributor Information

Ghodsieh Ghanbari, Department of Biomedical Informatics, University of California San Diego (UCSD) School of Medicine, La Jolla, CA 92093, United States.

Joseph C Ahn, Division of Gastroenterology & Hepatology, Mayo Clinic, Rochester, MN 55905, United States.

Eileen Kim, Department of Emergency Medicine, University of California San Diego School of Medicine, La Jolla, CA 92093, United States.

Ruth Laverde, Department of Emergency Medicine, University of California San Diego School of Medicine, La Jolla, CA 92093, United States.

Zach Pope, Department of Emergency Medicine, University of California San Diego School of Medicine, La Jolla, CA 92093, United States.

Shamim Nemati, Department of Biomedical Informatics, University of California San Diego (UCSD) School of Medicine, La Jolla, CA 92093, United States.

Author contributions

Ghodsieh Ghanbari (Conceptualization, Data curation, Formal analysis, Methodology, Project administration, Software, Validation, Visualization, Writing—original draft, Writing—review & editing), Joseph C. Ahn (Formal analysis, Investigation, Methodology, Validation, Writing—review & editing), Eileen Kim (Formal analysis, Investigation, Validation, Writing—review & editing), Ruth Laverde (Formal analysis, Investigation, Validation, Writing—review & editing), Zach Pope (Formal analysis, Investigation, Validation, Writing—review & editing), and Shamim Nemati (Conceptualization, Formal analysis, Funding acquisition, Investigation, Methodology, Project administration, Resources, Software, Supervision, Validation, Writing—review & editing)

Funding

G.G. acknowledges grant funding from the National Library of Medicine (award number T15LM011271). S.N. acknowledges grant funding from the National Library of Medicine (award number R01LM013998) and the National Institute of General Medical Sciences (award number R35GM143121).

Conflicts of interest

S.N. is a co-founder of Clairyon Inc., a predictive analytics startup. The terms of this arrangement have been reviewed and approved by the UC San Diego in accordance with its conflict-of-interest policies. All other authors declare no competing interest.

Data availability

No public repository exists for the study data, as it contains protected health information. Further inquiries can be directed to the corresponding author.

References

  • 1. Bapoje SR, Gaudiani JL, Narayanan V, Albert RK.  Unplanned transfers to a medical intensive care unit: causes and relationship to preventable errors in care. J Hosp Med. 2011;6:68-72. 10.1002/jhm.812 [DOI] [PubMed] [Google Scholar]
  • 2. Churpek MM, Wendlandt B, Zadravecz FJ, Adhikari R, Winslow C, Edelson DP.  Association between intensive care unit transfer delay and hospital mortality: a multicenter investigation. J Hosp Med. 2016;11:757-762. 10.1002/jhm.2630 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 3. Delgado MK, Liu V, Pines JM, Kipnis P, Gardner MN, Escobar GJ.  Risk factors for unplanned transfer to intensive care within 24 hours of admission from the emergency department in an integrated healthcare system. J Hosp Med. 2013;8:13-19. 10.1002/jhm.1979 [DOI] [PubMed] [Google Scholar]
  • 4. Escobar GJ, Greene JD, Gardner MN, Marelich GP, Quick B, Kipnis P.  Intra-hospital transfers to a higher level of care: contribution to total hospital and intensive care unit (ICU) mortality and length of stay (LOS). J Hosp Med. 2011;6:74-80. 10.1002/jhm.817 [DOI] [PubMed] [Google Scholar]
  • 5. Kause J, Smith G, Prytherch D, Parr M, Flabouris A, Hillman K; Australian and New Zealand Intensive Care Society Clinical Trials Group. A comparison of antecedents to cardiac arrests, deaths and emergency intensive care admissions in Australia and New Zealand, and the United Kingdom—the ACADEMIA study. Resuscitation. 2004;62:275-282. 10.1016/j.resuscitation.2004.05.016 [DOI] [PubMed] [Google Scholar]
  • 6. Liu V, Kipnis P, Rizk NW, Escobar GJ.  Adverse outcomes associated with delayed intensive care unit transfers in an integrated healthcare system. J Hosp Med. 2012;7:224-230. 10.1002/jhm.964 [DOI] [PubMed] [Google Scholar]
  • 7. Smith AF, Wood J.  Can some in-hospital cardio-respiratory arrests be prevented? A prospective survey. Resuscitation. 1998;37:133-137. 10.1016/S0300-9572(98)00056-2 [DOI] [PubMed] [Google Scholar]
  • 8. Jones D, Mitchell I, Hillman K, Story D.  Defining clinical deterioration. Resuscitation. 2013;84:1029-1034. 10.1016/j.resuscitation.2013.01.013 [DOI] [PubMed] [Google Scholar]
  • 9. Khalifa M, Zabani I.  Improving utilization of clinical decision support systems by reducing alert fatigue: strategies and recommendations. Stud Health Technol Inform. 2016;226:51-54. [PubMed] [Google Scholar]
  • 10. Patterson E.  Navigating alert fatigue: a case study in electronic health record alert design optimization. Stud Health Technol Inform. 2024;315:447-451. 10.3233/SHTI240188 [DOI] [PubMed] [Google Scholar]
  • 11. Kipnis P, Turk BJ, Wulf DA, et al.  Development and validation of an electronic medical record-based alert score for detection of inpatient deterioration outside the ICU. J Biomed Inform. 2016;64:10-19. 10.1016/j.jbi.2016.09.013 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12. Steitz BD, McCoy AB, Reese TJ, et al.  Development and validation of a machine learning algorithm using clinical pages to predict imminent clinical deterioration. J Gen Intern Med. 2024;39:27-35. 10.1007/s11606-023-08349-3 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 13. Churpek MM, Carey KA, Snyder A, et al.  Multicenter development and prospective validation of eCARTv5: a gradient-boosted machine-learning early warning score. Crit Care Explor. 2025;7:e1232. 10.1097/CCE.0000000000001232 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 14. Escobar GJ, LaGuardia JC, Turk BJ, Ragins A, Kipnis P, Draper D.  Early detection of impending physiologic deterioration among patients who are not in intensive care: development of predictive models using data from an automated electronic medical record. J Hosp Med. 2012;7:388-395. 10.1002/jhm.1929 [DOI] [PubMed] [Google Scholar]
  • 15. Escobar GJ, Liu VX, Schuler A, Lawson B, Greene JD, Kipnis P.  Automated identification of adults at risk for in-hospital clinical deterioration. N Engl J Med. 2020;383:1951-1960. 10.1056/NEJMsa2001090 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16. Churpek MM, Yuen TC, Edelson DP.  Predicting clinical deterioration in the hospital: the impact of outcome selection. Resuscitation. 2013;84:564-568. 10.1016/j.resuscitation.2012.09.024 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 17. Rothman MJ, Rothman SI, Beals J.  Development and validation of a continuous measure of patient condition using the electronic medical record. J Biomed Inform. 2013;46:837-848. 10.1016/j.jbi.2013.06.011 [DOI] [PubMed] [Google Scholar]
  • 18. Mitchell OJL, Dewan M, Wolfe HA, et al.  Defining physiological decompensation: an expert consensus and retrospective outcome validation. Crit Care Explor. 2022;4:e0677. 10.1097/CCE.0000000000000677 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19. Shashikumar SP, Le JP, Yung N, et al.  Development and validation of a deep learning model for prediction of adult physiological deterioration. Crit Care Explor. 2024;6:e1151. 10.1097/CCE.0000000000001151 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 20. Sutton RS.  Learning to predict by the methods of temporal differences. Mach Learn. 1988;3:9-44. 10.1007/BF00115009 [DOI] [Google Scholar]
  • 21. Yan MY, Gustad LT, Nytrø Ø.  Sepsis prediction, early detection, and identification using clinical text for machine learning: a systematic review. J Am Med Inform Assoc. 2022;29:559-575. 10.1093/jamia/ocab236 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22. Gao Z, Liu X, Kang Y, et al.  Improving the prognostic evaluation precision of hospital outcomes for heart failure using admission notes and clinical tabular data: multimodal deep learning model. J Med Internet Res. 2024;26:e54363. 10.2196/54363 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 23. Wang Y, Yin C, Zhang P.  Multimodal risk prediction with physiological signals, medical images and clinical notes. Heliyon. 2024;10:e26772. 10.1016/j.heliyon.2024.e26772 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 24. Shashikumar SP, Mohammadi S, Krishnamoorthy R, et al.  Development and prospective implementation of a large language model based system for early sepsis prediction. Npj Digit Med. 2025;8:1-10. 10.1038/s41746-025-01689-w [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 25. Edelson DP, Churpek MM, Carey KA, et al.  Early warning scores with and without artificial intelligence. JAMA Netw Open. 2024;7:e2438986. 10.1001/jamanetworkopen.2024.38986 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 26. Hager P, Jungmann F, Holland R, et al.  Evaluation and mitigation of the limitations of large language models in clinical decision-making. Nat Med. 2024;30:2613-2622. 10.1038/s41591-024-03097-1 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 27. Ullah E, Parwani A, Baig MM, Singh R.  Challenges and barriers of using large language models (LLM) such as ChatGPT for diagnostic medicine with a focus on digital pathology–a recent scoping review. Diagn Pathol. 2024;19:43. 10.1186/s13000-024-01464-7 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 28. Williams CYK, Miao BY, Kornblith AE, Butte AJ.  Evaluating the use of large language models to provide clinical recommendations in the emergency department. Nat Commun. 2024;15:8236. 10.1038/s41467-024-52415-1 [DOI] [PMC free article] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Data Availability Statement

No public repository exists for the study data, as it contains protected health information. Further inquiries can be directed to the corresponding author.


Articles from JAMIA Open are provided here courtesy of Oxford University Press

RESOURCES