Abstract
Clinical decision-making is inherently complex and fast-paced, particularly in emergency departments (EDs) where rapid and high-stakes decisions are made. Clinical Decision Rules (CDRs) are standardized evidence-based tools that combine signs, symptoms, and clinical variables into decision trees to make consistent and accurate diagnoses. CDR usage is often hindered by the clinician’s cognitive load, limiting their ability to quickly recall and apply the appropriate rules. We introduce CDR-Agent, a novel LLM-based system designed to enhance ED decision-making by autonomously identifying and applying the most appropriate CDRs based on unstructured clinical notes. To validate CDR-Agent, we curated two novel ED datasets: synthetic and CDR-Bench, although CDR-Agent is applicable to non ED clinics. CDR-Agent achieves a 56.3% (synthetic) and 8.7% (CDR-Bench) accuracy gain relative to the standalone LLM baseline in CDR selection, with overall prediction accuracy improvements of 134.0% (synthetic) and 20.4% (CDR-Bench). Moreover, CDR-Agent significantly reduces computational overhead.
Introduction
Clinical Decision Rules (CDRs) are standardized tools designed to assist clinicians by combining signs, symptoms, and clinical features into decision trees, enabling accurate and consistent, evidence-based bedside decisions.1 Currently, over 700 CDRs span nearly all medical specialties, covering diverse clinical scenarios.2, 3 These rules utilize patient data to generate composite scores or assessments that stratify patient risk for disease onset, progression, or clinical outcomes, helping clinicians deliver expert-level care regardless of their level of experience.4, 5
However, across healthcare settings, clinicians face increasing time pressure and cognitive burden when evaluating patients. Whether in outpatient clinics, inpatient wards, or telemedicine encounters, clinicians must rapidly assess complex presentations with limited time and incomplete information. In these environments, tools that can surface the right decision aid at the right time can help mitigate errors and democratize care quality. This challenge is even more pronounced in high-stakes trauma care, where rapid and accurate decisions are essential to avoid the harms associated with missed critical injuries or unnecessary imaging and interventions.6 Although trauma care has been regionalized to concentrate expertise, injured patients often initially present to emergency departments (EDs) without trauma specialization, further emphasizing the critical need for universally applicable decision-making aids like CDRs.7, 8
Despite the importance of CDRs, clinicians often face significant challenges recalling and applying them appropriately due to the sheer number of rules tailored to specific clinical conditions, organ systems, and patient populations.9 This challenge is not confined to emergency care – clinicians across healthcare settings routinely make rapid decisions under time constraints, often with incomplete data. Whether in prehospital, outpatient, or inpatient medicine, clinicians must balance efficiency with accuracy, making real-time access to relevant CDRs critically important.
Recent advancements have highlighted the integration of Large Language Models (LLMs) into Clinical Decision Support Systems (CDSS), significantly enhancing clinical decision-making processes. LLMs augmented with external medical knowledge, such as literature databases and clinical guidelines, demonstrate superior performance compared to standalone models in various clinical tasks, including COVID-19 outpatient care, medication prescriptions, and diagnostic accuracy.10, 11, 12 Additionally, explanations generated by LLMs based on patient notes have been shown to improve clinician agreement rates, emphasizing their potential utility in clinical contexts.13 Nevertheless, existing approaches often suffer from limited generalizability due to their application within narrow, specialized healthcare settings.14 Particularly relevant to our work is the study by Zakka et al., which introduced an LLM-based agent with external internet access and an extensive database of clinical calculators derived from MDCalc, coupled with a ClinicalQA dataset, achieving substantial performance improvements over plain LLMs.15
In this paper, we focus on trauma care as a case study – not because it is exclusive to the ED, but because it uniquely spans multiple disciplines and involves numerous CDRs across different organ systems and clinical specialities (e.g. trauma surgery, orthopedics, critical care, radiology, etc.). This allows us to explore whether an LLM-based agent can support holistic clinical reasoning by recognizing when and how to apply a diverse set of CDRs within a single patient scenario – a challenge that remains largely unmet in current AI systems.
Compared to existing research focused on automated diagnosis,16, 17, 18, 19, 20 knowledge-intensive medical question answering,21, 22, 23 clinical data analysis,24 and medication management,11 emergency department decision-making uniquely prioritizes accuracy and efficiency amidst challenges posed by incomplete patient information. Prior studies that augment LLMs with external medical knowledge typically employ retrieval-augmented generation (RAG) methods, which select relevant resources to contextualize the LLM’s responses. Although these methods improve accuracy and efficiency relative to standalone models, this alone is insufficient in emergency department scenarios. Existing RAG approaches commonly lack transparency regarding how the LLM integrates retrieved information or verifies the appropriateness of selected resources, limiting interpretability. This absence of clear rationale undermines clinical trust, as emergency clinicians rely heavily on transparent, interpretable decision-making to manage rapid, accurate patient care amidst incomplete and evolving clinical information.
As a remedy, we propose CDR-Agent in the “Method” section. CDR-Agent is an LLM-based system designed to autonomously select and apply the most suitable clinical decision rules (CDRs) based on given clinical notes in real ED situations. Its operation involves three primary steps, closely mirroring the decision-making process of clinicians. First, it identifies relevant CDRs by measuring semantic similarity between an input clinical note and available CDR descriptions. To maintain computational efficiency, this step leverages an embedding model without invoking LLM queries. The system then extracts specific indicator values from the clinical note, which serve as input variables for the selected CDRs followed by a verification of exclusion criteria to determine the CDRs valid to proceed. If certain necessary information is missing, clinicians may either collect the information or infer it using their clinical expertise. Finally, CDR-Agent precisely executes the selected CDRs by running the Python script corresponding to each of them to generate decision outcomes.
Due to the lack of public benchmarks for validating our CDR-Agent, we evaluate its effectiveness and efficiency on two curated ED datasets in the “Dataset Construction” section: CDR-Bench and a synthetic dataset derived from Pediatric Emergency Care Applied Research Network (PECARN).25, 26 The synthetic dataset provides a controlled evaluation setting, where each note is generated from PECARN tabular data and contains a single CDR ground truth label with no missing values. This ensures a structured assessment of model performance in unambiguous clinical scenarios. CDR-Bench consists of real-world clinical notes from various public datasets, annotated with clinician-labeled CDRs. These notes exhibit diverse writing styles, contain noise, and often have missing values, making CDR-Bench a challenging real-world robustness test. CDR-Bench features an adaptive number of CDRs per note, better reflecting real-world clinical decision-making complexity. Together, these datasets offer a comprehensive evaluation of CDR-Agent’s ability to interpret clinical narratives and apply trauma-related CDRs accurately. Both datasets are released with code and detailed documentation, ensuring reproducibility and facilitating further research in LLM-driven clinical decision support.
In the “Results and Findings” section, we evaluate CDR-Agent on the two datasets compared with the baseline approach that directly queries the same LLM for CDR execution results using the clinical notes and the list of available CDRs. CDR-Agent achieves 56.3% and 8.7% absolute gain in CDR selection accuracy1 on the synthetic data and CDR-Bench, respectively; while for CDR execution, CDR-Agent achieves comparable or better performance than the baseline on both datasets. Moreover, the entire procedure of CDR selection and execution costs 1.22 and 1.16 seconds on the synthetic data and CDR-Bench, respectively, much faster than the baseline with 4.12 and 8.7 seconds on the two datasets.
In summary, we propose CDR-Agent, the first automated, LLM-based system that assists with expert-level decision-making using CDRs in ED scenarios. We also created two datasets to facilitate research on ED decision-making involving CDR. Finally, our in-depth case studies underscore the potential to foster collaboration between clinicians and technicians in this field. The code for CDR-Agent can be found at https://github.com/zhenxianglance/medagent-cdr-agent.
Problem Description
To set up the task, we consider a clinical note x that describes patient information such as demographics, chief complaints, and physical exam results, and a set of predefined CDRs = {c1, ⋯, cN}. Each CDR ci is a decision tree function that maps from a specific variable space to a set of predefined outcomes. Here, the variables are the indicators in the clinical notes required by the CDR to arrive at a decision outcome. For example, the NEXUS criteria for C-spine imaging require five binary indicators: presence of focal neurologic deficit, midline spinal tenderness, altered level of consciousness, intoxication, and distracting injury, respectively, to determine an outcome that can either be “imaging recommended” (if any of the indicators present) or “imaging not necessary” (if no indicator is found), as shown in Figure 1. Our design objective is to build an automated system that, for any input clinical note x, identifies the most appropriate CDRs from , if there are any, and then execute the identified CDRs to get a set of final decisions.
Figure 1:
An example CDR for C-spine imaging. Top left: the variables/indicators required by the CDR and their definitions. Bottom left: the rule deciding whether the patient requires imaging. Right: the Python script for the CDR with automated imputation of missing variables.
Methods
In this section, we introduce our CDR-Agent framework proposed to address the problem described above, focusing on the design details and underlying rationales.
Overview Our system design takes into account real-world clinical scenarios with the following challenges. First, clinical notes often vary significantly in writing style and terms, making it difficult to accurately identify relevant indicators or variables for CDRs through simple word matching. Second, some critical information may be incomplete, particularly in cases involving unconscious, disabled, or pediatric patients, or cases involving progress notes2. Third, the execution of CDRs is prone to errors – both by humans, due to the complexity of the rules, and by automated systems such as LLMs, which may misinterpret textual rule descriptions and generate incorrect decisions.
CDR-Agent follows a three-step workflow designed to address these challenges, and we descibe the deatils of each step in the following paragraphs. The complete pipeline of CDR-Agent is shown in Figure 2.
Figure 2:
Illustration of the three-step workflow of CDR-Agent. For any input clinical note, CDR-Agent first selects a number of relevant CDRs that have high semantic similarity to the clinical note. Variables required by each selected CDR are then extracted from the clinical note using an LLM, with a set of exclusion rules applied to filter invalid CDRs. Finally, CDR-Agent executes the Python code for each valid CDR for decisions.
CDR Selection Step The goal of the first step is to identify the most relevant CDRs for a given clinical note. We use semantic similarity as a proxy to measure the relevance between the clinical note and CDRs. To measure semantic similarity, we first converted the CDR functions into textual descriptions using GPT-4o followed by clinician inspection to ensure the correctness, and used a pretrained word embedding model, text-embedding-ada-002 by OpenAI, to map both the clinical notes and the textual description of CDRs into a shared embedding vector space.
Formally, for each CDR ci, we compute an embedding3 vector E(x) for the clinical note x and an embedding vector E(ti) for the textual description ti of the CDR, respectively. Here, the embedding model is trained by OpenAI on a large corpus of text data, including medical literature, ensuring that semantically similar content is mapped to vectors with higher cosine similarity. A straightforward approach to select CDRs would be to compute the cosine similarity s(i) = cosim(E(x), E(ti)) for each CDR and select the top-k CDRs with the highest similarity scores for a predefined k. However, setting a value of k or a naive threshold on similarity scores for CDR selection is unrealistic and challenging in practice because of the variability in similarity score ranges across different clinical notes.
As a solution, we propose an anomaly detection approach to identify for each clinical note the CDRs with abnormally large similarity scores for selection in a principled way. Specifically, for each clinical note, we fit a Gaussian distribution using the computed similarity scores – the choice of Gaussian is validated in Section “Design Choices and Further Improvements”. A CDR ci is selected as an anomaly if p(s(i)|μ, σ2) > α, where μ and σ are the estimated mean and variance for the Gaussian distribution, and α is a predefined significance level. Drawing inspiration from hypothesis testing, we set α = 0.05 in all our experiments, which achieves an effective balance between minimizing false positives and enabling the detection of real outliers. If no CDR is identified as an anomaly, it indicates that none of the available CDRs are applicable to the given clinical scenario.
In addition to the above-mentioned base design, we will demonstrate in Section “Design Choices and Further Improvements” that multiple features can be added to the framework to further improve the CDR selection accuracy.
Variable Extraction Step CDR-Agent uses an LLM to extract the variables required by each selected CDR from the given clinical note, with a prompt shown in Figure 3. The key to this prompt design is the clear specification of variable definitions and formatting. For each selected CDR, the prompt includes the definition of each variable (see top-left of Figure 1 for example) to help the LLM accurately interpret its meaning while searching for the corresponding values from the clinical note. Additionally, the extracted variables are structured in a list format, where each variable name is followed by its corresponding extracted value.
Figure 3:
Prompt to LLM for variable extraction from a given clinical note. The prompt includes variables required by the selected CDR and their definitions, and the formatting requirements for the extracted variable values.
To handle potential missing values, we include only the determined variables in the output list. In practice, absent variables can be filled in by consulting clinicians for additional information collection from patients. However, in our experiments with automated evaluation settings, this feature (which requires human participation) is not incorporated. Instead, we adopt a “negative” imputation approach, where missing variables are assigned default values that do not trigger positive clinical decisions, such as imaging recommendations. This practice prevents the LLM from arbitrarily assigning values that could lead to false positives. For future work, we will explore alternative imputation strategies that leverage external knowledge, such as population-level measurements, to enhance accuracy of CDR execution. After extracting the variables, a set of exclusion rules, according to the inclusion/exclusion criteria for each CDR, is applied to filter out invalid CDRs, such as excluding adult patients for PECARN CDRs.
CDR Execution Step CDR-Agent executes each valid CDR by running a predefined Python script (see the right of Figure 2 for an example Python script), using the variable values extracted from the clinical note in step 2. Before execution, the extracted variable values are converted to their predefined data types, including boolean, integer, float, and string, ensuring consistency and correctness. Formally, the decision outcome for a selected CDR ci is obtained as , where are the variable values after formatting and fci is the CDR function. Note that any execution failure will trigger an error message, allowing the agent to prompt manual intervention to verify the information processed in earlier steps. This design ensures that only outputs based on accurate variable extraction are produced, minimizing potential errors caused by LLM hallucinations. Once all valid CDRs have been executed, their decisions are aggregated into a final list as the next-step management for the patient.
Dataset Construction
As the first study to explore the automation of CDR selection and execution in ED clinical scenarios, we develop two datasets to evaluate our framework and serve as a benchmark for future research. The first is a synthetic dataset focused on pediatric trauma cases, with ground truth CDR labels provided. The second is a more challenging real-world dataset comprising diverse clinical notes, where ground truth CDRs are annotated by clinical experts. While the synthetic dataset enables controlled analysis of system behavior, the real-world dataset assesses the CDR-Agent’s robustness in handling the complexity and variability of actual clinical documentation. Collectively comprising 544 patient notes, our dataset significantly surpasses prior related datasets by an order of magnitude.
CDR Database Based on clinician recommendations, we selected 15 trauma-related CDRs from peer-reviewed publications to form our CDR database (see Fig. 4 for a complete list). Each rule was manually transcribed from its original published format into a structured Python function, enabling automated execution within our system. To facilitate accurate interpretation and execution, we also provided detailed textual descriptions for each input feature and for the rule’s clinical purpose. This standardized approach ensures that models evaluated on this dataset have sufficient context to correctly select and apply relevant rules. The resulting CDR database was utilized for our experiments on both CDR-Bench and the synthetic dataset, demonstrating the generalizability of our method.
Figure 4:
(Left) A detailed breakdown of CDR label composition in CDR-Bench. Approximately 36.8% of the notes have no applicable CDRs. (Right) Token length variations across different data sources in CDR-Bench. MIMIC-IV notes are significantly longer than those from MedQA and ACN, often containing more noise and distracting information. This highlights both the diversity captured in CDR-Bench and the challenge it presents for CDR selection.
Synthetic Dataset PECARN provides Clinical Decision Rules (CDRs) along with corresponding tabular datasets for pediatric traumatic brain injuries (TBI) and pediatric intra-abdominal injuries (IAI). To systematically evaluate our system under controlled conditions, we generated a synthetic dataset by converting tabular patient data from PECARN into realistic free-form clinical narratives. Specifically, we randomly selected 400 patients (200 TBI and 200 IAI), of which 20% required medical intervention, ensuring representation of clinically significant scenarios. We first use templates to convert binary patient features into structured sentences, then refine these into coherent, natural-sounding clinical notes using GPT-4o. The resulting clinical narratives averaged approximately 98 tokens in length.
CDR-Bench CDR-Bench is constructed using real-world clinical notes from three well-established public datasets: MIMIC-IV,27 MedQA,28 and Augmented Clinical Notes (ACN),29 each featuring diverse writing styles. To ensure a balanced dataset with both CDR-applicable and non-applicable cases, we filter for notes containing trauma-related keywords (e.g., fracture, injury, trauma). This step prevents an overwhelming number of notes without relevant CDRs, given the large dataset sizes.
We randomly selected 50 trauma-related notes from each source for annotation by four emergency medicine clinicians. Initially, each sample was independently labeled by two clinicians. In cases of disagreement (with an average inter-annotator agreement of 80%), a third, more senior clinician reviewed the discrepancies and provided a final adjudication. Recognizing that clinical decision-making often lacks a single ground truth, we retain all label sets from the annotators. During evaluation, we measure the agent’s maximum CDR selection accuracy across these label sets – meaning CDR-Agent is considered correct if it aligns with any clinician’s judgment. This approach reflects the inherent variability in medical decision-making while ensuring a robust evaluation of CDR selection performance.
After notes with incomplete information to label were removed, following the clinicians’ feedback, the final CDR-Bench consists of 144 samples (MIMIC-IV: 47, MedQA: 47, ACN: 50). On average, each note contains 233.9 tokens. Approximately 36.8% of the notes have no applicable CDRs, ensuring a mix of relevant and irrelevant cases for robustness testing. Among the labeled samples, the average number of CDRs per note is 2.83, highlighting the multi-label nature of clinical decision-making. These facts together make CDR-Bench a challenging dataset compared to the synthetic dataset. See Figure 4 for a detailed breakdown of CDR label composition and token length variations across different data sources in CDR-Bench.
Experimental Setup
We aim to evaluate the proposed CDR-Agent with a focus on two research questions: First, can CDR-Agent accurately and efficiently select all relevant CDRs for a given clinical note? Second, with the incorporation of LLMs, can CDR-Agent accurately and efficiently extract the required variables from the complex clinical note and execute the CDRs reliably? Note again that efficiency here refers to the computational time cost of the system.
To answer these questions, we use two sets of metrics to assess CDR selection and execution outcomes, respectively. Note that the CDR labeling for a clinical note can result in a single CDR, a set of CDRs, or “no applicable CDR”. We evaluate CDR selection using three metrics: 1) exact match (EA) accuracy, which measures the proportion of clinical notes where the selected CDR(s) exactly match the labeled CDR(s); 2) F1-score, which jointly assesses the precision and recall for CDR selection4; and 3) selection time, which quantifies the computational cost of CDR selection in seconds. For CDR execution, we focus on sensitivity and specificity: 1) sensitivity, which measures the sensitivity of outcome recommendation across all correctly selected CDRs; 2) specificity, which measures the specificity across all correctly selected CDRs; and 3) execution time, which measures the average time cost for each CDR execution, including both variable extraction and execution of Python scripts of CDRs.
To illustrate the benefits of our design over conventional LLM querying, we compare CDR-Agent with a baseline approach where an LLM is directly queried using the clinical note alongside all the CDR information to determine the final CDR execution outcomes. We report the main results in Table 1, where we used one of the state-of-the-art LLMs, GPT-4o, as the core LLM for both approaches. Note that for some other LLM choices, such as GPT-4, the baseline approach easily encounters maximum token limitations due to the extensive prompt length required by the inclusion of all CDRs. This issue will be further amplified as the number of CDRs increases. However, our CDR-Agent is not affected by this issue as it leverages an embedding model to generate embedding vectors to compute similarity scores for CDR selection, without the need of squeezing all CDR information in-context in a query. In addition to the main results, we also evaluate CDR-Agent, particularly its CDR execution performance on other LLM choices (as model choices don’t affect CDR selection performance), including GPT-4 and GPT-4o-mini, and report the results in Table 2.
Table 1:
Evaluation results for CDR-Agent on the synthetic data and CDR-Bench, compared with the baseline approach purely based on LLM querying. Evaluation metrics including the exact match (EA) accuracy, F1-score, and time cost (Tsel) for CDR selection, and the sensitivity, specificity, and time cost (Texe) for CDR execution. The total time costs (Ttot) for both methods are also reported. All time costs are in second. “n.a.” means “not applicable”.
Table 2:
Comparison of various LLM choices for CDR-Agent in CDR execution.
| Synthetic Data | CDR-Bench | |||||
|---|---|---|---|---|---|---|
| GPT-4o | GPT-4o-mini | GPT-4 | GPT-4o | GPT-4o-mini | GPT-4 | |
| Sensitivity | 0.983 | 0.987 | 0.966 | 0.683 | 0.717 | 0.786 |
| Specificity | 0.687 | 0.583 | 0.711 | 0.983 | 0.967 | 1.0 |
| T exe | 0.78 | 1.29 | 3.05 | 0.64 | 0.75 | 1.89 |
Results and Findings
Main Evaluation Results The comparison between CDR-Agent and the baseline approach in CDR selection and execution is presented in Table 1. On the synthetic dataset, CDR-Agent achieves an exact match accuracy of 0.983 and an F1-score of 0.994, both significantly outperforming the baseline. For CDR execution, CDR-Agent achieves a sensitivity close to 1.0, along with a reasonable specificity of 0.687. In contrast, the baseline approach is more conservative, with a sensitivity of 0.818 and a specificity of 0.708. Furthermore, CDR-Agent is highly efficient, requiring only 1.22 seconds on average to complete the entire procedure, compared to 4.12 seconds for the baseline. On CDR-Bench, CDR-Agent achieves a superior exact match accuracy of 0.513 compared to 0.426 for the baseline in CDR selection, while maintaining a comparable F1-score. In terms of CDR execution, CDR-Agent is more conservative than the baseline, with an average sensitivity of 0.683 and an average specificity of 0.983. Moreover, this performance can be further enhanced when using GPT-4 as the core LLM, with sensitivity increasing to 0.786 and sensitivity to 1.0 and only one additional second in the time cost. In contrast, the baseline does not benefit from this alternative core model due to its limited scalability with the number of CDRs, constrained by the model’s input limitations. Regarding total time cost, CDR-Agent is more than seven times faster than the baseline, showing its high efficiency in computation. Notably, the time cost associated with CDR-Agent (around a second) remains significantly lower than that of human experts, who must read clinical notes, recall relevant CDRs, and make informed decisions, which may take seconds or even minutes. In summary, CDR-Agent shows comparable or better capabilities than the baseline in both CDR selection and CDR execution, with significantly lower time cost, which highlights its suitability in ED decision-making.
Influence of LLM Choices Table 2 presents the performance of CDR-Agent in CDR execution using different core LLMs. The results for CDR selection are omitted here, as this step does not involve querying an LLM. On the synthetic dataset, CDR-Agent achieves comparable sensitivity and specificity when using GPT-4o and GPT-4, with GPT-4 exhibiting slightly more conservative behavior. However, CDR-Agent with GPT-4o demonstrates significantly betetr efficiency, requiring only about one-quarter of the execution time compared to GPT-4. GPT-4o-mini achieves sensitivity comparable to GPT-4o but with noticeably lower specificity and increased execution time. On CDR-Bench, both the sensitivity and specificity are clearly improved when changing the core LLM from GPT-4o or GPT-4o-mini to GPT-4. These findings suggest that the performance of CDR-Agent may be further improved with larger models in the future (potentially with stronger capabilities).
Design Choices and Further Improvements In the CDR selection step, we model the similarity scores of irrelevant CDRs using a Gaussian distribution. To validate this choice, we analyze randomly sampled clinical notes by reviewing the Q-Q plots of similarity scores for all CDRs deemed irrelevant to each note. As illustrated by the example Q-Q plot on the left of Figure 5, the similarity scores align closely with the normal reference line (shown in red), showing that the Gaussian distribution is a reasonable choice. However, accurate estimation of the mean and variance of the Gaussian distribution requires a sufficient number of similarity scores. To improve the robustness of CDR selection, we apply repeated random truncations to clinical notes when computing similarity measures. This approach not only increases the number of samples (though dependent) for more reliable estimation of the distribution parameters, but may also reduce redundancy in the notes arising from verbose documentation. As shown on the right of Figure 5, we vary the number of random sampling iterations across [2,5,8,10,30,50] while testing different note retention ratios. This experiment is conducted on a held-out set consisting of 20% of clinical notes randomly sampled from CDR-Bench, using the same CDR-Agent settings as in our main experiments. Across all retention ratios, increasing the number of iterations generally improves CDR selection performance, albeit at the cost of longer computation times – still well within a range comparable to that of human clinical decision-making. Notably, when focusing on smaller segments of notes (lower note retention ratio), performance improves more sharply as the number of iterations increases. These results suggest that, for real-world CDRs with noisy or verbose documentation, applying a greater number of random truncations – particularly with shorter note segments – can lead to more robust CDR selection, should the time budget allow.
Figure 5:
(Left) An example Q-Q plot demonstrating that a Gaussian distribution is a reasonable choice for modeling similarity scores of irrelevant CDRs. (Right) Trade-off between F1-score and computation time on a held-out set of CDR-Bench for varying numbers of random sampling iterations and note retention ratios.
Incorporation of Additional Knowledge To account for the variability in real-world clinical notes, where medical terms often have alternative phrasings and abbreviations, we explore enhancing CDR-Agent with additional knowledge. Specifically, we augment the textual descriptions of CDRs by appending a list of synonymous terms for key indicators, prefixed with “Keywords to consider often include:”. For example, for distal radial fracture, we include alternatives such as DRF, fx distal radius, distal radius fracture, and for wrist swelling, we add swollen wrist, wrist puffiness, enlarged wrist. These keyword expansions were generated using GPT-4o with a one-shot example and later verified by a clinician. Incorporating this additional knowledge improved exact match accuracy by 4% and F1 score by 2% on CDR-Bench. This demonstrates CDR-Agent’s flexibility in integrating domain-specific enhancements. In future work, we plan to extend this capability of CDR-Agent by integrating a structured, extensible knowledge base into CDR-Agent, with validation from clinical experts. This knowledge base will include up-to-date clinical concepts, terminologies, semantic relationships, and representative examples, making it more robust or stable to linguistic variability in clinical text.
Conclusion
In this study, we introduced CDR-Agent, an LLM-based system designed to autonomously select and execute CDRs based on clinical notes. We created two high-quality datasets to address the lack of public benchmarks in validating automated CDR selection and execution approaches, thereby facilitating future research in LLM-driven clinical decision support. Moreover, we used these two datasets to demonstrate the superior performance of CDR-Agent in both accuracy and efficiency compared to traditional LLM prompting approaches.
While our framework is generalizable across clinical scenarios, we chose to focus on trauma care as a compelling case study due to its multidisciplinary nature and the availability of multiple, organ-specific CDRs. By bridging AI capabilities with clinical expertise, CDR-Agent has the potential to support decision-making not only in specialized trauma centers but also in diverse clinical settings where expertise may be limited. This work represents an important step toward deploying safe, transparent, and context-aware AI tools in frontline multidisciplinary trauma care.
Acknowledgments
We gratefully acknowledge partial support from NSF grant DMS-2413265, NSF grant DMS 2209975, NSF grant 2023505 on Collaborative Research: Foundations of Data Science Institute (FODSI), the NSF and the Simons Foundation for the Collaboration on the Theoretical Foundations of Deep Learning through awards DMS-2031883 and 814639, NSF grant MC2378 to the Institute for Artificial CyberThreat Intelligence and OperatioN (ACTION), NSF grant 1910100, NSF grant 2046726, NSF AI Institute ACTION No. IIS-2229876, NIH (DMS/NIGMS) grant R01GM152718, AI Safety Fund, and a Berkeley Deep Drive (BDD) Grant from BAIR and a Dean’s fund from CoE, at UC Berkeley. AZ additionally acknowledges support from NSF RTG Grant 1745640. Dr. Aaron Kornblith is a consultant and co-founder for Capture Dx and supported by the Eunice Kennedy Shriver National Institute of Child Health and Human Development of the National Institutes of Health under award number K23HD110716 (AK). This information or content and conclusions are those of the author and should not be construed as the official position or policy of, nor should any endorsements be inferred by HRSA, HHS or the U.S. Government.
Footnotes
The “CDR selection accuracy” here refers to the Exact Match (EM) accuracy defined in the “Experimental Setup” section. It quantifies the proportion of instances where the selected CDR(s) exactly match the gold answer annotated by clinical experts. Additional evaluation metrics, including those used for assessing CDR execution outputs, are also detailed in the same section.
Notes may not be generated until relatively late in an episode of ED care.
The actual computation is to map each token in the text to an embedding vector and then take the average embedding vector over the entire text.
Here, we treat CDR selection for each clinical note as a binary classification problem over the set of all candidate CDRs, including “no applicable CDR”. The actual positives are the applicable CDRs (or “no applicable CDR”), and a true positive is counted when a selected CDR matches the ground truth.
Figures & Tables
References
- 1.Pines JM, Carpenter CR. Clinical Decision Rules. In: Pines JM, Bellolio F, Carpenter CR, Raja AS, editors. Evidence-Based Emergency Care: Diagnostic Testing and Clinical Decision Rules. 3rd ed. John Wiley & Sons Ltd.; 2023. pp. p. 43–52. [Google Scholar]
- 2.Obra JK, Singh C, Watkins K, Feng J, Obermeyer Z, Kornblith A. Systematic Bias in Clinical Decision Instrument Development: A Quantitative Meta-Analysis. medRxiv. 2025. Available from: https://www.medrxiv.org/content/early/2025/02/16/2025.02.12.25320965. [DOI] [PMC free article] [PubMed]
- 3.Heerink ORHRKHKR J S. Clinical decision rules in primary care: necessary investments for sustainable healthcare. Primary health care research development. 2023 May:24. doi: 10.1017/S146342362300021X. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 4.Holmes JF, Kuppermann N. When Do Clinical Decision Rules Improve Patient Care? Annals of Emergency Medicine. 2014 March;63(3):372–3. doi: 10.1016/j.annemergmed.2013.10.023. [DOI] [PubMed] [Google Scholar]
- 5.Chan TM, Mercuri M, Turcotte M, Gardiner E, Sherbino J, de Wit K. Making Decisions in the Era of the Clinical Decision Rule: How Emergency Physicians Use Clinical Decision Rules. Academic Medicine. 2020;95(8):1230–7. doi: 10.1097/ACM.0000000000003098. [DOI] [PubMed] [Google Scholar]
- 6.Neumann J, Vogel C, Kießling L, Hempel G, Kleber C, Osterhoff G, et al. TraumaFlow—development of a workflow-based clinical decision support system for the management of severe trauma cases. International Journal of Computer Assisted Radiology and Surgery. 2024 May;19:2399–409. doi: 10.1007/s11548-024-03191-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 7.Deshormes T, Freibec G, Yanchar N, Zemek R, Beaudin M, Stang A, et al. Low-Value Clinical Practices in Pediatric Trauma Care. JAMA Network Open. 2024 October;7(10):e2440983. doi: 10.1001/jamanetworkopen.2024.40983. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 8.Perry JJ, Stiell IG. Impact of clinical decision rules on clinical care of traumatic injuries to the foot and ankle, knee, cervical, spine and head Injury. Int J Care Injured. 2006 December;37(12):1157–65. doi: 10.1016/j.injury.2006.07.028. [DOI] [PubMed] [Google Scholar]
- 9.Kharel P, Zadro JR, Chen Z, Himbury MA, Traeger AC, Linklater J, et al. Awareness and use of five imaging decision rules for musculoskeletal injuries: a systematic review. International Journal of Emergency Medicine. 2023 November:16. doi: 10.1186/s12245-023-00555-4. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 10.Wang Y, Yao X, Wu X. Transforming Large Language Models into Superior Clinical Decision Support Tools by Embedding Clinical Practice Guidelines. Mayo Clinic Proceedings: Digital Health. 2024;2(3):491–2. [Google Scholar]
- 11.Ong JCL, Jin L, Elangovan K, Lim GYS, Lim DYZ, Sng GGR, et al. Development and Testing of a Novel Large Language Model-Based Clinical Decision Support Systems for Medication Safety in 12 Clinical Specialties. arXiv preprint arXiv:240201741. 2024.
- 12.Lammert J, Dreyer T, Mathes S, Kuligin L, Borm KJ, Schatz UA, et al. Expert-Guided Large Language Models for Clinical Decision Support in Precision Oncology. JCO Precision Oncology. 2024;8:e2400478. doi: 10.1200/PO-24-00478. [DOI] [PubMed] [Google Scholar]
- 13.Umerenkov D, Zubkova G, Nesterov A. Deciphering Diagnoses: How Large Language Models Explanations Influence Clinical Decision Making. arXiv preprint arXiv:231001708. 2023.
- 14.Panicker RO, George AE. Adoption of Automated Clinical Decision Support System: A Recent Literature Review and a Case Study. Archives of Medicine and Health Sciences. 2023;11(1):86–95. [Google Scholar]
- 15.Zakka C, Shad R, Chaurasia A, Dalal AR, Kim JL, Moor M, et al. Almanac — Retrieval-Augmented Language Models for Clinical Medicine. NEJM AI. 2024;1:25. doi: 10.1056/aioa2300068. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 16.Abbasian M, Azimi I, Rahmani AM, Jain R. Conversational Health Agents: A Personalized LLM-Powered Agent Framework. 2024. [DOI] [PMC free article] [PubMed]
- 17.Shi W, Xu R, Zhuang Y, Yu Y, Zhang J, Wu H, et al. EHRAgent: Code Empowers Large Language Models for Few-shot Complex Tabular Reasoning on Electronic Health Records. In: Al-Onaizan Y, Bansal M, Chen YN, editors. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. Miami, Florida, USA: Association for Computational Linguistics; 2024. pp. p. 22315–39. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 18.McDuff D, Schaekermann M, Tu T, Palepu A, Wang A, Garrison J, et al. Towards Accurate Differential Diagnosis with Large Language Models. 2023. [DOI] [PMC free article] [PubMed]
- 19.Tu T, Palepu A, Schaekermann M, Saab K, Freyberg J, Tanno R, et al. Towards Conversational Diagnostic AI. 2024.
- 20.Levra AG, Gatti M, Mene R, Shiffer D, Costantino G, Solbiati M, et al. A large language model-based clinical decision support system for syncope recognition in the emergency department: A framework for clinical workflow integration. European Journal of Internal Medicine. 2025;131:113–20. doi: 10.1016/j.ejim.2024.09.017. [DOI] [PubMed] [Google Scholar]
- 21.Singhal K, Azizi S, Tu T, Mahdavi SS, Wei J, Chung HW, et al. Large language models encode clinical knowledge. Nature (London) 2023;620(7972):172–80. doi: 10.1038/s41586-023-06291-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 22.Saab K, Tu T, Weng WH, Tanno R, Stutz D, Wulczyn E, et al. Capabilities of Gemini Models in Medicine. 2024.
- 23.Wang Y, Ma X, Chen W. Augmenting Black-Box LLMs with Medical Textbooks for Clinical Question Answering. arXiv preprint arXiv:230902233. 2023.
- 24.Xiao W, Oneal P, Wang M, Mehta NJ, Liu Q, Zhang R, et al. Machine Learning-Based Identification of Sickle Cell Disease Subphenotypes in Clinical Trial Data. medRxiv. 2025.
- 25.Holmes JF, Lillis K, Monroe D, Borgialli D, Kerrey BT, Mahajan P, et al. Identifying children at very low risk of clinically important blunt abdominal injuries. Annals of Emergency Medicine. 2013;62(2):107–16.e2. doi: 10.1016/j.annemergmed.2012.11.009. [DOI] [PubMed] [Google Scholar]
- 26.Kuppermann N, Holmes JF, Dayan PS, Hoyle JDJ, Atabaki SM, Holubkov R, et al. Identification of children at very low risk of clinically–important brain injuries after head trauma: a prospective cohort study. The Lancet. 2009;374(9696):1160–70. doi: 10.1016/S0140-6736(09)61558-0. [DOI] [PubMed] [Google Scholar]
- 27.Johnson AE, Bulgarelli L, Pollard TJ, Horng S, Celi LA, Mark RG. MIMIC-IV, a freely accessible electronic health record dataset. Scientific Data. 2023;10(1):1–12. doi: 10.1038/s41597-022-01899-x. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 28.Jin D, Pan E, Oufattole N, Weng WH, Fang H, Szolovits P. What Disease does this Patient Have? A Large-scale Open Domain Question Answering Dataset from Medical Exams. arXiv preprint arXiv:200913081. 2020.
- 29.Bonnet A. Augmented Clinical Notes Dataset. 2023. Accessed: 2025-03-25.






