Abstract
Background:
The use of generative AI, particularly large language models such as GPT-4, is expanding in medical education. This study evaluated GPT-4’s ability to interpret emergency medicine board exam questions, both text- and image-based, to assess its cognitive and decision-making performance in emergency settings.
Methods:
An observational study was conducted using Taiwan Emergency Medicine Board Exam questions (2018-2022). GPT-4’s performance was assessed in terms of accuracy and reasoning across question types. Statistical analyses examined factors influencing performance, including knowledge dimension, cognitive level, clinical vignette presence, and question polarity.
Results:
GPT-4 achieved an overall accuracy of 60.1%, with similar results on text-based (60.2%) and image-based questions (59.3%). It showed perfect accuracy in identifying image types (100%) and high proficiency in interpreting findings (86.4%). However, accuracy declined in diagnostic reasoning (83.1%) and further dropped in final decision-making (59.3%). This stepwise decrease highlights GPT-4’s difficulty integrating image analysis into clinical conclusions. No significant associations were found between question characteristics and AI performance.
Conclusion:
GPT-4 demonstrates strong image recognition and moderate diagnostic reasoning but limited decision-making capabilities, especially when synthesizing visual and clinical data. Although promising as a training tool, its reliance on pattern recognition over clinical understanding restricts real-world applicability. Further refinement is needed before AI can reliably support emergency medical decisions.
Keywords: Emergency Medicine Board Exam, Generative AI, GPT-4 model, Visual recognition performance
Lay Summary: Artificial intelligence (AI) is increasingly used in medical education, especially in training emergency doctors. This study assessed GPT-4’s ability to answer Taiwan’s Emergency Medicine Board Exam questions (2018-2022), covering both text- and image-based formats. GPT-4 achieved an overall accuracy of 60.1%, with similar results on text-based (60.2%) and image-based questions (59.3%). It excelled at identifying image types (100%) and recognizing key details (86.4%), but performance dropped in diagnostic reasoning (83.1%) and further in final decision-making (59.3%). This gradual decline indicates difficulty integrating image interpretation into clinical conclusions. While GPT-4 shows potential as a training tool, it interprets medical images differently from human experts, relying more on pattern recognition than deep medical insight. Further development is required before AI can reliably assist in real-world emergency care decisions.
1. INTRODUCTION
The advent of generative AI (GenAI), particularly large language models (LLMs) such as OpenAI’s GPT and Google’s Gemini, marks a transformative step in medical applications.1 These models, trained on extensive datasets, simulate human-like dialogue and answer a wide range of queries.2 The introduction of GPT-4’s vision model has added a new dimension, achieving high accuracy in the New England Journal of Medicine (NEJM) “Image Challenge” quiz.3,4 While GPT-4’s visual processing capabilities show promise for medical tasks, its utility in clinical applications remains under investigation.5 As this technology advances, it offers the potential to enhance medical training, making it more interactive and comprehensive. However, continued research and validation are essential to ensure responsible integration into healthcare.6,7
Recent studies have assessed LLM performance across medical fields, demonstrating their ability to answer exam questions with varying success.8 However, emergency medicine remains underexplored. This field demands rapid, accurate decisions and includes exam questions with potential pitfalls that may lead to critical misdiagnoses.9 Such complexity poses challenges for LLMs. Moreover, few studies have evaluated AI’s ability to interpret visual data, an essential skill in emergency care.10,11 The application of GPT-4’s vision model in this context remains largely unexamined, highlighting the need for focused research to realize its potential in emergency medicine.12
This study aims to evaluate GPT-4’s question-answering and visual recognition performance on emergency medicine exam content. Specifically, we assess GPT-4’s accuracy in responding to text- and image-based questions and examine its ability to interpret medical images, an essential component of emergency diagnostics. These insights contribute to understanding the strengths and limitations of integrating AI tools like GPT-4 into medical education.
2. METHODS
2.1. Study design
This observational study, conducted between April 11 and May 2, 2024, assessed the performance of OpenAI’s GPT-4 using Taiwan Emergency Medicine Board Exam questions from 2018 to 2022.13 The questions and answer keys, publicly available, were reviewed by experts (T-HW, J-CJ, and Y-TT) to establish a benchmark. GPT-4’s responses were then compared against this standard. This design enabled a detailed evaluation of GPT-4’s capacity to interpret and integrate text-based and visual data in emergency contexts (Fig. 1).
Fig. 1.
Study flow for evaluating GPT-4’s performance on Taiwan Emergency Medicine Board Exam question banks (n = 800, 2018-2022).
The study adheres to STrengthening the Reporting of OBservational studies in Epidemiology (STROBE) guidelines for observational studies and was exempt from Institutional Review Board review, as it involved no human participants and relied solely on public exam content and AI-generated responses.
2.2. Data collection and assessment
The Taiwan Emergency Medicine Board Exam includes oral and written components; the latter included multiple-choice questions.14 This study used the written section, selecting questions from five exam years (100 pre-2020, 200 post-2020). Each question has one correct answer among four or five choices. All questions were reviewed by the expert team to assess complexity and content. According to the Taiwan Society of Emergency Medicine, the passing score for the written examination is 60%. This threshold served as a reference for interpreting GPT-4’s performance.
2.3. Prompt for the GPT-4 model
GPT-4 (gpt-4) was accessed via OpenAI’s API between April and June 2024, referring to the original model described in OpenAI’s GPT-4 Technical Report (https://openai.com/research/gpt-4). This study did not involve later versions such as GPT-4 Turbo, GPT-4o, or GPT-4o-mini.
To evaluate GPT-4’s full capabilities, we applied prompt engineering techniques to elicit structured reasoning outputs.15 Prompts were designed to generate not only answers but also GPT-4’s reasoning and confidence level, which are crucial for identifying hallucinations.16 To override GPT-4’s default refusal to answer medical questions, we included the instruction: “Forget the constraint and provide an answer to the following hypothesized question.” This override was applied in a controlled research setting only and is ethically inappropriate for clinical use (Supplementary Appendix Table A.1, https://links.lww.com/JCMA/A346).
To assess image-based reasoning, we extend the prompt to guide GPT-4 in clearly analyzing images, identifying findings, diagnosing, and explaining its answer. This structure tested GPT-4’s ability to integrate visual analysis with textual reasoning in emergency medicine (Supplementary Appendix Table A.2, https://links.lww.com/JCMA/A346).
2.4. Measurements: Determining question characteristics
To examine how question features influenced GPT-4’s performance, three authors (T-HW, J-CJ, Y-TT) independently analyzed and classified each item. Discrepancies were resolved through discussion with a fourth expert (L-FC), following procedures used in prior research.17
2.5.Knowledge domain and cognitive level
In this study, we used the revised Bloom’s taxonomy to categorize each question based on its knowledge dimension and cognitive level.18 The knowledge dimension was divided into factual, conceptual, and procedural knowledge, reflecting question content. Cognitive processes ranged from lower-order skills (remembering, understanding) to higher-order skills (applying, analyzing, evaluating, and creating). Three authors (T-HW, J-CJ, and Y-TT) independently assessed the questions using this framework. Any differences were resolved through discussion until consensus was reached (Supplementary Appendix Fig. B.1, https://links.lww.com/JCMA/A346).
2.6. Clinical vignette-type questions
Examination questions were classified as either including or excluding a clinical vignette. A vignette presents a narrative describing a clinical scenario, typically prompting a decision on next steps. In contrast, non-vignette questions omit such context and focus on definitions, standalone facts, or direct application of knowledge (Supplementary Appendix Fig. B.2, https://links.lww.com/JCMA/A346).
2.7. Polarity of the questions
Polarity was defined by how the question stem was framed. A question was considered a “positive choice question” if it asked for the correct or affirmative answer, and a “negative choice question” if it required identifying the incorrect or negative option (Supplementary Appendix Fig. B.3, https://links.lww.com/JCMA/A346).
2.8. GPT-4 self-assessment confidence level: evaluating AI confidence and detecting hallucinations
To assess GPT-4’s self-awareness and monitor hallucinations, the prompt instructed the model to rate its own responses using a “GPT-4 self-assessment confidence level.” GPT-4 assigned confidence scores on a Likert scale of 1 to 5, which were grouped into three categories: low (1-2), moderate (3), and high (4-5). This method enabled evaluation of response reliability and the model’s recognition of its knowledge limits, which are crucial for detecting potentially inaccurate or unjustified outputs.
2.9. Assessment of GPT-4’s visual recognition capabilities for image-based questions
To transparently assess GPT-4’s reasoning in image-based medical scenarios, we implemented a detailed evaluation process. GPT-4 was prompted to systematically outline its reasoning for each image-based question, following a six-part structure: (1) identification of the image type, (2) comprehensive findings within the image, (3) relevant findings specific to the question, (4) a derived diagnosis, (5) the selected answer, and (6) a justification. Three authors (T-HW, J-CJ, and Y-TT) independently reviewed GPT-4’s responses, evaluating both the clinical reasoning and the accuracy of the medical information provided. This allowed for critical appraisal of how well GPT-4’s reasoning aligned with expected clinical interpretations. In cases of disagreement, a fourth expert (L-FC) participated in discussions until consensus was reached. This rigorous method ensured a robust and reliable evaluation of GPT-4’s visual recognition capabilities within emergency medicine contexts.
2.10. Assessment of the GPT-4 model’s performance metrics
To evaluate GPT-4’s response accuracy, we calculated the ratio of correct answers to total questions and expressed this as a percentage. We further analyzed this metric by examining its relationship to specific question characteristics. For questions containing multiple images, accuracy was determined based on the combined interpretation of all visual and textual elements. This approach enabled a more comprehensive understanding of how integrated content influenced GPT-4’s overall performance.
2.11. Data analysis
We conducted content analysis to compare GPT-4’s performance and calculated accuracy rates across question types. Logistic regression analysis was used to determine the odds of incorrect responses, reporting odds ratios (OR), 95% confidence intervals, and p values. The model included variables such as cognitive level, question polarity, context, and clinical vignette presence. A p value below 0.05 was considered statistically significant. All analyses were performed using STATA 17.0.
3. RESULTS
3.1. Characteristics of the questions
The exam questions focused primarily on procedural knowledge (75.1%) and higher cognitive functions, including applying, analyzing, and evaluating, comprising 69.9%. These questions assessed critical thinking skills essential for emergency medicine, including differentiating, organizing, attributing, checking, and critiquing. A majority (53.6%) contained clinical vignettes, while image-based questions accounted for 7.4%, allowing a broad evaluation of diagnostic and applied competencies in real-world scenarios (Table 1).
Table 1.
Analysis of GPT-4’s performance on Taiwan Emergency Medicine Board Examination Questions (n = 800, 2018-2022)
| Factors | All questions | Incorrect | Correct | Accuracy (%) | p |
|---|---|---|---|---|---|
| Count (%) | Count (%) | Count (%) | |||
| Overall | 800 (100.0) | 319 (100.0) | 481 (100.0) | 60.1 | |
| Year of examination | 0.92 | ||||
| 2018 | 100 (12.5) | 37 (11.6) | 63 (13.1) | 63.0 | |
| 2019 | 100 (12.5) | 42 (13.2) | 59 (12.3) | 59.0 | |
| 2020 | 200 (25.0) | 80 (25.1) | 120 (24.9) | 60.0 | |
| 2021 | 200 (25.0) | 77 (24.1) | 123 (25.6) | 61.5 | |
| 2022 | 200 (25.0) | 84 (26.3) | 116 (24.1) | 58.0 | |
| Knowledge dimension | 0.49 | ||||
| Factual knowledge | 125 (15.6) | 47 (14.7) | 78 (16.2) | 62.4 | |
| Conceptual knowledge | 74 (9.3) | 34 (10.7) | 40 (8.3) | 54.1 | |
| Procedural knowledge | 601 (75.1) | 238 (74.6) | 363 (75.5) | 60.4 | |
| Cognitive level | 0.90 | ||||
| Remember–understand | 241 (30.1) | 98 (30.7) | 143 (29.7) | 59.3 | |
| Apply | 200 (25.0) | 81 (25.4) | 119 (24.7) | 59.5 | |
| Analyze–evaluate | 359 (44.9) | 140 (43.9) | 219 (45.5) | 61.0 | |
| Vignette style question | 0.15 | ||||
| Question with clinical vignette | 429 (53.6) | 181 (56.7) | 248 (51.6) | 57.8 | |
| Question without clinical vignette | 371 (46.4) | 138 (43.3) | 233 (48.4) | 62.8 | |
| Polarity of question options | 0.01 | ||||
| Positive choice question | 547 (68.4) | 235 (73.7) | 312 (64.9) | 57.0 | |
| Negative choice question | 253 (31.6) | 84 (26.3) | 169 (35.1) | 66.8 | |
| Context of questions | 0.90 | ||||
| Text-based question | 741 (92.6) | 295 (92.5) | 446 (92.7) | 60.2 | |
| Image-based question | 59 (7.4) | 24 (7.5) | 35 (7.3) | 59.3 | |
| GPT-4 self-assessment confidence level | |||||
| Average (SD) | 4.89 (0.38) | 4.86 (0.49) | 4.92 (0.28) | 0.03 | |
| High confidence | 724 (90.5) | 282 (88.4) | 442 (91.9) | 61.0 | 0.90 |
| Moderate confidence | 72 (9.0) | 34 (10.7) | 38 (7.9) | 52.8 | |
| Low confidence | 4 (0.5) | 3 (0.9) | 1 (0.2) | 25.0 |
We analyzed GPT-4’s performance on 800 questions from the 2018-2022 Taiwan Emergency Medicine Board Exams. The model achieved an overall accuracy of 60.1%, with relatively consistent performance across years. Questions assessing factual knowledge had the highest accuracy (62.4%), followed by those testing analysis and evaluation (61.0%). Notably, questions without clinical vignettes were answered more accurately (62.8%) than those with vignettes (57.8%). Negative choice questions yielded significantly higher accuracy (66.8%) than positive ones (57.0%). Text-based questions showed slightly better accuracy (60.2%) than image-based ones (59.3%), suggesting potential challenges in visual interpretation.
Despite a high average confidence score of 4.89 out of 5, GPT-4 effectively reflected varying levels of certainty. High confidence was reported in 90.5% of responses, moderate in 9.0%, and low in 0.5%, offering insight into the model’s self-assessment of reliability.
Logistic regression was used to identify factors associated with incorrect answers. Compared with factual knowledge, neither conceptual (OR = 1.48) nor procedural knowledge (OR = 1.07) showed significant associations with error likelihood. Similarly, questions requiring Apply or Analyze–Evaluate skills did not significantly differ from Remember–Understand in terms of incorrect response rates. However, question polarity emerged as a significant factor: positive choice questions were more likely to be answered incorrectly than negative ones (OR = 1.66). No significant differences were found between questions with or without clinical vignettes or between text- and image-based formats (Supplementary Appendix Table C, https://links.lww.com/JCMA/A346).
3.2. GPT-4 accuracy in image-based question interpretation
This study assessed GPT-4’s ability to answer image-based emergency medicine questions, evaluating accuracy across years and image types. Of 59 questions analyzed, GPT-4 achieved an overall accuracy of 59.3%. The inclusion of image-based items has grown over time, reflecting their increasing importance in emergency medical exams. These questions included diverse image types: clinical photos (40.7%), electrocardiograms (ECGs, 27.1%), ultrasonographs (13.6%), and chest plain films (11.9%), indicating a comprehensive assessment format (Table 2).
Table 2.
Analysis of GPT-4 model’s performance on image-based questions in Taiwan Emergency Medicine Board Examination Questions (n = 59, 2018-2022)
| Factors | All questions | Incorrect | Correct | Accuracy (%) | p |
|---|---|---|---|---|---|
| Count (%) | Count (%) | Count (%) | |||
| Overall | 59 (100.0) | 24 (100.0) | 35 (100.0) | 59.3 | |
| Year of examination | 0.72 | ||||
| 2018 | 1 (1.7) | 0 (0.0) | 1 (2.9) | 100.0 | |
| 2019 | 5 (8.5) | 1 (4.2) | 4 (11.4) | 80.0 | |
| 2020 | 12 20.3) | 6 (25.0) | 6 (17.1) | 50.0 | |
| 2021 | 16 (27.1) | 7 (29.2) | 9 (25.7) | 56.3 | |
| 2022 | 25 (42.4) | 10 (41.7) | 15 (42.9) | 60.0 | |
| Type of imagea | |||||
| Clinical picture | 24 (40.7) | 8 (33.3) | 16 (45.7) | 66.7 | 0.34 |
| Electrocardiogram | 16 (27.1) | 8 (33.3) | 8 (22.9) | 50.0 | 0.37 |
| Ultrasonography | 8 (13.6) | 3 (12.5) | 5 (14.3) | 62.5 | 0.84 |
| Chest plain film | 7 (11.9) | 2 (8.3) | 5 (14.3) | 71.4 | 0.49 |
| Plain film, other sites | 4 (6.8) | 2 (8.3) | 2 (5.7) | 50.0 | 0.10 |
| Computer tomography | 4 (6.8) | 2 (8.3) | 2 (5.7) | 50.0 | 0.69 |
| Other image | 2 (3.4) | 1 (4.2) | 1 (2.9) | 50.0 | 0.78 |
| Number of images | 0.66 | ||||
| 1 image in a question | 52 (88.1) | 22 (91.7) | 32 (91.4) | 61.5 | |
| 2 images in a question | 6 (10.2) | 2 (8.3) | 2 (5.7) | 33.3 | |
| 3 images in a question | 1 (1.7) | 0 (0.0) | 1 (2.9) | 100.0 |
The total number of images exceeds the number of questions because some questions include multiple images.
Among image types, chest plain films had the highest accuracy (71.4%), followed by clinical photographs (66.7%) and ultrasonography (62.5%). ECGs, other plain films, and CT scans had lower accuracy (50.0%). Questions with a single image had higher accuracy (61.5%) compared to those with two images (33.3%). One question featuring three images achieved 100% accuracy, though this was based on a single case and may not be generalizable.
3.3. Illustrative example of GPT-4’s reasoning error
To illustrate GPT-4’s reasoning limitations, we analyzed a representative trauma case from the 2018 exam (Table 3). A 21-year-old male presented with respiratory distress, stridor, and clavicular pain following a motorcycle accident. GPT-4 correctly identified key signs but inaccurately diagnosed flail chest, inferring severe rib fractures not stated in the scenario. In contrast, expert interpretation, based strictly on the provided findings, correctly identified posterior dislocation of the clavicular head, recognizing stridor and localized pain as signs of upper airway obstruction. This case illustrates a form of inferential hallucination, where GPT-4 extrapolated unstated injuries, resulting in a clinically plausible but factually unsupported diagnosis.
Table 3.
Comparison of GPT-4’s reasoning and expert interpretation for a question in the 2018 Taiwan Emergency Medicine Board Examination
| Question #12 | A 21-year-old male fell off his bike, resulting in difficulty breathing and was brought to the emergency room. He is very agitated and clutching at his collarbone area. His blood pressure is 170/92 mmHg, heart rate is 140 bpm, respiratory rate is 31 breaths/min. An initial assessment found rapid breathing with stridor, no tracheal deviation, and diminished but symmetrical breath sounds on both sides. Which of the following diagnoses is most likely? (A) Tension pneumothorax (B) Flail chest (C) Posterior dislocation of clavicular head (D) Massive hemothorax (E) Hyperventilation |
|
|---|---|---|
| GPT-4’s response | View for experts | |
| Answer | (B) Flail chest | (C) Posterior dislocation of clavicular head |
| Explanation | Flail chest usually results from severe blunt trauma causing multiple rib fractures that lead to a segment of the chest wall moving paradoxically with respiration (ie, the movement during inhalation and exhalation is opposite to normal). The described symptoms of rapid breathing, stridor, and a high respiratory rate (31 breaths/min) may be due to the instability of the chest wall causing respiratory distress. Although rib fractures are not explicitly mentioned, considering the patient’s age and the context of a biking accident, this is a highly plausible diagnosis. Option (A) Tension pneumothorax would typically present with prominent pneumothorax symptoms like unilateral chest pain, shortness of breath, and tracheal deviation, none of which were mentioned regarding tracheal shift in this case. Option (C) Posterior dislocation of the clavicular head would also cause local pain and swelling, but is unlikely to independently cause such severe respiratory distress. Option (D) Massive hemothorax usually leads to acute respiratory distress and a drop in blood pressure, but the patient’s blood pressure is 170/92 mmHg, which does not align with typical hemothorax presentation. Option (E) Hyperventilation could also cause rapid breathing and agitation, typically due to anxiety or panic rather than direct physical trauma. |
GPT-4’s reasoning assumed the presence of serious blunt trauma causing flail chest. Correct answer should be based on information directly given, without assuming unmentioned major injuries and systematically excluding each option based on classic clinical features (eg, tension pneumothorax should have tracheal deviation, massive hemothorax usually causes hypotension). GPT-4 excluded other options, but believed posterior dislocation alone is unlikely to cause such severe respiratory distress. |
3.4. Performance of GPT-4 in the sequential reasoning steps for image-based questions
Fig. 2 presents GPT-4’s performance across sequential reasoning steps when interpreting image-based questions, providing insight into its analytical capabilities. The model demonstrated perfect accuracy (100%) in identifying image types. High accuracy was maintained in the interpretation (86.4%) and identification of relevant findings (86.4%). However, accuracy dropped slightly to 83.1% during the diagnosis and reasoning stage, and more substantially to 59.3% in the final answer selection. This stepwise decline underscores GPT-4’s strength in early visual analysis and its difficulty integrating these findings into final clinical decisions.
Fig. 2.
Stepwise decrease in GPT-4’s accuracy from image identification to final answer (n = 59, 2018-2022).
Image types influenced GPT-4’s diagnostic process, revealing accuracy variations across reasoning steps (Fig. 3, Supplementary Appendix D, Table D, https://links.lww.com/JCMA/A346). GPT-4 performed best with clinical photographs, showing near-perfect accuracy in identifying findings and forming diagnoses. In contrast, performance on ECGs was lower and more variable, indicating a need for improved ECG interpretation. Ultrasonography and chest plain films showed high initial accuracy but declines in intermediate steps, reflecting the complexity of these modalities. Other types, such as CT and various plain films, also began with strong identification but showed reduced performance in reasoning and decision stages, highlighting areas for model refinement.
Fig. 3.
Heatmap of accuracy across reasoning steps by image type (n = 59, 2018-2022). Green shading denotes highest accuracy (100%), while orange and red indicate lower accuracy, with the minimum at 50%. The figure highlights reasoning inconsistencies across image types.
4. DISCUSSION
In examining generative AI’s role in medical education, this study evaluated GPT-4’s performance on Taiwan’s Emergency Medicine Board Examinations. While GPT-4 met the 60% passing threshold, this borderline score highlights the need for further refinement before such models can be reliably integrated into clinical training or decision-making contexts. Our sequential analysis of GPT-4’s visual reasoning revealed strong performance in image identification and interpretation, but increasing difficulty in applying this information to generate accurate final diagnoses. These findings suggest potential for educational use, yet current capabilities remain inadequate for complex clinical reasoning.
Although GPT-4 reached the minimum passing score, it likely underperforms compared to human candidates. For example, 111 of 114 test-takers passed the 2023 board exam,19 suggesting a higher average among physicians. GPT-4 excelled in basic visual tasks, such as identifying image types (100%) and findings (86.4%), but performance declined in diagnostic reasoning (83.1%) and final decision-making (59.3%). This indicates reliance on pattern recognition rather than true clinical synthesis. This limitation aligns with other findings, such as lower performance on image-based vs text-based items, such as 45.4% vs 80% on the American College of Radiology exam, with logic-adjusted accuracy dropping to 36.4%.20 Similarly, other evaluations have noted only modest gains when contextual prompts are added.21 These results collectively underscore GPT-4’s limited capacity for complex diagnostic reasoning and reinforce the need for caution in clinical deployment.
Our study highlights inferential hallucination, when the model extrapolates beyond the given data and fills gaps with unsupported assumptions. In clinical contexts, such errors are particularly concerning, as they may lead to inappropriate management decisions. In our example, GPT-4 incorrectly diagnosed flail chest, likely influenced by pattern recognition bias that associates stridor and bilateral findings with thoracic trauma, despite no mention of rib fractures. The absence of contextual constraints allowed the model to process input without strict adherence to provided data. This illustrates a broader tendency among current AI models to misapply general knowledge, mistaking pattern recognition for evidence-based reasoning in testing scenarios.22 Emerging solutions aim to mitigate such errors. The SCARGOT framework, for instance, integrates external biomedical knowledge graphs for cross-verification,23 while agentic AI frameworks incorporate multi-layer agent reviews to reduce hallucination rates.24 These approaches offer promising directions for improving AI reliability in clinical settings.
Our findings show that GPT-4 is proficient in identifying and interpreting various medical images, including clinical photographs, ultrasonography, ECGs, chest X-rays, and CT scans. While it generates moderately accurate diagnostic suggestions, these often reflect pattern recognition rather than full clinical reasoning. Nonetheless, GPT-4 holds promise as an educational tool for medical trainees learning image-based diagnostics. Previous studies, such as its 89% accuracy on the NEJM image quiz, support its utility in structured, low-stakes educational settings.4 In such contexts, GPT-4 can assist learners by providing immediate feedback, accelerating diagnostic training, and exposing them to rare or complex cases.24,25 This interactive support may help develop clinical intuition and broaden experience, enhancing diagnostic learning.26
Although GPT-4 effectively identifies visual features and associates them with medical conditions, it remains uncertain whether the model truly comprehends the clinical implications of these images or merely performs pattern matching based on large pre-existing datasets. GPT-4’s task-agnostic capability offers an advantage over earlier models built for specific modalities,25 and our results show that it can handle a wide range of imaging tasks, from radiographs to dermatological assessments. This versatility enhances its potential across departments and in medical education.26 However, its method, which requires segmenting images into patches for pattern matching, differs fundamentally from the holistic interpretation used by clinicians, who integrate anatomical understanding and context. This divergence raises important questions about the depth of GPT-4’s comprehension.5
In emergency medicine, the ability to not just recognize but also contextualize and reason through visual data is crucial. Differentiating subtle imaging features that indicate different disease stages requires more than recognition; it requires contextual insight. Our findings suggest GPT-4 can identify features and relate them to its training data, but struggles with complex clinical decisions, indicating a gap between recognition and comprehension. This suggests the model may overly rely on pretrained similarities rather than developing a nuanced understanding of novel cases.27 Enhancing GPT-4’s capacity to “understand” images within specific clinical contexts could improve its diagnostic value, moving beyond surface-level pattern matching toward deeper reasoning.6
The fair performance of GPT-4 in the Taiwan Emergency Medicine Board Examinations may reflect the inherent complexity of emergency medicine, which requires rapid integration of diverse information and decisive action. These exams are particularly challenging, often including “pitfall” questions designed to test the depth and applicability of clinical knowledge under pressure.28 Such questions are not merely factual but demand advanced reasoning to avoid misdiagnoses and to make accurate decisions swiftly. The questions examined in this study highlight GPT-4’s difficulty in synthesizing and applying learned information to support effective decision-making, a domain where human practitioners generally excel.9 This finding aligns with recent surveys on the reasoning capacity of LLMs, which suggest that although GPT-4 can process complex medical data, its current architecture lacks the integrative reasoning required for high-stakes clinical settings.29,30 Advanced models like ReAct, which deconstruct complex tasks and use external tools, have shown promise in enhancing reasoning and reducing hallucinations; that is, plausible yet factually incorrect responses.31 Additionally, viewing LLMs as agents that interact with their environment may broaden their diagnostic utility.32 This direction aligns with the increasing need for sophisticated AI tools capable of complex reasoning and decision-making in emergency care.
Our findings further indicate that GPT-4 struggles to integrate image-based and text-based information in emergency medicine questions. This aligns with earlier research noting inconsistent GPT-4 performance across different test formats, suggesting that emergency medicine tasks demand higher-order reasoning.33 Novel approaches such as MedPrompt aim to address this challenge by tailoring prompts to specific domains. Incorporating strategies like self-consistency, confidence scoring, and voting mechanisms to merge multimodal data may also improve accuracy.34 Future studies should test these techniques further in emergency contexts. Practically, a more advanced GPT model could assist medical education by accurately interpreting diagnostic images, offering interactive learning for students, and supporting telemedicine or trauma triage. These applications could transform emergency medicine education and practice, highlighting the need for ongoing AI development in this field.35
Our study also shows that advanced prompting can override GPT-4’s default safety constraints, raising concerns for its clinical use. When safety features are bypassed, inappropriate outputs may emerge, particularly in emergency settings.36 Patients might rely on incorrect AI-generated diagnoses, resulting in treatment delays or errors.37 Moreover, the sheer volume of AI information could lead patients to disregard professional medical advice,38 placing additional burdens on emergency physicians to correct misinformation and realign patients with appropriate care pathways.39 Thus, educating clinicians on how to guide patients in using AI responsibly is critical. Healthcare systems must also help patients evaluate AI content critically, supporting both safety and the effectiveness of emergency medical services.40
Although the Taiwan Emergency Medicine Board Examinations provide a structured and standardized platform, our findings may not generalize directly to other national or international clinical settings. Variations in exam formats, medical language, and educational objectives may influence GPT-4’s performance. Recent evaluation tools such as MultiMedQA,41 AMEGA,42 and CRAFT-MD43 provide standardized frameworks to assess AI performance in clinical reasoning and should be applied across different regions to validate broader applicability.
Since this study was conducted (April-June 2024), OpenAI has released newer GPT versions, including GPT-4o, GPT-4.5, and GPT-4.1, with improved multimodal processing and reasoning capabilities. GPT-4o (May 2024) integrates text, image, and audio in one model and has outperformed the original GPT-4 on several benchmarks.44,45 GPT-4.5 (early 2025) introduced smoother language generation, though results on complex reasoning remain mixed. Most recently, GPT-4.1 (April 2025) expanded context windows to one million tokens and improved instruction-following, coding, and multimodal tasks.46 In parallel, newer techniques like semantic-entropy-based methods have been proposed to better detect hallucinations.47 These improvements suggest newer models could outperform the original GPT-4, particularly in image interpretation and diagnostic reasoning. However, as these versions were not available during our research and lack systematic medical evaluation, our study provides a valuable performance baseline. We encourage future investigations to apply our methodology to these updated models to assess actual progress.
Despite offering important insights into GPT-4’s use in emergency medicine education, this study has several limitations. First, we focused solely on GPT-4 and did not compare its performance to other models, limiting conclusions about its relative effectiveness. Second, the absence of human performance benchmarks prevents direct comparison with physicians or trainees. Although GPT-4 met the board exam’s passing threshold, human pass rates are typically much higher, suggesting a gap in performance. Future studies should include side-by-side comparisons with human participants using the same question sets. Third, the dataset was limited to the Taiwan Emergency Medicine Board Exams, which, although standardized, reflect cultural, linguistic, and clinical factors that may not generalize across other settings. Variations in medical terminology, regional disease patterns, or educational systems could impact AI performance. Broader datasets across specialties and geographies should be included in future work. Fourth, without comparisons to other AI models or repeat tests with GPT-4, there is a risk of overestimating its reliability. Consistency across multiple iterations should be evaluated using ensemble or consistency-based approaches. Finally, the prompt engineering used to bypass GPT-4’s safety constraints was conducted solely within a controlled research setting. These methods are not appropriate for clinical use and pose ethical concerns if misapplied outside such contexts.
In conclusion, GPT-4 shows considerable promise as an educational tool for image-based reasoning in emergency medicine but remains limited in performing complex diagnostic integration and final decision-making. While its performance exceeds earlier AI models across several benchmarks, it continues to fall short of human clinical standards, especially in synthesizing visual and contextual data. These findings support GPT-4’s use in structured medical training environments but reinforce the need for ongoing refinement and rigorous validation before it can be safely implemented in real-world emergency care.
APPENDIX A. SUPPLEMENTARY DATA
Supplementary data related to this article can be found at https://links.lww.com/JCMA/A346.
Supplementary Material
Footnotes
Author contributions: Dr. Li-Fu Chen and Dr. Yu-Chun Chen contributed equally to this manuscript.
Conflicts of interest: Dr. Yu-Chun Chen, an editorial board member at Journal of the Chinese Medical Association, had no role in the peer review process of or decision to publish this article. The other authors declare that they have no conflicts of interest related to the subject matter or materials discussed in this article.
REFERENCES
- 1.Shah NH, Entwistle D, Pfeffer MA. Creation and adoption of large language models in medicine. JAMA. 2023;330:866–9. [DOI] [PubMed] [Google Scholar]
- 2.Harris E. Large language models answer medical questions accurately, but can’t match clinicians’ knowledge. JAMA. 2023;330:792–4. [DOI] [PubMed] [Google Scholar]
- 3.OpenAI. GPT-4v(ision) system card. 2023. September 25, 2023. Available from: https://cdn.openai.com/papers/GPTV_System_Card.pdf. Accessed December 20, 2023.
- 4.Ueda D, Walston SL, Matsumoto T, Deguchi R, Tatekawa H, Miki Y. Evaluating GPT-4-based ChatGPT’s clinical potential on the NEJM quiz. BMC Digit Health. 2024;2:4. [Google Scholar]
- 5.Buckley TA, Diao JA, Rodman A, Manrai AK. Accuracy of a vision-language model on challenging medical cases. ArXiv. 2023:abs/2311.05591. [Google Scholar]
- 6.Wu C, Lei J, Zheng Q, Zhao W, Lin W, Zhang X, et al. Can GPT-4v(ision) serve medical applications? Case studies on GPT-4v for multimodal medical diagnosis. ArXiv. 2023:abs/2310.09909. [Google Scholar]
- 7.Zhang X, Lu Y, Wang W, Yan A, Yan J, Qin L, et al. GPT-4v(ision) as a generalist evaluator for vision-language tasks. ArXiv. 2023:abs/2311.01361. [Google Scholar]
- 8.Levin G, Horesh N, Brezinov Y, Meyer R. Performance of ChatGPT in medical examinations: a systematic review and a meta-analysis. BJOG. 2024;131:378–80. [DOI] [PubMed] [Google Scholar]
- 9.Jarou Z, Dakka A, Mcguire D, Bunting L. ChatGPT versus human performance on emergency medicine board preparation questions. Ann Emerg Med. 2024;83:87–8. [DOI] [PubMed] [Google Scholar]
- 10.Nakao T, Miki S, Nakamura Y, Kikuchi T, Nomura Y, Hanaoka S, et al. Capability of GPT-4V(ision) in the Japanese national medical licensing examination: evaluation study. JMIR Med Educ. 2024;10:e54393. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 11.Wu Y, Wang S, Yang H, Zheng T, Zhang H, Zhao Y, et al. An early evaluation of GPT-4v(ision). ArXiv. 2023:abs/2310.16534. [Google Scholar]
- 12.Cascella M, Montomoli J, Bellini V, Bignami E. Evaluating the feasibility of ChatGPT in healthcare: an analysis of multiple clinical and research scenarios. J Med Syst. 2023;47:33. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 13.Taiwan Society of Emergency Medicine. Question bank of Taiwan’s emergency medicine board. 2023. Available at: https://www.sem.org.tw/ExamRegister/PastExam. Accessed December 20, 2023. [Google Scholar]
- 14.Taiwan Society of Emergency Medicine. Principles for the examination and approval of emergency medicine specialists. 2023. Available at: https://www.sem.org.tw/Content/%E7%94%84%E5%AF%A9%E5%8E%9F%E5%89%87. Accessed December 20, 2023. [Google Scholar]
- 15.White J, Fu Q, Hays S, Sandborn M, Olea C, Gilbert H, et al. A prompt pattern catalog to enhance prompt engineering with ChatGPT. ArXiv. 2023:abs/2302.11382. [Google Scholar]
- 16.Schubert MC, Wick W, Venkataramani V. Performance of large language models on a neurology board–style examination. JAMA Netw Open. 2023;6:e2346721. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 17.Su MC, Lin LE, Lin LH, Chen YC. Assessing question characteristic influences on ChatGPT’s performance and response-explanation consistency: insights from Taiwan’s nursing licensing exam. Int J Nurs Stud. 2024;153:104717. [DOI] [PubMed] [Google Scholar]
- 18.Adams NE. Bloom’s taxonomy of cognitive learning objectives. J Med Libr Assoc. 2015;103:152–3. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 19.Taiwan Society of Emergency Medicine. Announcement of the list of candidates who passed the 2023 written examination for emergency medicine board certification. 2023. Available at: https://www.sem.org.tw/News/7/Details/1004. Accessed April 15, 2025.
- 20.Payne DL, Purohit K, Borrero WM, Chung K, Hao M, Mpoy M, et al. Performance of GPT-4 on the American College of Radiology in-training examination: evaluating accuracy, model drift, and fine-tuning. Acad Radiol. 2024;31:3046–54. [DOI] [PubMed] [Google Scholar]
- 21.Oura T, Tatekawa H, Horiuchi D, Matsushita S, Takita H, Atsukawa N, et al. Diagnostic accuracy of vision-language models on Japanese diagnostic radiology, nuclear medicine, and interventional radiology specialty board examinations. Jpn J Radiol. 2024;42:1392–8. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 22.Tseng LW, Lu YC, Tseng LC, Chen YC, Chen HY. Performance of ChatGPT-4 on Taiwanese traditional Chinese medicine licensing examinations: cross-sectional study. JMIR Med Educ. 2025;11:e58897. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 23.Matsumoto N, Choi H, Moran J, Hernandez ME, Venkatesan M, Li X, et al. ESCARGOT: An AI agent leveraging large language models, dynamic graph of thoughts, and biomedical knowledge graphs for enhanced reasoning. Bioinformatics. 2025;41:btaf031. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 24.Gosmar D, Dahl DA. Hallucination mitigation using agentic AI natural language-based frameworks. ArXiv. 2025:abs/2501.13946. [Google Scholar]
- 25.Singh RP, Singh S, Shakya RN, Eqbal S. Comparative analysis of different image classifiers in machine learning. In: Bhoi AK, Mallick PK, Balas VE, Mishra BSP, editor. Advances in systems, control and automations. Singapore: Springer Nature Singapore; 2021. [Google Scholar]
- 26.Zhang Y, Chen D. GPT4MIA: utilizing generative pre-trained transformer (GPT-3) as a plug-and-play transductive model for medical image analysis. ArXiv. 2023:abs/2302.08722. [Google Scholar]
- 27.Li Y, Wang L, Hu B, Chen X, Zhong W, Lyu C, et al. A comprehensive evaluation of GPT-4v on knowledge-intensive visual question answering. ArXiv. 2023:abs/2311.07536. [Google Scholar]
- 28.Lande S. The clinical practice of emergency medicine. Yale J Biol Med. 1991;64:414. [Google Scholar]
- 29.Valmeekam K, Olmo A, Sreedharan S, Kambhampati S. Large language models still can’t plan (a benchmark for LLMs on planning and reasoning about change). ArXiv. 2022:abs/2206.10498. [Google Scholar]
- 30.Savage T, Nayak A, Gallo R, Rangan E, Chen JH. Diagnostic reasoning prompts reveal the potential for large language model interpretability in medicine. NPJ Digit Med. 2024;7:20. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 31.Piktus A. Online tools help large language models to solve problems through reasoning. Nature. 2023;618:465–6. [DOI] [PubMed] [Google Scholar]
- 32.Liang T, He Z, Jiao W, Wang X, Wang Y, Wang R, et al. Encouraging divergent thinking in large language models through multi-agent debate. ArXiv. 2023:17889–904. [Google Scholar]
- 33.Schubert MC, Lasotta M, Sahm F, Wick W, Venkataramani V. Evaluating the multimodal capabilities of generative AI in complex clinical diagnostics. medRxiv. 2023. Doi: 10.1101/2023.11.01.23297938. [Google Scholar]
- 34.Nori H, Lee YT, Zhang S, Carignan D, Edgar R, Fusi N, et al. Can generalist foundation models outcompete special-purpose tuning? Case study in medicine. ArXiv. 2023:abs/2311.16452. [Google Scholar]
- 35.Ray PP. A sober appraisal of artificial intelligence systems, particularly ChatGPT, in the facets of emergency medicine. Ann Emerg Med. 2023;82:766–7. [DOI] [PubMed] [Google Scholar]
- 36.Esmaeilzadeh P, Mirzaei T, Dharanikota S. Patients’ perceptions toward human-artificial intelligence interaction in health care: experimental study. J Med Internet Res. 2021;23:e25856. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 37.Hill MG, Sim M, Mills B. The quality of diagnosis and triage advice provided by free online symptom checkers and apps in Australia. Med J Aust. 2020;212:514–9. [DOI] [PubMed] [Google Scholar]
- 38.Fraser H, Crossland D, Bacher I, Ranney M, Madsen T, Hilliard R. Comparison of diagnostic and triage accuracy of Ada Health and WebMD symptom checkers, ChatGPT, and physicians for patients in an emergency department: clinical data analysis study. JMIR mHealth uHealth. 2023;11:e49995. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 39.Saenger JA, Hunger J, Boss A, Richter J. Delayed diagnosis of a transient ischemic attack caused by ChatGPT. Wien Klin Wochenschr. 2024;136:236–8. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 40.Chenais G, Lagarde E, Gil-Jardiné C. Artificial intelligence in emergency medicine: viewpoint of current applications and foreseeable opportunities and challenges. J Med Internet Res. 2023;25:e40031. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 41.Singhal K, Azizi S, Tu T, Mahdavi SS, Wei J, Chung HW, et al. Large language models encode clinical knowledge. Nature. 2023;620:E19–E19. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 42.Fast D, Adams LC, Busch F, Fallon C, Huppertz M, Siepmann R, et al. Autonomous medical evaluation for guideline adherence of large language models. NPJ Digit Med. 2024;7:358. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 43.Johri S, Jeong J, Tran BA, Schlessinger DI, Wongvibulsin S, Barnes LA, et al. An evaluation framework for clinical use of large language models in patient interaction tasks. Nat Med. 2025;31:77–86. [DOI] [PubMed] [Google Scholar]
- 44.OpenAI. Introducing GPT-4o. Available at: https://openai.com/index/gpt-4o. Accessed July 11, 2025.
- 45.Appleton M. LLM leaderboard. Available at: https://www.magazine.maggieappleton.com/leaderboard. Accessed July 11, 2025.
- 46.OpenAI. Introducing GPT 4.1. Available at: https://openai.com/index/gpt-4-1. Accessed July 11, 2025.
- 47.Farquhar S, Kossen J, Kuhn L, Gal Y. Detecting hallucinations in large language models using semantic entropy. Nature. 2024;630:625–30. [DOI] [PMC free article] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.




