Skip to main content
Cureus logoLink to Cureus
editorial
. 2024 Jan 15;16(1):e52318. doi: 10.7759/cureus.52318

Pioneering the Integration of Artificial Intelligence in Medical Oral Board Examinations

Satoshi Hanada 1,, Yuri Hayashi 1,2, Sudhakar Subramani 1, Kokila Thenuwara 1
Editors: Alexander Muacevic, John R Adler
PMCID: PMC10866608  PMID: 38357084

Abstract

We evaluated the use of ChatGPT-4, an advanced artificial intelligence (AI) language model, in medical oral examinations, specifically in anesthesiology. Initially proven adept in written examinations, ChatGPT-4's performance was tested against oral board sample sessions of the American Board of Anesthesiology. Modifications were made to ensure responses were concise and conversationally natural, simulating real patient consultations or oral examinations. The results demonstrate ChatGPT-4's impressive adaptability and potential in oral board examinations as a training and assessment tool in medical education, indicating new avenues for AI application in this field.

Keywords: anesthesiology, medical education technology, chatgpt-4, oral board examination, artificial intelligence (ai)

Editorial

Artificial intelligence (AI) has recently made remarkable strides in medicine [1-3], leaving indelible footprints, particularly with the emergence of ChatGPT [4,5]. The advanced language model, ChatGPT-4, has demonstrated a striking aptitude for emulating human conversation since its launch in early 2023. This AI's prowess is notable in written examinations [6-8], such as the U.S. Medical Licensing Examination [9] and specialty board certification written examinations, including the Royal College of Anaesthetists written examination [10].

Nevertheless, a hitherto unexplored territory is the potential of ChatGPT in oral examinations. We delved into the performance of ChatGPT-4 in oral board examinations, which evaluate an examinee’s clinical judgment, adaptability to unanticipated clinical changes, and proficiency in organizing and presenting information. Remarkably, when published American Board of Anesthesiology (ABA) sample sessions [11] were presented to ChatGPT-4, it generated answers that met, or even exceeded, the passing criteria for all questions, as judged by two ABA board-certified anesthesiologists (SH and KT). The initial responses were thorough and comprehensive; however, they were lengthy and lacked a natural conversational tone (Figure 1).

Figure 1. Response from ChatGPT-4 to the following selected question from the ABA oral board sample sessions: A 65-year-old man underwent an uncomplicated CABG 16 hours earlier and was extubated four hours ago. In the past hour, his BP fell from 110/70 to 70/50, and the CVP rose from 8 to 22 mmHg. If tamponade is suspected and mediastinal exploration is required, how would you provide anesthesia?

Figure 1

Thus, we implemented constraints, instructing AI to limit responses to a word count, as if a person had paused to think and respond in the allotted time (Figure 2).

Figure 2. The subsequent response from ChatGPT-4 to the following instruction: Answer within 100 words.

Figure 2

We then added instructions to emulate a conversational tone and to simulate an anesthesiologist's consultation with a patient or a board examination scenario. This resulted in concise, information-dense answers that adhere closely to human conversation patterns, such as in a real oral examination (Figures 3A-3C).

Figure 3. The subsequent responses from ChatGPT-4 to the following instruction: (A) Answer within 100 words in a human conversational manner. (B) Answer within 100 words as if you are the anesthesiologist taking care of this patient. (C) Answer within 100 words as if you were a candidate for the anesthesiology oral board examination.

Figure 3

AI has access to a wealth of information; still, it requires clear instructions to provide the expected response. Our expertise lies in anesthesiology; thus, we chose an anesthesia example. However, this model could be applied to diverse medical specialties. The responses provided by AI are truly impressive, complete with clinical decision-making and the reasoning behind it. This technology could serve as a tool to test both the validity and reliability of question design and also to assist candidates in preparing for oral board examinations. Moreover, the adaptability of ChatGPT-4 suggests its potential role in teaching and preparing medical professionals for real-world clinical scenarios, thereby enhancing their communication and decision-making skills. These facets are new frontiers in the application of AI in medical education and examination, paving the way for further advancements in this rapidly evolving field.

Acknowledgments

The authors used ChatGPT for this article.

The authors have declared financial relationships, which are detailed in the next section.

Kokila Thenuwara, MD declare(s) support for travel and an honorarium have been provided for the service as a board examiner from American Board of Anesthesiology (ABA). Kokila Thenuwara, MD, is a written and applied board examiner for the ABA. The contents, opinions, and conclusions in this article are those of the authors and do not represent those of the ABA . Satoshi Hanada, MD, FASE declare(s) support is provided for attending the ABA Winter 2024 Item Writing Workshop from American Board of Anesthesiology (ABA). Satoshi Hanada, MD, FASE, is a member of the Maintenance of Certification in Anesthesiology Minute Adult Cardiac Anesthesiology Committee for the ABA. The contents, opinions, and conclusions in this article are those of the authors and do not represent those of the ABA

Author Contributions

Concept and design:  Satoshi Hanada, Yuri Hayashi, Sudhakar Subramani, Kokila Thenuwara

Acquisition, analysis, or interpretation of data:  Satoshi Hanada, Yuri Hayashi, Sudhakar Subramani, Kokila Thenuwara

Drafting of the manuscript:  Satoshi Hanada, Yuri Hayashi, Sudhakar Subramani, Kokila Thenuwara

Critical review of the manuscript for important intellectual content:  Satoshi Hanada, Yuri Hayashi, Sudhakar Subramani, Kokila Thenuwara

Supervision:  Satoshi Hanada, Kokila Thenuwara

References

  • 1.Reporting guideline for the early stage clinical evaluation of decision support systems driven by artificial intelligence: DECIDE-AI. Vasey B, Nagendran M, Campbell B, et al. BMJ. 2022;377:0. doi: 10.1136/bmj-2022-070904. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 2.Artificial intelligence-guided screening for atrial fibrillation using electrocardiogram during sinus rhythm: a prospective non-randomised interventional trial. Noseworthy PA, Attia ZI, Behnken EM, et al. Lancet. 2022;400:1206–1212. doi: 10.1016/S0140-6736(22)01637-3. [DOI] [PubMed] [Google Scholar]
  • 3.Adding artificial intelligence to gastrointestinal endoscopy. Berzin TM, Topol EJ. Lancet. 2020;395:485. doi: 10.1016/S0140-6736(20)30294-4. [DOI] [PubMed] [Google Scholar]
  • 4.Using ChatGPT to write patient clinic letters. Ali SR, Dobbs TD, Hutchings HA, Whitaker IS. Lancet Digit Health. 2023;5:179–181. doi: 10.1016/S2589-7500(23)00048-1. [DOI] [PubMed] [Google Scholar]
  • 5.Artificial hallucinations in ChatGPT: implications in scientific writing. Alkaissi H, McFarlane SI. Cureus. 2023;15:0. doi: 10.7759/cureus.35179. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6.Large language models answer medical questions accurately, but can't match clinicians' knowledge. Harris E. JAMA. 2023;330:792–794. doi: 10.1001/jama.2023.14311. [DOI] [PubMed] [Google Scholar]
  • 7.Performance of ChatGPT and GPT-4 on neurosurgery written board examinations. Ali R, Tang OY, Connolly ID, et al. Neurosurgery. 2023;93:1353–1365. doi: 10.1227/neu.0000000000002632. [DOI] [PubMed] [Google Scholar]
  • 8.The performance of ChatGPT on orthopaedic in-service training exams: A comparative study of the GPT-3.5 turbo and GPT-4 models in orthopaedic education. Rizzo MG, Cai N, Constantinescu D. J Orthop. 2024;50:70–75. doi: 10.1016/j.jor.2023.11.056. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 9.How does ChatGPT perform on the United States Medical Licensing Examination? The implications of large language models for medical education and knowledge assessment. Gilson A, Safranek CW, Huang T, Socrates V, Chi L, Taylor RA, Chartash D. JMIR Med Educ. 2023;9:0. doi: 10.2196/45312. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 10.Artificial intelligence and anaesthesia examinations: exploring ChatGPT as a prelude to the future. Aldridge MJ, Penders R. Br J Anaesth. 2023;131:0–7. doi: 10.1016/j.bja.2023.04.033. [DOI] [PubMed] [Google Scholar]
  • 11.Sample standardized oral exam questions. The American Board of Anesthesiology. https://www.theaba.org/wp-content/uploads/2022/12/SOE_Questions.pdf. Amr Bor Aesth. 2023 [Google Scholar]

Articles from Cureus are provided here courtesy of Cureus Inc.

RESOURCES