Skip to main content
BMC Medical Education logoLink to BMC Medical Education
. 2026 Feb 24;26:412. doi: 10.1186/s12909-026-08821-8

Comparison of large language models for clinical scenario generation in medical education: a mixed-methods study

S Öncü 1,✉, F Torun 2, H H Ülkü 3
PMCID: PMC12980910  PMID: 41736016

Abstract

Background

In undergraduate medical education, the ability to manage clinical-cases is a core competency expected of future physicians. Traditionally, this skill is developed through repeated exposure to real patient encounters in clinical settings. However, increasing patient safety concerns, limited clinical opportunities, and faculty workload constraints have made it increasingly difficult for students to access sufficient clinical practice. As a result, innovative solutions such as AI-based simulations are being explored to supplement clinical training. Among these, large language models (LLMs) offer promising potential for generating diverse, interactive, and context-specific clinical scenarios that can support competency-based education.

This study aims to evaluate and compare the effectiveness and educational utility of four widely used and accessible LLMs; ChatGPT-4o, Claude 3.7 Sonnet, Gemini 1.5, and DeepSeek (Chat), in generating clinical scenarios for Turkish undergraduate medical education, and to identify the model that produces the most accurate, understandable, and pedagogically appropriate content aligned with national medical education standards.

Methods

A convergent parallel mixed methods design was employed. Using standardized prompts based on Türkiye’s National Core Undergraduate Medical Education Program-2020, scenarios on three common infectious diseases were generated by each LLM. Twenty-five senior medical students and five expert clinicians evaluated the Turkish-language scenarios using structured rating forms and open ended feedback. Quantitative data were analyzed with Friedman and Wilcoxon tests; qualitative data underwent thematic analysis.

Results

Claude received the highest ratings for clarity, realism, and support for clinical reasoning. Statistically significant differences favored Claude over Gemini and DeepSeek (p < 0.05). Qualitative feedback supported these results, highlighting Claude’s educational value and linguistic precision. ChatGPTperformed moderately, while Gemini and DeepSeek exhibited issues with realism and coherence.

Conclusions

In this study, Claude was rated highest for generating Turkish-language scenarios perceived as clinically appropriate and pedagogically useful for undergraduate medical education in Türkiye. Overall, the findings provide preliminary evidence on perceived scenario quality across models and support further multicenter and outcomes-focused studies to evaluate feasibility, implementation, and educational impact in diverse settings. Future research should also examine how LLM-generated scenarios can be used as supplementary materials in simulation-based learning.

Supplementary Information

The online version contains supplementary material available at 10.1186/s12909-026-08821-8.

Keywords: Artificial intelligence, AI-Based scenario, Clinical scenario generation, Large language models, Medical education, Simulation-based learning

Introduction

Medical education requires bridging theory and real-world practice, as a good physician is defined not only by what they know, but by their ability to apply it effectively and with empathy. According to the World Health Organization (WHO), a competent physician must be clinically knowledgeable, ethically responsible and capable of working effectively within teams to ensure safe and quality care [1].

Clinical scenario-based learning is a cornerstone of modern medical education, providing learners with the opportunity to apply theoretical knowledge to realistic situations in a safe and structured environment. It promotes the development of essential clinical competencies, including diagnostic reasoning, decision-making, communication, and teamwork, which are critical for ensuring patient safety and delivering high-quality care. Scenario-based approaches are widely used in simulations, objective structured clinical examinations (OSCEs), and case-based learning. By presenting authentic clinical challenges, these methods help prepare students for the unpredictable and high-stakes nature of real-world medical practice. Scenario-based teaching has been shown to enhance learner motivation and clinical thinking while supporting the development of core clinical competencies, including clinical reasoning, decision-making, teamwork, and communication [2–6].

However, creating high-quality scenarios is time-consuming, requires subject-matter expertise and often places a heavy burden on educators. In this context, artificial intelligence-powered large language models (LLMs) are increasingly being explored as promising tools to support the development of clinical scenarios in medical education. By enhancing the scalability and accessibility of clinical scenario generation, LLMs have the potential to support broader educational goals while reducing workload for faculty. LLMs such as ChatGPT, Claude, Gemini and DeepSeek are used in medical education for tasks including rapid information retrieval, multiple-choice question creation and scenario-based simulation design. These models can generate realistic, diverse and pedagogically aligned clinical scenarios in a fraction of the time required by manual efforts [7–10].

The use of AI-generated clinical scenarios in medical education aligns closely with Kolb’s Experiential Learning Theory and Dewey’s experiential approach to education. Kolb conceptualizes learning as a cyclical process involving four stages: concrete experience, reflective observation, abstract conceptualization, and active experimentation (Fig. 1) [11, 12]. Dewey’s foundational view emphasizes learning through meaningful experiences, active participation, and reflection rather than passive learning [13]. In this context, AI-generated scenarios function as pedagogical tools that provide structured, simulated clinical encounters. These scenarios support real-world reasoning, encourage reflection, and help bridge the gap between theoretical instruction and practical application, fostering deeper understanding and applied competence in a safe, controlled environment.

Fig. 1.

Fig. 1

Kolb’s experiential learning cycle in AI-assisted scenario-based learning

This study aims to evaluate and compare the effectiveness and educational utility of four large language models; ChatGPT-4o/GPT-4o, Claude 3.7 Sonnet, Google Gemini 1.5, DeepSeek (Chat) in generating clinically valuable scenarios for medical education. It investigates which model produces the most accurate, understandable, and pedagogically appropriate content aligned with national medical education standards, based on feedback from expert clinicians and senior medical students. These models were selected due to their public accessibility, frequent mention in recent literature, and demonstrated potential for clinical content generation [8, 14–16]. Additionally, the study aimed to examine the language adaptability of LLMs by evaluating their performance in generating Turkish-language clinical scenarios, an area that remains underexplored in current medical education research.

Methods

Study design and participants

This study employed a convergent parallel mixed methods design, where qualitative and quantitative data were collected simultaneously and merged during interpretation to provide a comprehensive understanding of the research problem [17]. Clinical scenarios were generated using the same prompt by four LLMs. Model access and versioning: All LLMs were accessed via their respective web interfaces with default settings in early April 2025. Model names/labels were recorded as displayed in the interfaces (ChatGPT-4o, Claude 3.7 Sonnet, Gemini 1.5, and DeepSeek (Chat)). Hereafter, the models are referred to as ChatGPT, Claude, Gemini, and DeepSeek. Scenarios were evaluated by expert clinicians and senior medical students using structured forms. Quantitative and qualitative methods were used to assess scenario quality, clinical realism, and educational relevance.

The study was conducted at Aydın Adnan Menderes University Faculty of Medicine. Expert clinicians were five academic physicians specializing in Infectious Diseases or Pediatric Infectious Diseases Department, each with at least 10 years of clinical and teaching experience. Twenty-five senior medical students (5th-year medical students, who had completed clinical training) were included as participants. The sample size of 25 students was determined based on feasibility, diversity of feedback, and the exploratory nature of the study. Small-to-moderate sample sizes are commonly used and accepted in mixed-methods medical education research, especially when both quantitative and qualitative data are collected from the same group [18]. All participants were voluntary and they were informed about the study.

Scenario selection and prompt development

To represent a range of clinical reasoning challenges and ensure content diversity, three distinct clinical scenarios were used, allowing assessment of the consistency of participants’ responses across different case types. The clinical scenarios were selected based on three common infectious diseases that physicians in Türkiye are expected to manage competently, as outlined in the National Core Undergraduate Medical Education Program (UÇEP 2020) [19]. These included: Crimean-Congo Hemorrhagic Fever (CCHF) (diagnosis, treatment, prevention); Pneumonia (diagnosis, treatment, prevention); Gastroenteritis (diagnosis, treatment, emergency management, and prevention competencies) [19]. CCHF was included due to its endemic nature in the city of the medical faculty, while the other two represent frequently encountered infectious diseases in clinical practice in our country.

A single standardized prompt was prepared for each of the three clinical conditions (CCHF, pneumonia, gastroenteritis), resulting in three separate prompts in total. Each prompt was entered identically into all four LLMs, producing four scenarios per clinical case (12 scenarios in total). Details of the prompts are shown in Fig. 2. All prompts were written in Turkish, and all outputs were used exactly as generated, without any post-editing, correction, or restructuring. This ensured full standardization and prevented evaluator bias.

Fig. 2.

Fig. 2

AI prompt framework for clinical scenario generation

To improve methodological transparency, the complete original prompts and the full AI-generated scenarios (Turkish and English versions) are provided in Supplementary File 1. Each prompt included the target learner level, clinical context, expected competencies, and word limit, as illustrated in Fig. 2. All LLMs received identical prompts for each disease, and each scenario was generated independently. Outputs were checked only for completeness (i.e., presence of required sections) and were otherwise used as generated, without any content, language, or structural edits.

Scenario coding and blinding

All scenarios were anonymized and randomly coded before distribution. Participants, both students and expert clinicians, were blinded to the identity of the language model that generated each scenario, ensuring objective assessment. To ensure blinding and facilitate organized evaluation, scenarios were coded using a structured format:

S.X.Y (S: Scenario, X: Scenario number (1 = CCHF, 2 = Pneumonia, 3 = Gastroenteritis) and Y refers to the LLM used (AI-1 = ChatGPT, AI-2 = Claude, AI-3 = Gemini, AI-4 = DeepSeek). For example, S.1.2 refers to; CCHF case generated by AI-2; Claude.

Data collection tools

Data were collected using two structured evaluation forms, designed by the researchers for this study; one for medical students and the other for expert clinicians. The forms were administered in Turkish, the participants’ native language, because medical students have varying levels of English proficiency and conduct patient encounters, clinical reasoning, and case discussions primarily in Turkish. Using the native language ensured accurate comprehension, reduced potential bias, and reflected the authentic clinical learning environment. All responses were collected via Google Forms. Both instruments were pilot-tested prior to distribution for clarity and usability.

Expert evaluation form

Clinicians completed a 14-item evaluation form assessing the clinical accuracy, realism, diagnostic coherence, safety considerations, and instructional validity of the scenarios. These items provided an expert-level appraisal of whether each scenario was medically sound and suitable for educational use. The full expert questionnaire (Turkish + English) is available in Supplementary File 2.

Student evaluation form

Students completed a 10-item scenario evaluation form designed to assess the linguistic clarity and educational value of each scenario. Items measured clarity of language, readability, logical flow, level appropriateness, cognitive engagement, willingness to discuss, and perceived educational usefulness. The full student questionnaire (Turkish + English) is available in Supplementary File 3. Student forms were completed in a supervised classroom setting to ensure full participation and consistency.

Common structure

Each participant evaluated 12 coded clinical scenarios, presented in randomized order. For every scenario, participants rated items using a 5-point Likert scale (1 = strongly disagree, 5 = strongly agree). Student data were collected in a supervised classroom setting, whereas expert clinicians completed the forms individually and independently.

Open-ended questions

Expert clinicians also responded to two open-ended questions: (1) “Which scenario do you think is most suitable for educational use? Please explain” and (2) “What do you think about using such scenarios in clinical education? Would it contribute to your learning process?”.

Students responded to two open-ended questions: (1) “Which scenario did you like most, and why?” and (2) “What are your views on using such AI-generated scenarios in clinical education?”.

Data analysis

Quantitative data were analyzed using descriptive statistics (frequency, mean, and standard deviation) and non-parametric tests. Due to the ordinal nature of the data and the violation of normality assumptions, the Friedman test was chosen as a suitable method for comparing repeated measures. The Friedman test was used to compare group differences across scenarios. When significant differences were detected, the Wilcoxon signed-rank test was applied to determine the direction of these differences. A p-value < 0.05 was considered statistically significant.

Qualitative analyses were conducted separately for students and instructors due to differences in data volume and response characteristics. Student responses were analysed using thematic content analysis, whereas instructor responses, being fewer and highly similar, were summarized descriptively. Two medical education researchers independently coded the student data and resolved discrepancies through discussion and consensus. A third researcher with expertise in educational technologies assisted by checking code consistency but did not generate conceptual codes. Representative quotations were used to illustrate each theme, and student responses were labelled sequentially (e.g., st01, st02) to ensure anonymity.

Results

Quantitative findings

Student evaluations of AI-generated clinical scenarios were analyzed across three different clinical scenarios (Table 1). Claude was consistently rated highest across multiple criteria, including clarity of language (M = 4.17), clinical realism (M = 4.05), and support for clinical reasoning (M = 3.95), where M represents the mean score on the 5-point Likert scale (Add. File 1. Table 1). It also received the lowest score for the reverse-coded sloppiness item (M = 2.05), indicating careful construction. DeepSeek and ChatGPTearned slightly lower but comparable ratings (Table 1). In addition to student evaluations, instructors also rated Claude most favorably in terms of medical accuracy, content relevance, and educational value, further reinforcing its overall effectiveness in clinical scenario design (Add. File 2. Table 2). Claude also achieved the highest overall student evaluation across all ten questionnaire items (Mean ± SD = 3.76 ± 0.99), outperforming other models in clarity, educational effectiveness, and structure (Add. File 1. Table 1). DeepSeek and ChatGPT followed closely (3.51 ± 0.98), while Gemini scored the lowest (3.38 ± 1.04) (Table 1).

Table 1.

Comparison of AI Models Based on Student Evaluations (Mean ± SD)

AI Model
Scenario 1 Scenario 2 Scenario 3 Overall
Claude 3.80 ± 0.94 3.71 ± 1.02 3.76 ± 1.00 3.76 ± 0.99
ChatGPT 3.55 ± 1.00 3.63 ± 0.96 3.36 ± 1.02 3.51 ± 0.99
DeepSeek 3.62 ± 0.94 3.46 ± 1.00 3.46 ± 0.99 3.51 ± 0.98
Gemini 3.24 ± 1.14 3.48 ± 0.96 3.42 ± 1.01 3.38 ± 1.04

The highest-rated item across both students and instructors was Q1 “The language of the scenario was clear and understandable”, underscoring Claude’s superiority in linguistic clarity and accessibility (Add. File 1. Table 1, Add. File 2.Table 2). The results indicate significant differences in scenario quality ratings by both students and instructors, with Claude being the most consistently preferred model across groups (Tables 1, and 2).

Table 2.

Statistical analysis for instructors

Mean ± SD χ2 df p-value
Scenario AI1 AI2 AI3 AI4
S1 4.44 ± 0.43 4.37 ± 0.59 3.76 ± 0.81 4.09 ± 0.41 2.200 3 0.532
S2 4.01 ± 0.25 4.40 ± 0.47 3.81 ± 0.29 4.09 ± 0.66 4.538 3 0.209
S3 3.56 ± 0.65 4.19 ± 0.28 3.51 ± 0.61 2.90 ± 1.37 8.200 3 0.042 *
Total 4.00 ± 0.13 4.32 ± 0.41 3.70 ± 0.45 3.69 ± 0.58 4.837 3 0.184

*P < 0.05 is statistically significant

Statistical analysis using Friedman and Wilcoxon tests revealed significant differences in scenario-specific ratings by both students and instructors. Among students, significant differences were detected in Scenarios 1 (χ2 = 11.564, p = 0.009); 3 (χ2 = 17.344, p = 0.001); and total scores (χ2 = 12.590, p = 0.006) (Table 3).

Table 3.

Statistical Analysis for Students

Scenario Mean ± SD χ2 df p-value
AI1 AI2 AI3 AI4
S1 3.55 ± 0.67 3.80 ± 0.73 3.24 ± 0.77 3.62 ± 0.61 11.564 3 0.009 *
S2 3.63 ± 0.70 3.71 ± 0.79 3.48 ± 0.69 3.46 ± 0.66 4.242 3 0.236
S3 3.36 ± 0.72 3.76 ± 0.74 3.42 ± 0.64 3.46 ± 0.69 17.344 3 0.001*
Total 3.51 ± 0.99 3.76 ± 0.99 3.38 ± 1.04 3.51 ± 0.98 12.590 3 0.006*

*P < 0.05 is statistically significant

For instructors, S3 (Scenario 3) (χ2 = 8.200, p = 0.042; Table 2), with AI2 (Claude) rated significantly higher than AI3 (Gemini) (Z = –2.032, p = 0.042; Table 4). Wilcoxon pairwise comparisons (Table 4) showed that Claude was rated significantly higher than Gemini (S3.3, p = 0.006; S3.4, p = 0.002) for Scenario 3 and also higher than other models in Scenario 1 (S1.2 vs. S1.3, p = 0.000; S1.2 vs. S1.4, p = 0.027). At the model level, Claude (AI2.Total) received significantly higher overall ratings than AI1 (p = 0.036), AI3 (p = 0.001), and AI4 (p = 0.003).

Table 4.

Wilcoxon pairwise comparison statistics

Participants Comparison Z p-value Participants Comparison Z p-value
Instructors S.3.3 – S.3.2 −2,032 0.042* Students S.3.3 – S.3.2 −2,767 0.006*
Students S.1.3 – S.1.2 −3,644 0.000* Students S.3.4 – S.3.2 −3,136 0.002*
Students S.1.4 – S.1.2 −2,215 0.027* Students AI2total—AI1total −2,093 0.036*
Students S.1.4 – S.1.3 −2,390 0.017* Students AI3total—AI2total −3,211 0.001*
Students S.3.2 – S.3.1 −3,209 0.001* Students AI4total—AI2total −2,999 0.003*

*P < 0.05 is statistically significant

Qualitative findings

Open-ended responses provided by the expert clinicians were limited in number and therefore analysed descriptively rather than thematically. Overall, experts indicated a clear preference for Scenario 2, noting that it contained more complete clinical information and clearer instructional value compared with the other scenarios. Some experts commented that all scenarios were generally suitable for educational use, but Scenario 2 was perceived as the most detailed and pedagogically useful. Representative responses included: “All scenarios may be appropriate for the intended purpose, but if I must choose one, Scenario 2 provides clearer guidance,” and “Scenario 2 contains more comprehensive information than the others.” No additional qualitative themes emerged due to the small number and uniform nature of the expert responses.”

A total of 25 students’ open-ended responses were analyzed using thematic content analysis to explore perceptions regarding the use of AI-generated clinical scenarios in medical education. Because a single open-ended response could reflect multiple thematic categories, individual responses were allowed to be coded under more than one theme. Thus, the frequency values represent coded meaning units rather than unique participants. Three primary codes emerged based on students' qualitative responses: educational applications, opportunities of AI use, and challenges and suggestions (Fig. 3).

Fig. 3.

Fig. 3

Thematic analysis for students’ feedback

These were organized thematically and supported with direct quotations:

  • Educational Applications (n = 18): Students frequently emphasized that AI-generated clinical scenarios contributed significantly to their understanding of clinical reasoning, diagnostic approaches, and practical application of theoretical knowledge.

“These cases prepare us better for real clinical practice than traditional theory” (st04).

“I think it is easier to understand the subject with such clinical scenarios in the course work. So we can understand practically rather than theoretically how real-life cases come” (st02).

  • Opportunities of AI Use (n = 5): Some responses highlighted that AI-generated scenario, when properly structured, could save time and offer valuable practice opportunities.

“Even if created by AI, they can still be useful if organized properly” (st10).

“They help us think critically while being time-efficient” (st11).

  • Challenges and Suggestions (n = 7): Students provided constructive feedback for enhancing the clarity, realism, and educational value of the scenarios. A few students expressed their doubts about AI-generated scenarios.

“It feels like the scenario moves too fast, missing steps we would cover in real practice” (st06).

“It would be better if it included physical examination and history-taking” (st08).

  • Preferred Scenarios by Participants: Across all three cases, both students and instructors most frequently preferred scenarios generated by Claude. The alignment between qualitative themes and quantitative data was evident. High scores on Q1 (clarity), Q9 (clinical thinking), and Q10 (overall impression) supported the theme of educational contribution, while lower Q6 (“The scenario feels human-written”) ratings for Gemini mirrored students’ concerns about realism, demonstrating internal consistency between students’ quantitative ratings and open-ended feedback.

To strengthen mixed-methods integration, we mapped each quantitative evaluation item (Q1–Q10) to its corresponding qualitative theme. This summary matrix is provided in Supplementary File 4 – Table 1.

Discussion

This study evaluated the perceived educational quality of clinical scenarios generated by four LLMs (ChatGPT, Claude, Gemini, and DeepSeek) based on feedback from medical students and clinical experts. Using three distinct scenarios allowed us to observe whether participants' preferences and criticisms were consistent across different clinical contexts.

The findings indicate that participants consistently rated Claude-generated scenarios highest across key domains such as linguistic clarity, educational value, realism, and overall satisfaction. The most highly rated item for both students and instructors was “The language of the scenario was clear and understandable,” emphasizing Claude’s strength in producing accessible content aligned with the cognitive level of senior medical students. Notably, some students perceived the scenarios as potentially authored by human educators, suggesting that well-designed prompts can yield authentic instructional materials. These findings also suggest that LLMs may support efficient development of draft scenario materials.

In this study, Claude’s high performance suggests that it may be suitable for medical training. These results align with prior findings on Claude’s accuracy and content richness in complex medical areas like neuroscience and uveitis [14, 20]. ChatGPT also demonstrated strengths in clarity and flow but was less effective in promoting reasoning, suggesting it provides accurate, accessible content but may not sufficiently foster critical thinking or clinical reflection [21]. Gemini and DeepSeek performed similarly but at a lower level across most criteria. While both offer readable outputs, they lack the depth and realism necessary for advanced medical education [1].

Similar patterns have been observed in the literature regarding the performance and reliability of LLMs in medical education. In the study of Mavrych et al., Claude achieved the highest accuracy (Claude: 83%, ChatGPT: 81.7%, Gemini: 53.6%) on USMLE-style neuroscience questions [20]. This aligns well with Claude's superior ratings in 'Clinical realism' and 'Helps clinical skills' in our study. There are some studies highlighting Claude’s accuracy [15, 16] Zhao et al. evaluated LLMs on uveitis-related clinical questions and found that Claude (96.3% 'Excellent' accuracy) and ChatGPT-4 (88.9%) outperformed Gemini significantly. This reinforces Claude and ChatGPT’s strength in 'Logical info flow' and 'Level appropriateness’ [14]. Ratnagandhi et al. highlighted that ChatGPT had the best readability scores among LLMs, although it lacked clinical depth. This mirrors its high 'clarity of language' rating in our analysis, but low 'encourages reasoning’ [21]. Temsah et al. described DeepSeek's advantage in open-source flexibility and customizability for medical contexts, although they noted the need for better clinical validation, consistent with DeepSeek's lower realism and reasoning scores here [10].

The highest-rated item in the study concerned the clarity of language. One important factor that may have influenced the results is the language of interaction, Turkish. Claude demonstrated strong adaptability to the Turkish language, effectively managing complex sentence structures and contextual nuances, which are essential in medical education scenarios. In contrast, other LLMs showed varying limitations: ChatGPT exhibited reduced fluency and contextual depth; Gemini was less reliable, particularly in clinical accuracy; and DeepSeek struggled to maintain coherence in medical contexts. Therefore, Claude’s superior performance likely reflects not only its reasoning capacity but also its advanced Turkish language processing, supporting its potential as a valuable tool for Turkish-language medical education. In this study, Claude scenarios were perceived as more structured, understandable, and pedagogically useful, suggesting that such scenarios may be used as supplementary teaching materials to support clinical reasoning and active learning. Although research on Claude’s performance in Turkish is limited, cross-lingual studies of LLMs on morphologically rich languages—including Turkish—show variation in fluency and coherence [22]. This variation may help explain the clearer and more coherent outputs observed for Claude in our study.

Both students and expert evaluators rated the educational value of AI-generated scenarios highly. Specifically, instructors' responses to the item "educational value" indicated a strong perception of these scenarios as valuable learning tools. This alignment suggests that the AI-generated scenarios were perceived as clinically appropriate, understandable and pedagogically effective in supporting medical education objectives.

The results of the study align with established educational theories. AI-generated clinical scenarios, particularly those created by Claude, support experiential learning by providing realistic, context-rich cases that promote reflection and active engagement, core components of Kolb’s learning cycle. Similarly, Dewey’s emphasis on problem-based, real-world learning is echoed in participants’ appreciation for the scenarios’ authenticity and practical relevance. By bridging theoretical knowledge with practice, these scenarios offer learners opportunities to build clinical reasoning skills in a structured yet learner-driven manner. The consistency in perceptions between students and instructors further reinforces the relevance of these AI-driven materials within a constructivist educational framework.

The strengths of this study include its mixed-methods design, the inclusion of both student and instructor perspectives, and a comparative evaluation of four LLMs across three UCEP-aligned clinical scenarios. Moreover, its focus on Turkish-language outputs contributes valuable insights into the language adaptability of LLMs, a relatively underexplored area in medical education. While LLMs are often optimized for English, their performance in structurally different languages—such as Turkish, which is agglutinative and morphologically rich—may vary substantially. Therefore, conducting similar studies across typologically diverse languages would help determine whether the patterns observed here are language-specific or generalizable.

The study explores the potential use of LLM-generated scenarios as supplementary teaching materials and compares their perceived usability and educational relevance. Overall, the ratings suggest that some LLMs can produce realistic and well-structured scenarios that may support clinical reasoning and active learning. The study is limited by its single-center context and relatively small sample size. Additionally, the exclusive use of Turkish may have influenced performance differentially across models. Another potential source of bias is participants’ prior familiarity or attitudes toward specific AI systems, which may have influenced their evaluations. Despite these limitations, the findings highlight the value of expert review and careful model selection when considering LLM-generated scenarios as supplementary educational materials. Because we used three standardized disease-specific prompts, each applied to all four LLMs, and did not systematically vary prompt wording, structure, or quality, we could not evaluate the effects of prompt variability or prompt quality-control strategies on scenario quality. Future multicenter studies across different learner levels and clinical disciplines are recommended to validate and extend these findings.

Conclusions

This study is among the first structured investigations to evaluate the perceived educational applicability utility of LLMs in medical education in Türkiye. Through scenario-based comparisons of four advanced LLMs, participants consistently rated Claude as the most effective model in terms of clarity, clinical realism, logical flow, and overall educational value. Its ability to generate Turkish-language scenarios with high linguistic and contextual accuracy was a critical factor in its performance, highlighting the importance of language compatibility in AI-based educational tools. Our findings suggest that Claude-generated scenarios may be used as pedagogically valuable materials for promoting clinical reasoning, reflection, and active learning. The perceived authenticity of these AI-generated cases aligns well with experiential learning theory and constructivist educational principles.

Positive reception from both students and expert educators suggests that LLMs, when guided by thoughtful prompt design and expert oversight, may be useful as supplementary educational materials; however, feasibility and learning outcomes were not evaluated in this study. These findings provide preliminary evidence to inform future work on developing language-sensitive, AI-supported strategies for scenario-based teaching and assessment. Because the scenarios in this study were generated and evaluated in Turkish, further research comparing LLM-generated clinical scenarios across different languages would help clarify whether model performance patterns are language-dependent or broadly generalizable. Future studies across different clinical domains, learner levels, and institutional settings will also be valuable for better understanding the opportunities and limitations of LLMs in diverse educational environments.

Supplementary Information

Supplementary Material 1. (17.7KB, docx)

Acknowledgements

The authors would like to thank the participants for their participation in this study.

Abbreviations

AI

Artificial Intelligence

CCHF

Crimean-Congo Hemorrhagic Fever

WHO

World Health Organization

LLM

Large Language Models

OSCE

Objective Structured Clinical Examinations

Q

Item-Question

S

Scenario

st

Student

UÇEP

National Core Undergraduate Medical Education Program

Authors’ contributions

S.Ö. designed and managed the project, analyzed data, prepared tables and drafted the main manuscript. F.T. supported technical process, conducted data analysis and contributed tables and figures and contributed to manuscript writing. H.H.Ü. conducted the qualitative data and contributed to manuscript writing. All authors reviewed and approved the final version of the manuscript.

Funding

Not applicable.

Data availability

All data generated or analyzed during this study are included in this published article and its supplementary information files.

Declarations

Ethics approval and consent to participate

This study was performed in line with the principles of the Declaration of Helsinki. The approval was granted by the Non-Interventional Ethics Committee of the Faculty of Medicine, Aydın Adnan Menderes University (Prot.No:2025/154). The participants were informed about the survey's voluntary nature and use for research purposes before their participation and informed consent was obtained. The participants were assured that their findings would remain confidential. We confirm that all methods were carried out in accordance with relevant guidelines and regulations.

Consent for publication

Not applicable.

Competing interests

The authors declare no competing interests.

Footnotes

Publisher’s Note

Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.

References

  • 1.World Health Organization, Transforming and scaling up health professionals’ education and training: World Health Organization guidelines. 2013, Geneva: World Health Organization.URL: https://iris.who.int/handle/10665/93635 [PubMed]
  • 2.Shrivastava SRB, Shrivastava P. Inclusion of scenario-based teaching in undergraduate medical education. MRIMS J Health Sci. 2021;9:147–50. 10.4103/mjhs.mjhs_26_21. [Google Scholar]
  • 3.Tian Y. Establishment and actualization of clinical scenario-based learning for the improvement of professionalism of medical students. Chin J Med Educ. 2013;33:708–11. 10.3760/CMA.J.ISSN.1673-677X.2013.05.022. [Google Scholar]
  • 4.Alirezaei S, Sadeghnezhad M, Ramezani M. Evaluating the effect of scenario-based learning on the knowledge, attitude, and perception of nursing and midwifery students about patient safety. J Adv Med Educ Prof. 2024;12:243–50. 10.30476/jamp.2024.101869.1947. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 5.Barry Issenberg S, Mcgaghie WC, Petrusa ER, Lee Gordon D, Scalese RJ. Features and uses of high-fidelity medical simulations that lead to effective learning: a BEME systematic review. Med Teach. 2005;27(1):10–28. 10.1080/01421590500046924. [DOI] [PubMed] [Google Scholar]
  • 6.Cook DA, Hamstra SJ, Brydges R, Zendejas B, Szostek JH, Wang AT, et al. Comparative effectiveness of instructional design features in simulation-based education: systematic review and meta-analysis. Med Teach. 2013;35(1):e867–98. 10.3109/0142159X.2012.714886. [DOI] [PubMed] [Google Scholar]
  • 7.Maaz S, et al. A guide to prompt design: foundations and applications for healthcare simulationists. Front Med. 2025. 10.3389/fmed.2024.1504532. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 8.Öncü S, Torun F, Ülkü HH. AI-powered standardised patients: evaluating ChatGPT-4o’s impact on clinical case management in intern physicians. BMC Med Educ. 2025;25(1):278. 10.1186/s12909-025-06877-6. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 9.Saowaprut P., et al., Evaluation of Large Language Models in Thailands National Medical Licensing Examination. 2024. 10.1101/2024.12.20.24319441
  • 10.Temsah A, et al., DeepSeek in healthcare: revealing opportunities and steering challenges of a new open-source artificial intelligence frontier. Cureus, 2025. 17. 10.7759/cureus.79221 [DOI] [PMC free article] [PubMed]
  • 11.Kolb DA. Experiential learning: experience as the source of learning and development. 2014: FT press. ISBN: 0133892506.
  • 12.Kolb AY, Kolb DA. Experiential learning theory as a guide for experiential educators in higher education. Exp Learn Teach High Educ. 2017;1(1):7–44 (ISSN: 2474-3410). [Google Scholar]
  • 13.Linh D. Applying John Dewey’s experiential learning model to organize life skills education activities for elementary school students. Eur J Theor Appl Sci. 2024;2:760–9. 10.59324/ejtas.2024.2(4).65. [Google Scholar]
  • 14.Zhao F-F, et al. Benchmarking the performance of large language models in uveitis: a comparative analysis of ChatGPT-3.5, ChatGPT-4.0, Google Gemini, and Anthropic Claude3. Eye. 2024. 10.1038/s41433-024-03545-9. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15.Gao T, et al. A Comparison of DeepSeek and Other LLMs. 2025.URL: https://consensus.app/papers/a-comparison-of-deepseek-and-other-llms-gao-jin/f33d8637e05e58ca926ffe62524dc185/.
  • 16.Fujimoto M, et al. Evaluating large language models in dental anesthesiology: a comparative analysis of ChatGPT-4, Claude 3 Opus, and Gemini 1.0 on the Japanese Dental Society of Anesthesiology Board Certification Exam. Cureus. 2024. 10.7759/cureus.70302. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 17.Creswell JW, Tashakkori A. Differing perspectives on mixed methods research. J Mixed Methods Res. 2007;1(4):303–8. 10.1177/1558689807306132. [Google Scholar]
  • 18.A Typology of Mixed Methods Sampling Designs in Social Science Research. The Qualitative Report, 2007;12: p. 281-316. URL: https://consensus.app/papers/a-typology-of-mixed-methods-sampling-designs-in-social/9702f262b6a75d83bbb3da712d52d914/.
  • 19.Ulusal Cep-2020 UCG, Ulusal Cep-2020 UYVYCG, Ulusal Cep-2020 DSBBCG. Medical Faculty- National Core Curriculum 2020. Tıp Eğitimi Dünyası, 2020;19(57-1): 1-146. 10.25282/ted.716873.
  • 20.Mavrych V, Yaqinuddin A, Bolgova O. Claude, ChatGPT, Copilot, and Gemini performance versus students in different topics of neuroscience. Adv Physiol Educ. 2025. 10.1152/advan.00093.2024. [DOI] [PubMed] [Google Scholar]
  • 21.Ratnagandhi JA, et al. Enhancing anesthetic patient education through the utilization of large language models for improved communication and understanding. Anesth Res. 2025. 10.3390/anesthres2010004. [Google Scholar]
  • 22.Xia C, Wu Q, Guan H, Tian S, Hao Y, Wu X. Evaluating modern large language models on low-resource and morphologically rich languages: a cross-lingual benchmark across Cantonese, Japanese, and Turkish. 2025. 10.48550/arXiv.2310.06347.

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Supplementary Material 1. (17.7KB, docx)

Data Availability Statement

All data generated or analyzed during this study are included in this published article and its supplementary information files.


Articles from BMC Medical Education are provided here courtesy of BMC

RESOURCES