Abstract
Background
Artificial intelligence (AI) has demonstrated remarkable capabilities across diverse medical applications, potentially revolutionizing healthcare delivery systems. This systematic review and meta-analysis investigated the comparative effectiveness of generative artificial intelligence (GAI)-based teaching methodologies versus conventional pedagogical approaches on educational outcomes among medical students.
Methods
We conducted a comprehensive literature search across multiple electronic databases, including PubMed, Cochrane Library, EMBASE, and Web of Science, encompassing studies published from January 2014 through January 2025. The review focused on randomized controlled trials (RCTs) that compared GAI-based teaching interventions with traditional instructional teaching methods in medical students.
Results
The meta-analysis incorporated 11 eligible RCTs, comprising 786 medical students. Pooled analysis revealed no statistically significant difference in knowledge acquisition scores between GAI-based and traditional teaching approaches (standardized mean difference [SMD] 0.27, 95% confidence interval [CI] -0.31 to 0.85; p = 0.36). However, subgroup analysis indicated enhanced knowledge performance in the GAI group specifically for extended learning periods (exceeding one week) and practice-oriented courses. GAI-based instruction demonstrated superior outcomes in practical skill development compared to conventional methods (SMD 0.63, 95% CI 0.10–1.16; p = 0.02). Students in the GAI group reported significantly higher satisfaction scores with their learning experience.
Conclusion
While theoretical knowledge acquisition remains comparable between teaching modalities, the distinctive advantages of GAI-based approaches in practical skill development warrant their integration into medical curricula. Future research should focus on optimizing the integration of GAI-based teaching methods, standardizing implementation protocols, and evaluating long-term educational outcomes.
Trial registration
This protocol was registered on the International Platform of Registered Systematic Review and Meta-analysis Protocols (INPLASY) with the registration number INPLASY202510006. Registered on 2 January 2025.
Supplementary Information
The online version contains supplementary material available at 10.1186/s12909-025-07750-2.
Keywords: Artificial intelligence, Generative artificial intelligence, Medical education
Introduction
Medical education has traditionally relied on lecture-based instruction supplemented with multimedia resources and clinical demonstrations [1]. Recent decades have witnessed evidence-based innovations in medical education. Seminar-based teaching has demonstrated considerable success in enhancing students’ active learning capabilities [2], while simulation-based education has shown marked improvements in practical skills and knowledge application [3]. Nevertheless, simulation-based education presents substantial challenges, including significant financial investments in infrastructure and equipment maintenance. Furthermore, these simulations are limited in their ability to comprehensively replicate the full spectrum of clinical scenarios, particularly in emergency and critical care settings [4, 5].
The exponential advancement of artificial intelligence (AI) has transformed various aspects of healthcare delivery, from clinical diagnosis support to patient education [6], necessitating its integration into medical education curricula. A cross-sectional survey spanning 48 countries revealed widespread enthusiasm among medical students regarding AI applications in healthcare, with strong interest in expanding their AI-related educational opportunities [7]. Generative artificial intelligence (GAI), a subset of AI technology based on large language models, creates diverse content through iterative learning from extensive datasets [8, 9]. GAI systems like ChatGPT have demonstrated remarkable capabilities, successfully passing the United States Medical Licensing Examination (USMLE) and achieving performance levels comparable to senior medical students across various licensing exams [10–12]. Research has consistently demonstrated that active learning methodologies substantially enhance educational outcomes [13]. GAI technology offers unique advantages in this context, providing personalized, interactive learning experiences through simulated clinical scenarios and real-time feedback mechanisms. This approach potentially fosters autonomous learning capabilities and sustained educational engagement among students [14–16]. However, the implementation of GAI in medical education has prompted legitimate concerns among educators, particularly regarding potential overreliance on technology and academic integrity issues, including plagiarism and data fabrication [8, 15].
Despite these developments, the empirical evidence supporting the effectiveness of GAI-based teaching methodologies in medical education remains inconclusive. Therefore, this meta-analysis aimed to systematically evaluate the comparative effectiveness of GAI-based teaching approaches versus traditional pedagogical methods in medical education.
Methods
This systematic review and meta-analysis adhered to the PRISMA (Preferred Reporting Items for Systematic Reviews and Meta-Analyses) statement [17]. Ethical approval and patient consent were not required as all analyses were based on previously published studies.
Search strategy and criteria
A comprehensive literature search was conducted by two researchers (L.J. and Y.K.Y.) across multiple electronic databases: PubMed, Cochrane Library, EMBASE, and Web of Science. The search period spanned from January 2014 to January 2025. The analysis included randomized controlled trials (RCTs) that compared GAI-based teaching methodologies with traditional teaching approaches in medical education. Reference lists of eligible studies and relevant review articles were manually screened to identify additional pertinent studies. The detailed search strategy is provided in Appendix 1.
Studies were included if they met all of the following criteria: (1) enrolled medical students as participants; (2) implemented GAI-based teaching methods as the primary intervention; (3) utilized traditional teaching approaches in control groups; (4) reported quantifiable outcomes including knowledge or skill assessment scores; and (5) employed a randomized controlled trial design. Studies were excluded if they were duplicate publications, involved non-medical student populations, lacked genuine GAI tool integration in interventions, did not include a traditional teaching control group, used non-RCT study designs (such as reviews, commentaries, or observational studies), or failed to provide complete and extractable outcome data. Traditional teaching methods are defined as instructor- or expert-led pedagogical approaches where content is delivered primarily through conventional lectures and didactic instruction, without the integration of artificial intelligence technologies in any aspect of the teaching process. GAI-based teaching methods, in contrast, are characterized by the systematic incorporation of large language model-powered tools (e.g., ChatGPT, conversational AI agents) into the learning environment, where these technologies actively facilitate content delivery, provide personalized feedback, and enhance student-instructor interaction.
Two reviewers (L.J. and Y.K.Y.) independently screened titles and abstracts using EndNote X9 software for reference management. Following initial screening, full-text articles were evaluated for final inclusion. Any disagreements were resolved through consensus discussions with the senior research team (Y.D.W, C.D.X. and J.X.Q.).
Data extraction
Data were extracted as follows: study characteristics (authors, publication date, and country); student characteristics (sample size, grade, and region); GAI measures (tools related to GAI); course features (course type, learning duration, and assessment timing); and outcomes. The primary outcomes were knowledge scores and skill scores after exams. The teaching satisfaction of students was considered as the secondary outcome.
Risk of bias assessment
The methodological quality of included RCTs was evaluated using the Cochrane Collaboration Risk of Bias tool [18]. Six domains were assessed: selection, performance, detection, attrition, reporting, and other biases. Risk levels were categorized as high, low, or unclear. Two reviewers (L.J. and Y.K.Y.) conducted independent assessments, with discrepancies resolved through research team consultation (Y.D.W, C.D.X. and J.X.Q.).
The quality of evidence assessment
Evidence quality was evaluated using the GRADE (Grading of Recommendations Assessment, Development, and Evaluation) framework [19]. Outcomes were classified as very low, low, moderate, or high quality. Disagreements were resolved through the previously described consensus process.
Statistical analysis
Statistical analyses were performed using Review Manager (Version 5.4). Continuous variables were expressed as means with standard deviation (SD). Standardized mean difference (SMD) with 95% confidence interval (CI) was calculated for all continuous variables. The I² statistic was employed to assess heterogeneity, with I² > 50% indicating substantial heterogeneity [20]. A random-effects model was employed for outcomes showing substantial heterogeneity (I² > 50%) or when meaningful methodological heterogeneity was present (e.g., variations in skill assessment tools and methods), as determined by investigator assessment; otherwise, a fixed-effects model was used. Predetermined subgroup analyses examined knowledge score variations across four parameters: geographical location (Asia versus Europe), course classification (theory versus practical), learning duration (< 1 week versus ≥ 1 week), and assessment timing (immediate versus delayed). Practice courses were defined as those conducted during clinical rotations or internships, while theoretical courses encompassed all other educational formats. Statistical significance was set at p < 0.05.
Results
Literature search
The initial database search yielded 1200 records. Following the removal of 158 duplicates, 1042 titles and abstracts underwent preliminary screening by two independent reviewers. After excluding 1010 ineligible records, 32 full-text articles were assessed for eligibility. Ultimately, 11 studies met the inclusion criteria and were included in the meta-analysis (Fig. 1). A comprehensive list of exclusion reasons is available in the Appendix 3.
Fig. 1.
Flowchart for the literature search and exclusion criteria
Study and participant characteristics
The meta-analysis encompassed 11 RCTs with a total of 786 medical students, comprising 395 participants in GAI-based intervention groups and 391 in traditional teaching control groups [21–31]. All included studies were published between 2024 and 2025, with a predominant representation from Asian institutions (8 studies, 72.7%). The study population primarily consisted of undergraduate medical students, with only one study including postgraduate participants. ChatGPT was the primary GAI tool, utilized in 9 studies (81.8%). 5 studies (45.5%) implemented interventions exceeding one week. 6 studies (54.5%) focused on theoretical courses, while 5 studies (45.5%) addressed practical courses. Detailed baseline characteristics of all included studies are presented in Table 1.
Table 1.
Baseline characteristics of all included studies
| Study | Country | Tradition (n) | GAI (n) | GAI measures | Students | Course type | Learning duration | Assessment timing | Primary Outcomes |
|---|---|---|---|---|---|---|---|---|---|
| Ba et al. 2024 [21] | China | 38 | 39 | ChatGPT | Undergraduate | Practice course | 2 weeks | Immediate | Knowledge scores and skill assessment |
| Bhatia et al. 2024 [22] | India | 50 | 50 | ChatGPT | Undergraduate | Theory course | 30 min | Immediate | Knowledge scores |
| Çiçek et al. 2024 [23] | Turkiye | 65 | 64 | ChatGPT | Undergraduate | Theory course | 5 days |
Immediate 10 days later |
Knowledge scores |
| Gan et al. 2024 [24] | China | 56 | 54 | ChatGPT | Undergraduate | Theory course | 1 week | Immediate | Knowledge scores |
| Huang et al. 2024 [25] | China | 32 | 32 | ChatGPT | Undergraduate | Practice course | Not mentioned | Immediate | Knowledge scores |
| Kavadella et al. 2024 [28] | Europe | 38 | 39 | ChatGPT | Undergraduate | Theory course | 4 weeks | Immediate | Knowledge scores |
| Jiang et al. 2024 [27] | China | 30 | 31 | The chatbot utilizing standardized patient | Undergraduate | Practice course | 2 h | Immediate | Knowledge scores |
| Svendsen et al. 2024 [29] | Norway | 16 | 15 | ChatGPT | Undergraduate Postgraduate | Theory course | 45 min | Immediate | Knowledge scores |
| Tabuchi et al. 2024 [30] | Japan | 28 | 27 | AI-generated image | Undergraduate | Theory course | 10 min | Immediate | The accuracy rates of test |
| Wu et al. 2024 [31] | China | 30 | 31 | ChatGPT | Undergraduate | Practice course | 4 weeks | 3 days later | Knowledge and skill scores |
| Zeng et al. 2025 [26] | China | 21 | 21 | ChatGPT | Undergraduate | Practice course | 2 weeks | 3 days later | Knowledge and skill scores |
GAI Generative Artificial Intelligence, ChatGPT Chat Generative Pre-trained Transformer
Risk of bias and quality assessment
Regarding random sequence generation, 5 studies were assessed as low risk [21, 22, 24, 26, 29] and one study as high risk [30], with the remaining studies being unclear. For allocation concealment, only one study demonstrated low risk [24], while others were classified as unclear due to insufficient reporting of concealment methods. With respect to blinding, 3 studies reported their blinding methodology [21, 22, 24], though only one study implemented blinding of outcome assessment [23]. Additionally, incomplete outcome data and selective reporting were identified as high-risk domains in 2 studies each [27, 28]. The comprehensive risk of bias assessment for individual studies is visualized in Fig. 2. The overall quality of evidence, as evaluated using the GRADE framework, was categorized as low (Appendix 2). This classification was primarily attributed to substantial heterogeneity across studies, inadequate reporting of randomization procedures, and limited implementation of blinding methods.
Fig. 2.
Risk of bias summary for included studies
Meta-analysis
Data on knowledge scores were extracted from all eleven studies (395 GAI-based students and 391 traditional teaching students). Meta-analysis with a random-effects model showed that there was no significant difference in knowledge scores between the groups (SMD 0.27, 95% CI −0.31 to 0.85; p = 0.36; I² = 93%; Fig. 3). Funnel plots showed that no publication bias was observed (Fig. 4). A low quality of evidence was observed as assessed using the GRADE method (Appendix 2). Subgroup analysis revealed that a learning duration of more than one week with GAI could improve knowledge scores (p = 0.02; Table 2). Moreover, learning with GAI in practice courses resulted in higher knowledge scores (p < 0.01; Table 2), whereas it had little effect on theory courses. Neither the geographical location of instruction implementation nor the timing of student achievement assessment influenced the comparative effectiveness of the two teaching methods on theoretical achievement.
Fig. 3.
Forest plot of knowledge scores, skill scores and teaching satisfaction
Fig. 4.
Funnel plot of knowledge scores
Table 2.
Subgroup analysis for the knowledge scores
| Subgroups | Studies | GAI (n) | Tradition (n) | SMD | 95% CI | Model | I2 | P | Subgroup Differences | ||
|---|---|---|---|---|---|---|---|---|---|---|---|
| Chi2 | I2 | P | |||||||||
| Geographical location | |||||||||||
| Asia | 8 | 285 | 285 | 0.25 | [−0.55, 1.04] | Random | 95% | 0.55 | 0.01 | 0% | 0.90 |
| Europe | 3 | 110 | 106 | 0.31 | [−0.30, 0.92] | Random | 77% | 0.32 | |||
| Learning duration | |||||||||||
| < 1 week | 5 | 179 | 183 | −0.34 | [−1.26, 0.58] | Random | 94% | 0.47 | 3.20 | 68.8% | 0.07 |
| ≥ 1 week | 5 | 184 | 176 | 0.64 | [0.09, 1.19] | Random | 84% | 0.02 | |||
| Course classification | |||||||||||
| Theory course | 6 | 241 | 240 | −0.25 | [−1.00, 0.50] | Random | 93% | 0.51 | 5.19 | 80.7% | 0.02 |
| Practice course | 5 | 154 | 151 | 0.89 | [0.25, 1.54] | Random | 86% | < 0.01 | |||
| Assessment timing | |||||||||||
| Immediate | 9 | 343 | 340 | 0.04 | [−0.55, 0.64] | Random | 93% | 0.88 | 1.39 | 28.0% | 0.24 |
| Delayed | 3 | 108 | 110 | 0.82 | [−0.32, 1.96] | Random | 93% | 0.16 | |||
GAI Generative Artificial Intelligence, SMD Standardized Mean Differences, CI Confidence Interval
Additionally, the meta-analysis, which involved only two studies (52 GAI-based students and 51 traditional teaching students), showed that the GAI-based teaching had significantly higher skill scores compared to the traditional teaching method (SMD 0.63, 95% CI 0.10 to 1.16; p = 0.02; I² = 43%; Fig. 3), under the random-effects model. There were three studies conducted on the survey on teaching satisfaction in both student groups, and the GAI group reported higher satisfaction than the traditional group (SMD 1.18, 95% CI 0.50 to 1.87; p < 0.01; I² = 79%; Fig. 3), with the difference being statistically significant. The GRADE quality of evidence is presented in Appendix 2.
Discussion
Although GAI-based teaching methods have not significantly improved theoretical scores compared with traditional teaching, GAI teaching methods can better improve theoretical scores by extending the teaching time of GAI and adjusting the practical content of GAI courses. As expected, GAI-based teaching methods significantly improve skills and teaching satisfaction.
Although various approaches are employed to evaluate medical students, examination scores remain the most direct indicator of their knowledge acquisition, skills development, and overall learning outcomes [32]. Our analysis revealed no significant difference in knowledge scores between the two teaching methods; however, this finding should not diminish the potential effectiveness of GAI-based instruction. GAI teaching methodologies, being relatively novel, are still evolving in their pedagogical applications. Notably, our subgroup analysis indicated that GAI demonstrated particular efficacy in practical courses. The limited frequency and duration of GAI implementation in the analyzed studies suggest that extended application periods and enhanced integration may be necessary for comprehensive evaluation. Furthermore, the results may be influenced by confounding factors, including students’ established learning patterns [33]. The absence of standardized instructor training protocols and a robust framework for assessing teaching effectiveness presents additional methodological considerations [34]. These identified gaps in current research warrant further investigation and methodological refinement.
In the domain of skill assessment, GAI-based teaching demonstrated superior performance compared to traditional teaching methodologies. Although the analysis of skill scores encompassed only two studies, the significant potential of GAI in medical education warrants attention. Consistent with our subgroup analysis findings, GAI-based methods yielded improved knowledge scores in practical courses relative to theoretical instruction. In contrast to conventional passive knowledge acquisition approaches, GAI facilitates active learning by promoting independent information seeking, inquiry, and problem-solving [35]. This interactive pedagogical approach enhances comprehension while fostering essential cognitive competencies, including critical thinking, problem-solving abilities, and self-directed learning capabilities. GAI addresses traditional constraints of instructor availability and standardized curricula by providing individualized learning experiences and iterative practice opportunities tailored to specific student needs, thereby enhancing learner confidence [36]. These elements are fundamental to the acquisition of complex clinical competencies. Furthermore, GAI effectively simulates diverse clinical scenarios, offering students broader exposure to practical cases and clinical experiences that are typically challenging to replicate through conventional teaching methods [37, 38]. On the other hand, Cook also highlighted that with thoughtfully designed and appropriately implemented AI-related instructional models, immersive learning experience can be transformed into an accessible, cost-effective, and sustainable educational approach [39]. Consequently, the interactive and adaptive capabilities of GAI should be strategically integrated into guided learning curriculum design to optimize comprehension and knowledge retention.
It is crucial to acknowledge that GAI currently functions primarily as a complementary educational tool rather than a replacement for traditional teaching methodologies in medical education [8, 14, 40]. Without robust curricular integration, GAI’s impact on knowledge acquisition and skill development may be constrained, underscoring the necessity for strategic implementation within existing educational frameworks. To address these limitations, researchers have investigated the synergistic potential of AI systems with innovative pedagogical approaches, including flipped classrooms, case-based learning (CBL), and gamified instruction, which may enhance information processing and reinforce foundational knowledge [41–43]. At the same time, the establishment of a practical framework for AI literacy in medical education is imperative for guiding the design, implementation, and evaluation of curriculum content, resource allocation, and intended learning outcomes [44]. Nevertheless, this endeavor places substantial demands on the digital competencies of medical educators. Systematic training for both students and educators is essential to ensure effective GAI utilization and successful adaptation to evolving educational paradigms [45, 46]. Further rigorous research remains necessary to evaluate the longitudinal effects of GAI implementation and establish evidence-based best practices for meaningful integration across diverse curricular contexts.
This meta-analysis exhibits several notable limitations that merit careful consideration. The statistical power of our findings may be constrained by the restricted number of included studies and their relatively modest sample sizes. Despite conducting extensive subgroup analyses that resulted in reduced heterogeneity, the remaining substantial heterogeneity warrants acknowledgment. A significant methodological concern stems from the quality of included studies, many of which lacked comprehensive documentation regarding randomization procedures and blinding protocols, potentially introducing systematic bias. The limited prior exposure of participants to AI-based educational tools represents an additional confounding factor, as unfamiliarity with these technologies may have influenced learning efficiency and intervention outcomes. Future research should address these limitations through well-designed, large-scale studies employing standardized protocols, while simultaneously investigating the longitudinal impact of GAI integration in educational settings.
Conclusions
This meta-analysis provided evidence regarding the role of GAI-based teaching methodologies in medical education. While our findings indicated comparable theoretical knowledge acquisition between GAI and traditional teaching approaches, GAI demonstrated significant advantages in practical skill development and student engagement. The enhanced outcomes in extended learning periods and practice-oriented courses suggest that GAI’s effectiveness is particularly pronounced in specific educational contexts.
However, successful implementation requires careful consideration of several factors: robust curricular integration, adequate technological support, and systematic training for both educators and students. Future research should focus on conducting large-scale, longitudinal studies with standardized methodologies to better understand the long-term impact of GAI integration and establish evidence-based best practices across diverse medical education settings.
Supplementary Information
Acknowledgements
None.
Abbreviations
- AI
Artificial Intelligence
- ChatGPT
Chat Generative Pre-trained Transformer
- CI
Confidence Interval
- GAI
Generative Artificial Intelligence
- GRADE
Grading of Recommendations Assessment, Development, and Evaluation
- PRISMA
Preferred Reporting Items for Systematic Reviews and Meta-Analyses
- RCTs
Randomized Controlled Trials
- SMD
Standardized Mean Difference
Authors’ contributions
L.J. and Y.K.Y. provided study design, literature screening, and data extraction. All authors contributed to manuscript writing, review, and editing and approved the final submitted version.
Funding
This study was supported by the Science and Technology Department of Sichuan Province (No. 2024NSFSC1679 to Dong Xu Chen) and Chengdu Science and Technology Department - Technological Innovation R&D Project (General Project. No. 2024-YF05-00339-SN to Dong Xu Chen).
Data availability
The datasets and analyzed during the current study are available from the corresponding author on reasonable request.
Declarations
Ethics approval and consent to participate
Ethical approval and patient consent were not required as all analyses were based on previously published studies.
Consent for publication
Not applicable.
Competing interests
The authors declare no competing interests.
Footnotes
Publisher’s Note
Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
Juan Li and Kaiyu Yin contributed equally to this work.
Contributor Information
Xiaoqin Jiang, Email: 1598862657jxq@scu.edu.cn.
Dongxu Chen, Email: scucdx@foxmail.com.
References
- 1.Wang B, Jin S, Huang M, Zhang K, Zhou Q, Zhang X, et al. Application of lecture-and-team-based learning in stomatology: in-class and online. BMC Med Educ. 2024;24(1):264. 10.1186/s12909-024-05235-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 2.Zeng HL, Chen DX, Li Q, Wang XY. Effects of seminar teaching method versus lecture-based learning in medical education: a meta-analysis of randomized controlled trials. Med Teach. 2020;42(12):1343–9. 10.1080/0142159x.2020.1805100. [DOI] [PubMed] [Google Scholar]
- 3.Saragih ID, Suarilah I, Hsiao CT, Fann WC, Lee BO. Interdisciplinary simulation-based teaching and learning for healthcare professionals: a systematic review and meta-analysis of randomized controlled trials. Nurse Educ Pract. 2024;76:103920. 10.1016/j.nepr.2024.103920. [DOI] [PubMed] [Google Scholar]
- 4.Truchot J, Boucher V, Li W, Martel G, Jouhair E, Raymond-Dufresne É, et al. Is in situ simulation in emergency medicine safe? A scoping review. BMJ Open. 2022;12(7):e059442. 10.1136/bmjopen-2021-059442. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 5.Lee AJ, Goodman S, Corradini B, Cohn S, Chatterji M, Landau R. A serious video game—EmergenCSim™—for novice anesthesia trainees to learn how to perform general anesthesia for emergency Cesarean delivery: a randomized controlled trial. Anesthesiol Perioper Sci. 2023;1(2):14. 10.1007/s44254-023-00016-4. [Google Scholar]
- 6.Ng JY, Cramer H, Lee MS. Traditional, complementary, and integrative medicine and artificial intelligence: novel opportunities in healthcare. Integr Med Res. 2024;13(1):101024. 10.1016/j.imr.2024.101024. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 7.Busch F, Hoffmann L, Truhn D, Ortiz-Prado E, Makowski MR, Bressem KK, et al. Global cross-sectional student survey on AI in medical, dental, and veterinary education and practice at 192 faculties. BMC Med Educ. 2024;24(1):1066. 10.1186/s12909-024-06035-4. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 8.Preiksaitis C, Rose C. Opportunities, challenges, and future directions of generative artificial intelligence in medical education: scoping review. JMIR Med Educ. 2023;9:e48785. 10.2196/48785. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 9.Waldock WJ, Zhang J, Guni A, Nabeel A, Darzi A, Ashrafian H. The accuracy and capability of artificial intelligence solutions in health care examinations and certificates: systematic review and meta-analysis. J Med Internet Res. 2024;26:e56532. 10.2196/56532. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 10.Davies NP, Wilson R, Winder MS, Tunster SJ, McVicar K, Thakrar S, et al. ChatGPT sits the DFPH exam: large Language model performance and potential to support public health learning. BMC Med Educ. 2024;24(1):57. 10.1186/s12909-024-05042-9. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 11.Gilson A, Safranek CW, Huang T, Socrates V, Chi L, Taylor RA, et al. How does ChatGPT perform on The united States medical licensing examination (USMLE)? The implications of large Language models for medical education and knowledge assessment. JMIR Med Educ. 2023;9:e45312. 10.2196/45312. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 12.Skalidis I, Cagnina A, Luangphiphat W, Mahendiran T, Muller O, Abbe E, et al. ChatGPT takes on the European exam in core cardiology: an artificial intelligence success story? Eur Heart J Digit Health. 2023;4(3):279–81. 10.1093/ehjdh/ztad029. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 13.Freeman S, Eddy SL, McDonough M, Smith MK, Okoroafor N, Jordt H, et al. Active learning increases student performance in science, engineering, and mathematics. Proc Natl Acad Sci U S A. 2014;111(23):8410–5. 10.1073/pnas.1319030111. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 14.Boscardin CK, Gin B, Golde PB, Hauer KE. ChatGPT and generative artificial intelligence for medical education: potential impact and opportunity. Acad Med. 2024;99(1):22–7. 10.1097/acm.0000000000005439. [DOI] [PubMed] [Google Scholar]
- 15.Wu Y, Zheng Y, Feng B, Yang Y, Kang K, Zhao A. Embracing ChatGPT for medical education: exploring its impact on Doctors and medical students. JMIR Med Educ. 2024;10:e52483. 10.2196/52483. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 16.Xu T, Weng H, Liu F, Yang L, Luo Y, Ding Z, et al. Current status of ChatGPT use in medical education: potentials, challenges, and strategies. J Med Internet Res. 2024;26:e57896. 10.2196/57896. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 17.Page MJ, McKenzie JE, Bossuyt PM, Boutron I, Hoffmann TC, Mulrow CD, et al. The PRISMA 2020 statement: an updated guideline for reporting systematic reviews. Rev Esp Cardiol. 2021;74(9):790–9. 10.1016/j.rec.2021.07.010. [DOI] [PubMed] [Google Scholar]
- 18.Higgins JP, Altman DG, Gøtzsche PC, Jüni P, Moher D, Oxman AD, et al. The Cochrane collaboration’s tool for assessing risk of bias in randomised trials. BMJ. 2011;343:d5928. 10.1136/bmj.d5928. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 19.Brożek JL, Akl EA, Compalati E, Kreis J, Terracciano L, Fiocchi A, et al. Grading quality of evidence and strength of recommendations in clinical practice guidelines part 3 of 3. The GRADE approach to developing recommendations. Allergy. 2011;66(5):588–95. 10.1111/j.1398-9995.2010.02530.x. [DOI] [PubMed] [Google Scholar]
- 20.Cumpston M, Li T, Page MJ, Chandler J, Welch VA, Higgins JP, et al. Updated guidance for trusted systematic reviews: a new edition of the Cochrane handbook for systematic reviews of interventions. Cochrane Database Syst Rev. 2019;10(10):ED000142. 10.1002/14651858.Ed000142. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 21.Ba H, Zhang L, Yi Z. Enhancing clinical skills in pediatric trainees: a comparative study of ChatGPT-assisted and traditional teaching methods. BMC Med Educ. 2024;24(1):558. 10.1186/s12909-024-05565-1. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 22.Bhatia AP, Lambat A, Jain T. A comparative analysis of conventional and chat-generative pre-trained transformer-assisted teaching methods in undergraduate dental education. Cureus. 2024;16(5):e60006. 10.7759/cureus.60006. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 23.Çiçek FE, Ülker M, Özer M, Kiyak YS. ChatGPT versus expert feedback on clinical reasoning questions and their effect on learning: a randomized controlled trial. Postgrad Med J. 2024;458–63. 10.1093/postmj/qgae170. [DOI] [PubMed]
- 24.Gan W, Ouyang J, Li H, Xue Z, Zhang Y, Dong Q, et al. Integrating ChatGPT in orthopedic education for medical undergraduates: randomized controlled trial. J Med Internet Res. 2024;26:e57037. 10.2196/57037. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 25.Huang Y, Xu BB, Wang XY, Luo YC, Teng MM, Weng X. Implementation and evaluation of an optimized surgical clerkship teaching model utilizing ChatGPT. BMC Med Educ. 2024;24(1):1540. 10.1186/s12909-024-06575-9. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 26.Zeng H, Zhu ZW, Hu J, Cui Y. Application of ChatGPT-assisted problem-based learning teaching method in clinical medical education. BMC Med Educ. 2025;25(1):50. 10.1186/s12909-024-06321-1. [DOI] [PMC free article] [PubMed]
- 27.Jiang Y, Fu X, Wang J, Liu Q, Wang X, Liu P, et al. Enhancing medical education with chatbots: a randomized controlled trial on standardized patients for colorectal cancer. BMC Med Educ. 2024;24(1):1511. 10.1186/s12909-024-06530-8. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 28.Kavadella A, Dias da Silva MA, Kaklamanos EG, Stamatopoulos V, Giannakopoulos K. Evaluation of chatgpt’s real-life implementation in undergraduate dental education: mixed methods study. JMIR Med Educ. 2024;10:e51344. 10.2196/51344. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 29.Svendsen K, Askar M, Umer D, Halvorsen KH. Short-term learning effect of ChatGPT on pharmacy students’ learning. Explor Res Clin Soc Pharm. 2024;15:100478. 10.1016/j.rcsop.2024.100478. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 30.Tabuchi H, Nakajima I, Day M, Yoneda T, Tanabe M, Strang N, et al. Comparative educational effectiveness of AI generated images and traditional lectures for diagnosing Chalazion and sebaceous carcinoma. Sci Rep. 2024;14(1):29200. 10.1038/s41598-024-80732-4. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 31.Wu C, Chen L, Han M, Li Z, Yang N, Yu C. Application of ChatGPT-based blended medical teaching in clinical education of hepatobiliary surgery. Med Teach. 2024;47(3):445–9. 10.1080/0142159X.2024.2339412. [DOI] [PubMed] [Google Scholar]
- 32.Kreiter CD, Green J, Lenoch S, Saiki T. The overall impact of testing on medical student learning: quantitative Estimation of consequential validity. Adv Health Sci Educ Theory Pract. 2013;18(4):835–44. 10.1007/s10459-012-9395-7. [DOI] [PubMed] [Google Scholar]
- 33.George Pallivathukal R, Kyaw Soe HH, Donald PM, Samson RS, Hj Ismail. A R. ChatGPT for academic purposes: survey among undergraduate healthcare students in Malaysia. Cureus. 2024;16(1):e53032. 10.7759/cureus.53032. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 34.Cabellos B, de Aldama C, Pozo JI. University teachers’ beliefs about the use of generative artificial intelligence for teaching and learning. Front Psychol. 2024;15:1468900. 10.3389/fpsyg.2024.1468900. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 35.Borg A, Jobs B, Huss V, Gentline C, Espinosa F, Ruiz M, et al. Enhancing clinical reasoning skills for medical students: a qualitative comparison of LLM-powered social robotic versus computer-based virtual patients within rheumatology. Rheumatol Int. 2024;44(12):3041–51. 10.1007/s00296-024-05731-0. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 36.Arun G, Perumal V, Urias F, Ler YE, Tan BWT, Vallabhajosyula R, et al. ChatGPT versus a customized AI chatbot (Anatbuddy) for anatomy education: a comparative pilot study. Anat Sci Educ. 2024;17(7):1396–405. 10.1002/ase.2502. [DOI] [PubMed] [Google Scholar]
- 37.Holderried F, Stegemann-Philipps C, Herschbach L, Moldt JA, Nevins A, Griewatz J, et al. A generative pretrained transformer (GPT)-powered chatbot as a simulated patient to practice history taking: prospective, mixed methods study. JMIR Med Educ. 2024;10:e53961. 10.2196/53961. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 38.Komasawa N, Yokohira M. Simulation-based education in the artificial intelligence era. Cureus. 2023;15(6):e40940. 10.7759/cureus.40940. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 39.Cook DA. Creating virtual patients using large Language models: scalable, global, and low cost. Med Teach. 2025;47(1):40–2. 10.1080/0142159x.2024.2376879. [DOI] [PubMed] [Google Scholar]
- 40.Xu Y, Jiang Z, Ting DSW, Kow AW, C, Bello F, Car J, et al. Medical education and physician training in the era of artificial intelligence. Singap Med J. 2024;65(3):159–66. 10.4103/singaporemedj.SMJ-2023-203. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 41.Driesnack S, Rücker F, Dietze-Jergus N, Bondarenko A, Pletz MW, Viehweger A. A practice-based approach to teaching antimicrobial therapy using artificial intelligence and gamified learning. JAC Antimicrob Resist. 2024;6(4):dlae099. 10.1093/jacamr/dlae099. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 42.Ossa LA, Rost M, Lorenzini G, Shaw DM, Elger BS. A smarter perspective: learning with and from AI-cases. Artif Intell Med. 2023;135:102458. 10.1016/j.artmed.2022.102458. [DOI] [PubMed] [Google Scholar]
- 43.Vertemati M, Zuccotti GV, Porrini M. Enhancing anatomy education through flipped classroom and adaptive learning a pilot project on liver anatomy. J Med Educ Curric Dev. 2024;11:23821205241248023. 10.1177/23821205241248023. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 44.Gordon M, Daniel M, Ajiboye A, Uraiby H, Xu NY, Bartlett R, et al. A scoping review of artificial intelligence in medical education: BEME guide 84. Med Teach. 2024;46(4):446–70. 10.1080/0142159x.2024.2314198. [DOI] [PubMed] [Google Scholar]
- 45.Ullah M, Bin Naeem S, Kamel Boulos MN. Assessing the guidelines on the use of generative artificial intelligence tools in universities: a survey of the world’s top 50 universities. Big Data Cogn Comput. 2024;194. 10.3390/bdcc8120194.
- 46.Stogiannos N, Skelton E, Kumar S, Ahmed S, Amedu C, Vince C, et al. Evaluation of a customised, AI-focused educational seminar delivered to final year undergraduate radiography students in the UK: a cross-sectional study. Radiography. 2025;31(3):102926. 10.1016/j.radi.2025.102926. [DOI] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Data Availability Statement
The datasets and analyzed during the current study are available from the corresponding author on reasonable request.




