Abstract
Background
Generative artificial intelligence (GenAI) technologies —exemplified by large language models (LLMs) such as ChatGPT and DeepSeek—have attracted substantial interest in medical education for their capacity to generate content, adapt to individual learners, and support interactive dialogue. Whether these technologies produce measurable improvements in educational outcomes remains uncertain, and this evidence gap has constrained broader adoption. We conducted this systematic review and meta-analysis of randomized controlled trials (RCTs) to evaluate the effects of GenAI-assisted teaching on knowledge acquisition, clinical skills, and learner attitudes.
Methods
We conducted a systematic review and meta-analysis of randomized controlled trials (RCTs) following the PRISMA 2020 guidelines. The study protocol was registered on PROSPERO (CRD420261277187).Six databases (PubMed, Cochrane Library, Embase, Web of Science, CNKI, and Wanfang) were searched from January 2020 to February 2026. We included RCTs in which GenAI-integrated medical education was compared against traditional teaching or non-GenAI digital teaching, with participants being medical students, interns, or residents. Risk of bias was evaluated with the Cochrane RoB 2 tool; the GRADE framework was used to judge evidence certainty. Pooling was done through random-effects models with REML estimation.
Results
We identified 20 eligible RCTs (1413 participants in total). Post-intervention knowledge scores were higher in GenAI groups (SMD = 0.99, 95% CI: 0.49–1.49, P = 0.001; k = 11; I2 = 86.7%; GRADE: low certainty), and so were clinical skill scores (SMD = 1.09, 95% CI: 0.81–1.38, P < 0.001; k = 8; I2 = 41.0%; GRADE: moderate certainty). At follow-up, knowledge retention remained significant (SMD = 0.51, 95% CI: 0.24–0.78, P = 0.010; k = 4; I2 = 0%; GRADE: moderate certainty). Publication bias was not detected by Egger’s test (P = 0.104); trim-and-fill adjustment gave a similar pooled estimate (SMD = 0.90). Subgroup comparisons by AI tool type, learner level, interaction modality, and risk of bias were all non-significant.
Conclusion
GenAI-assisted teaching was associated with moderate-to-large pooled standardised mean differences for clinical skills (SMD = 1.09; moderate certainty) and post-intervention knowledge (SMD = 0.99; low certainty). However, the prediction interval for knowledge crossed zero (− 0.57 to 2.56), indicating that the direction of benefit cannot be assumed in future contexts. In four studies, knowledge gains were observed at follow-up (SMD = 0.51; moderate certainty), though the small number of studies limits confidence in durability. Evidence certainty is low to moderate, owing to high heterogeneity, risk of bias, and the absence of blinding. Current evidence is insufficient to draw definitive conclusions; GenAI may serve as a useful supplement to conventional instruction but should not be adopted as a replacement without evidence from well-powered, pre-registered, multi-centre trials.
Trial registration
PROSPERO registration: CRD420261277187.
Supplementary Information
The online version contains supplementary material available at 10.1186/s12909-026-09320-6.
Keywords: Generative artificial intelligence, Large language model, ChatGPT, Medical education, Meta-analysis, Randomized controlled trial
Introduction
Over the past few years, generative artificial intelligence (GenAI) technologies, and large language models (LLMs) in particular such as ChatGPT (OpenAI) and DeepSeek, have attracted growing interest within medical education circles. GPT-3 appeared in 2020 [1]; ChatGPT followed in late 2022 and became widely accessible. The number of studies examining whether these tools can actually help medical students, interns, and residents learn better has been growing since. What makes GenAI different from the computer-assisted instruction tools that came before it is a combination of capabilities: natural language dialogue, real-time contextual adaptation, on-demand feedback, and the ability to create reasonably realistic clinical simulations [2]. Educators have begun exploring a range of applications based on these capabilities: virtual patients powered by AI, tutoring systems that adapt to individual learners, tools for problem-based learning, case-based reasoning support, among others.
Several structural features of clinical education make it a plausible domain for GenAI application. Learning opportunities are unevenly distributed and depend on patient case mix. Access to standardized patients is constrained. The quality of bedside supervision varies by rotation, and faculty time is increasingly limited as curricula expand [3]. Lecture- and textbook-based approaches remain the foundation of most programmes, yet they are recognised as suboptimal for developing clinical reasoning, integrating complex information, or tracking rapidly evolving evidence. GenAI addresses some of these gaps by enabling on-demand, personalised interactions that do not depend on instructor availability. An LLM-based virtual patient, for example, offers unlimited access, a reproducible clinical narrative, and adaptive difficulty calibrated to the learner’s level—features that are difficult to replicate in conventional clinical settings [4].
That said, the published evidence tells a mixed story. Certain RCTs have found large gains in knowledge scores and clinical skills, with effect sizes that in some cases went beyond one standard deviation. Other trials, no less carefully designed, have produced null or negative results. The implication is that outcomes depend heavily on the specifics: how the GenAI tool is built into the teaching, the quality of what the AI produces, and the particular educational setting. Nissen et al. [5] provide one example: they ran a crossover trial (n = 154) and found that GPT-4’s personalized feedback did not perform better than generic expert-written feedback on key-feature exam questions. The AI feedback quality was uneven; sometimes the explanations for wrong answers were too brief, and other times students were penalized for trivial spelling mistakes. Another example comes from Hirosawa et al. [6], who found lower clinical performance scores with AI-mediated encounters than with face-to-face standardized patient sessions. These studies raise a practical question: whether AI-generated content actually improves on well-prepared traditional materials in certain circumstances.
Systematic reviews that have looked at this area so far have mostly dealt with AI in education more broadly, or they have had to work with a limited number of eligible trials, non-randomized study designs, or a single AI tool as their focus [2, 3]Existing reviews have summarized AI-assisted medical education, and more recent evidence syntheses have begun to examine GenAI-based interventions. However, available reviews have generally been limited by mixed study designs, narrower search coverage, or fewer trials in Chinese-language databases. We therefore conducted an updated RCT-only systematic review and meta-analysis with broader database coverage and additional certainty assessments.
We set out to fill this gap. The primary aim was to pool the effects of GenAI-integrated teaching on knowledge and clinical skill scores after the intervention, using traditional or non-GenAI teaching as the comparator, and to judge how certain we can be about the evidence. Beyond this, we wanted to check whether any learning benefits last at follow-up, to look at what might moderate the treatment effects, to test for publication bias, and to map out where more research is most needed.
Methods
Study registration and reporting guidelines
This systematic review and meta-analysis was conducted in accordance with the Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA) 2020 guidelines [7] and the methodological standards of the Cochrane Handbook for Systematic Reviews of Interventions (version 6.4) [8]. The review protocol was registered with PROSPERO (CRD420261277187) on 10 January 2026, following completion of the initial literature search. The final updated database search was completed on 1 February 2026. All study selection, data extraction, and statistical analyses were conducted after registration. The registered protocol documents the pre-specified research questions, inclusion and exclusion criteria, outcome hierarchy, and statistical analysis plan, and no post-hoc changes were made to these elements. The retrospective nature of registration is acknowledged as a methodological limitation in the Limitations section.
Registration details:
PROSPERO registration number: CRD420261277187;
Date of registration:10 January 2026;
Title of registered protocol: “The Impact of Integrating Generative Artificial Intelligence into Medical Education on Short-Term Learning Outcomes: A Systematic Review and Meta-Analysis of Randomized Controlled Trials.”
The registered protocol outlines the primary research questions, inclusion/exclusion criteria, outcome measures, and statistical analysis plan.
Eligibility criteria
We used the PICOS framework to set the inclusion criteria. For population (P), we included medical students at any training stage, preclinical, clerkship, internship, as well as residents and clinical fellows. The intervention (I) was GenAI-integrated teaching; we defined this broadly as any program that used a large language model or similar generative system to produce text, dialogue, or multimedia as part of the educational process. Tools that qualified included ChatGPT (any version), DeepSeek, Coze-based platforms, custom LLM-based virtual patients, and others. The comparator (C) was teaching that did not use GenAI, whether traditional (lectures, bedside teaching, case discussions) or digital. For outcomes (O), studies needed to report at least one quantitative education measure: exam scores, knowledge tests, clinical or procedural skill scores, attitude questionnaires, efficiency metrics, etc. Study design (S) was restricted to RCTs, either parallel-group or crossover.
The time window was January 2020 to February 2026. We chose January 2020 as the start to avoid inadvertently excluding any early-stage generative AI studies; we acknowledge that GPT-3 [1] was not publicly accessible via API until June 2020, and educational applications of generative AI did not appear in the peer-reviewed literature until 2022–2023. Consistent with this, all 20 included studies were published in 2024–2025. The January 2020 start date therefore did not add studies from earlier years but provides a transparent, reproducible boundary. The final search was completed on February 1, 2026.
The exclusion criteria were: studies from which quantitative data could not be extracted; duplicate reports; interventions involving non-generative AI technologies; participants outside the medical training pipeline; and non-primary research publications including case reports, letters, editorials, commentaries, and narrative reviews.
Information sources and search strategy
A systematic search was carried out across six electronic databases: PubMed, the Cochrane Central Register of Controlled Trials (CENTRAL), Embase, Web of Science, the Chinese National Knowledge Infrastructure (CNKI), and Wanfang Data. Two categories of search terms were combined using Boolean operators: GenAI-related terms (“Generative AI,” “artificial intelligence,” “AI-assisted learning,” “large language model,” “ChatGPT,” “LLM,” “DeepSeek,” “GPT”) were intersected with medical education terms (“medical education,” “clinical teaching,” “medical students,” “residency training”). Reference lists of included studies and relevant systematic reviews were hand-searched to identify additional eligible studies. The final search was executed on February 1, 2026.
Study selection
All records were imported into reference management software, and duplicates were removed. Two reviewers independently screened titles and abstracts against the eligibility criteria. Full-text articles of potentially eligible studies were then independently assessed. Disagreements were resolved through discussion and, when necessary, adjudication by a third reviewer. The screening process was documented using a PRISMA flow diagram (Fig. 1).
Fig. 1.
PRISMA 2020 flow diagram. 5,389 records were identified; after screening, 20 RCTs were included
Data extraction
A total of 20 studies were included in this meta-analysis [5, 6, 9–26].Data were independently extracted by two reviewers using a standardized form, collecting: first author, year of publication, country, study design, participant characteristics, GenAI tool and model version, comparator description, study duration, outcome measures, and statistical data. For studies reporting medians and interquartile ranges, the methods of Wan et al. [27] were used to estimate means and standard deviations. For crossover trials, a conservative within-person correlation coefficient (r = 0.5) was assumed, with sensitivity analyses at r = 0.3 and r = 0.7. For multi-arm trials, shared control groups were appropriately split to avoid unit-of-analysis errors [8].
Outcome classification
Following Kirkpatrick’s hierarchy [28], we used a modified four-level framework as an organisational heuristic for outcome classification. All outcomes in this review correspond to Kirkpatrick Level 2 (Learning); no included study measured Level 3 (behavioural transfer to practice) or Level 4 (patient outcomes), which is an important limitation of the current evidence base. Within Level 2, we distinguished sub-domains by assessment instrument: knowledge (declarative exam and critical thinking tests), skills (performance-based assessments: Mini-CEX, OSCE, medical history-taking), attitudes (self-report rating scales capturing dispositions, motivation, and perceived abilities, including self-directed learning and clinical reasoning scales), and efficiency (task-completion metrics such as preparation time and time-to-diagnosis). We acknowledge that constructs such as self-directed learning ability and clinical reasoning ability span the boundary between attitudes and skills; their placement in the attitudes domain reflects the self-report instrument used rather than an ontological claim. Each outcome was tagged by measurement timing: post-intervention (within one week) or follow-up (at least two weeks after).
Risk of bias assessment
Methodological quality was independently evaluated using the Cochrane Risk of Bias 2 (RoB 2) tool [7], assessing five domains: randomization process (D1), deviations from intended interventions (D2), missing outcome data (D3), measurement of the outcome (D4), and selection of the reported result (D5). Each domain was judged as “low risk,” “some concerns,” or “high risk.” Disagreements were resolved by consensus.
Statistical analysis
We ran all meta-analyses in R (version 4.3) using the metafor package [29]. Because the included studies used different scales, we expressed effect sizes as standardized mean differences (SMD, Hedges’ g) and fitted random-effects models with restricted maximum likelihood (REML) estimation for the heterogeneity variance parameter τ2 throughout. Inference (confidence intervals and P-values) was based on the Hartung–Knapp (HK) adjustment, which uses a t-distribution with k − 1 degrees of freedom and is recommended for small meta-analyses to avoid anti-conservative coverage; the forest plots are accordingly labelled “random effects (HK)”. To quantify heterogeneity we used I2, Cochran’s Q, τ2, and 95% prediction intervals. Subgroup analyses, specified in advance, split the data by AI tool type (ChatGPT family, DeepSeek family, other GenAI), learner level (undergraduate vs. clinical trainee), interaction modality (text-only vs. multimodal), and overall risk of bias (high risk vs. some concerns). We also ran univariate meta-regression on total sample size and individual study risk-of-bias category as study-level covariates, leave-one-out sensitivity analyses, some-concerns-only analyses (restricting to studies rated “some concerns” overall, i.e., excluding the five high-risk studies), and fixed-effect comparisons. No study achieved overall low risk of bias, so a low-risk-only sensitivity analysis was not feasible. Given the number of pre-specified outcome domains and exploratory subgroup comparisons (approximately 20 statistical tests), findings from secondary analyses should be treated as hypothesis-generating rather than confirmatory. Outcome domains with only two contributing studies (k = 2) are presented narratively rather than as pooled estimates, given the extreme imprecision of between-study variance estimates at this sample size.
Publication bias was assessed for the primary outcome using Egger’s linear regression test [30] and the Duval and Tweedie trim-and-fill method [31].
Certainty of evidence (GRADE)
Evidence certainty was rated using the GRADE framework [32], looking at five domains: risk of bias, inconsistency, indirectness, imprecision, and publication bias. RCT evidence starts at “high” and gets downgraded when there are serious concerns. Final ratings were high, moderate, low, or very low.
Results
Study selection
Searching the six databases produced 5,389 records. We removed 969 duplicates, leaving 4,420 records for title-and-abstract screening. Of these, 4,280 did not meet the eligibility criteria and were excluded. The remaining 140 reports were subjected to further screening; 79 were excluded without full-text retrieval based on abstract-level review (63 with inconsistent study type, 11 with incomplete or unextractable data, and 5 with mismatch in theme or study subjects), leaving 61 full texts for detailed eligibility assessment. At the full-text stage we excluded 41: 19 because the population was outside our scope, 11 because the design was not an RCT, 3 because GenAI was not the intervention, 1 for an inappropriate comparator, 1 for insufficient data, 1 as a duplicate, and 5 for incomplete or unreliable data. The remaining 20 studies went into the quantitative synthesis (Fig. 1).
Study characteristics
The 20 RCTs were all published in 2024 or 2025 and had a combined enrollment of 1,413 participants. China accounted for 16 of them; the remaining four came from Italy, Japan, Germany, and the United States. Eighteen were parallel-group trials; two were crossover. Sample sizes ranged from 20 to 154, with a median of 60. Undergraduates or clinical interns made up the population in 13 studies; postgraduate residents in six; mixed groups in one. Specialties covered were broad: pediatrics, urology, hepatobiliary surgery, orthopedics, ophthalmology, nephrology, obstetrics, lung cancer targeted therapy, critical care, oncology, general medicine.
ChatGPT (versions 3.5 and 4.0) was the GenAI tool used most often (nine studies), followed by DeepSeek-R1 (three studies). Eight studies used other tools: the Coze AI learning platform, LLM-based digital patient systems, a custom GPT called LearnGuide, and multi-tool setups. What was used as the control differed from study to study: traditional lectures, bedside teaching, internet-based self-study, case-based learning without AI, and expert-written feedback all appeared. Interventions lasted anywhere from one session to 12 weeks. The detailed characteristics of all included studies are summarized in Table 1 [5, 6, 9–26].
Table 1.
Characteristics of included studies (n = 20 RCTs)
| Study | Country | Design | Population | N | GenAI Tool | Model | Comparator | RoB |
|---|---|---|---|---|---|---|---|---|
| Ba 2024 [19] | China | Parallel | Medical interns | 77 | ChatGPT | GPT-4.0 | Bedside teaching | SC |
| Chen 2025 [20] | China | Parallel | Undergraduates | 40 | Coze (AI-PLP) | BERT + DQN | Lecture-based | HR |
| Digiacomo 2025 [17] | Italy | Parallel (3-arm) | Undergraduates | 59 | ChatGPT | GPT-3.5 | Traditional lecture | SC |
| Gan 2024 [16] | China | Parallel | Undergraduates | 110 | ChatGPT | GPT-4.0 | Internet self-study | SC |
| Hirosawa 2025 [6] | Japan | Crossover | Residents (PGY1-2) | 20 | Custom GPTs | GPTs | Standardized patient | SC |
| Hui 2025 [21] | China | Parallel | Medical interns | 42 | ChatGPT | GPT-4.0 | Traditional teaching | SC |
| Kalam 2025 [18] | USA | Parallel (3-arm) | Undergraduates (M1) | 21 | ChatGPT | GPT-4.0 | Institutional resources | HR |
| Liang 2025 [22] | China | Parallel | Undergraduates | 34 | ChatGPT + DeepSeek | Multiple | Traditional teaching | HR |
| Liu 2025 [12] | China | Parallel | Clinical clerkship | 120 | DeepSeek-R1 | R1 v2.3.1 | Traditional teaching | SC |
| Luo 2025 [14] | China | Parallel | Undergraduates (Y4) | 84 | LLMDP | Baichuan-13B | Real patient training | SC |
| Nissen 2025 [5] | Germany | Crossover | Undergraduates (Y4) | 154 | GPT-4 | GPT-4-turbo | Expert feedback | SC |
| Tan 2025 [23] | China | Parallel | Postgraduates | 41 | ChatGPT | GPT-4.0 | Traditional FC | SC |
| Tong 2025 [10] | China | Parallel | Residents | 54 | ChatGPT | NR | Traditional teaching | SC |
| Wang 2025a [9] | China | Parallel | Mixed (UG + PG) | 103 | LearnGuide | Custom GPT | PBL without AI | SC |
| Wang 2025b [13] | China | Parallel | Undergraduates | 154 | Multiple GenAI | Multiple | Multimedia teaching | SC |
| Wang 2025c [24] | China | Parallel | Residents (M1) | 20 | LLM-VP system | NR | Traditional teaching | HR |
| Wu 2025a [11] | China | Parallel | ICU residents | 96 | DeepSeek-R1 | R1 (671B) | Traditional resources | SC |
| Wu 2025b [15] | China | Parallel | Medical interns | 61 | ChatGPT | GPT-3.5 | Traditional teaching | SC |
| Zhang 2025a [25] | China | Parallel | Medical interns | 80 | DeepSeek-R1 | R1 | Traditional teaching | SC |
| Zhang 2025b [26] | China | Parallel | Postgraduates | 43 | GenAI (NR) | NR | Traditional teaching | HR |
SC Some concerns, HR High risk, NR Not reported, FC Flipped classroom, UG Undergraduate, PG Postgraduate, M1 first-year master, Y4 fourth year, VP virtual patient
Risk of bias
According to the RoB 2 assessment, not one of the 20 studies qualified as low risk of bias overall. Fifteen (75%) had “some concerns” and five (25%) were “high risk” (Fig. 2). At the domain level: randomization (D1) was low risk in nine studies (45%) and had some concerns in 11 (55%), mostly because allocation concealment was not described well enough. The domain that performed worst was deviations from intended interventions (D2), with four studies (20%) rated high risk, 14 (70%) had some concerns, and the reason was almost always that it is hard to blind people to a GenAI-based intervention. Missing outcome data (D3) was low risk in 17 studies (85%). Outcome measurement (D4) raised some concerns in 11 studies (55%) because subjective assessments were used without blinding the evaluators. Reported result selection (D5) was high risk in two studies (10%); most studies had no pre-registered protocol (Fig. 2).
Fig. 2.
Risk of bias summary showing the proportion of studies judged as low risk, some concerns, and high risk for each domain of the Cochrane RoB 2 tool
Primary outcomes
Eleven studies (796 participants) contributed to the knowledge score meta-analysis. The pooled SMD was 0.99 (95% CI: 0.49 to 1.49, P = 0.001), a large average effect favouring GenAI-assisted teaching (Fig. 3). However, heterogeneity was substantial (I2 = 86.7%, τ2 = 0.44, Q = 75.0, P < 0.001), and the 95% prediction interval (− 0.57 to 2.56) crossed zero, indicating that the true effect is likely to vary considerably across settings and that a negative effect in future individual studies cannot be excluded. The point estimate should therefore be interpreted as an average over a heterogeneous set of comparisons rather than as a reliable expected effect in any specific context. A fixed-effect model yielded a lower estimate (SMD = 0.81, 95% CI: 0.66 to 0.96), consistent with smaller studies having disproportionately extreme effect sizes.
Fig. 3.
Forest plot of post-intervention knowledge scores. Random-effects model: pooled SMD = 0.99 (95% CI: 0.49–1.49). k = 11, n = 796
The clinical skills analysis included eight studies and 552 participants. The pooled SMD was 1.09 (95% CI: 0.81 to 1.38, P < 0.001), a large effect favouring GenAI-assisted teaching (Fig. 4). Heterogeneity was moderate (I2 = 41.0%, τ2 = 0.04, Q = 11.9, P = 0.105), and the 95% prediction interval (0.54 to 1.65) remained entirely above zero, indicating that a positive effect on clinical skills would be expected across a range of settings (Table 2). The fixed-effect number was close: SMD = 1.03 (95% CI: 0.86 to 1.19).
Fig. 4.
Forest plot of post-intervention clinical skill scores. Random-effects model: pooled SMD = 1.09 (95% CI: 0.81–1.38). k = 8, n = 552
Table 2.
Summary of pooled effect estimates for all outcome domains
| Outcome (Timing) | SMD | k | 95% CI | P | I2 (%) | τ2 | 95% PI |
|---|---|---|---|---|---|---|---|
| Knowledge (Post) | 0.99 | 11 | 0.49 to 1.49 | 0.001 | 86.7 | 0.44 | − 0.57 to 2.56 |
| Skills (Post) | 1.09 | 8 | 0.81 to 1.38 | < 0.001 | 41.0 | 0.04 | 0.54 to 1.65 |
| Knowledge (Follow-up) | 0.51 | 4 | 0.24 to 0.78 | 0.010 | 0 | 0 | 0.10 to 0.92 |
| Skills (Follow-up) | —† | 2 | —† | —† | —† | —† | —† |
| Attitude (Follow-up) | —† | 2 | —† | —† | —† | —† | —† |
| Efficiency (During) | —† | 2 | —† | —† | —† | —† | —† |
SMD Standardized mean difference (Hedges’ g), k = number of studies, CI Confidence interval, PI Prediction interval, Post Post-intervention
Secondary outcomes
To assess knowledge retention, we examined assessments administered at least two weeks after the teaching period ended. Four studies (212 participants) contributed follow-up knowledge data. The pooled SMD was 0.51 (95% CI: 0.24 to 0.78, P = 0.010); heterogeneity was negligible (I2 = 0%, Q = 1.31, P = 0.726). The prediction interval (0.10 to 0.92) was entirely positive, suggesting that knowledge gains are sustained beyond the immediate post-intervention period in the contexts studied, although the evidence base (k = 4) warrants caution in generalising this conclusion.
Only two studies (61 participants) reported skill assessments at follow-up. Because pooling two studies in a random-effects model yields an imprecise estimate of between-study variance, these data are presented narratively rather than as a pooled effect. Tan et al. [23] (n = 41) assessed literature review skills across six sub-dimensions at one month post-intervention, finding large effects consistently favouring GenAI (SMDs ranging from 1.73 to 2.68). Wang et al. [24] (n = 20) assessed medical record documentation quality at two months, also finding a large positive effect favouring GenAI (SMD ≈ 1.05). Both studies indicated sustained skill benefits, though the small sample sizes and the statistical limitations of k = 2 pooling preclude firm conclusions about long-term skill retention.
Two studies with 60 participants contributed data on learner attitudes assessed at follow-up. The pooled SMD was 0.61 (95% CI: − 3.62 to 4.85, P = 0.317), which was not statistically significant; the severe imprecision resulted in a very low GRADE certainty rating. One study (Ba et al., n = 77) that reported a dichotomous attitude measure at the post-intervention timepoint found an odds ratio of 3.09 (95% CI: 1.21 to 7.89) favoring the GenAI group, which may mean that learners view GenAI-assisted teaching more positively, although the evidence from pooled analyses remains inconclusive.
Two studies (161 participants) reported learning efficiency during the intervention. Given that only two estimates are available and heterogeneity was high (I2 = 78.6%), these findings are described narratively rather than pooled. Tan et al. [23] (n = 41) assessed efficiency-related outcomes within a flipped classroom ChatGPT curriculum for graduate medical education. Liu et al. [12] (n = 120) found that GenAI-assisted students acquired knowledge at a higher rate than controls (18.7 vs 13.4 items/hour) and completed clinical decision tasks more quickly (14.2 vs 21.8 min). These individual findings suggest potential efficiency benefits of GenAI that may be task-specific and require replication in adequately powered studies.
Subgroup analyses
We ran pre-specified subgroup analyses on both primary outcomes where enough data existed. Effect sizes were positive across every stratification, with no significant between-subgroup differences (Table 3).
Table 3.
Subgroup analyses for post-intervention knowledge and clinical skill scores
| Outcome | Subgroup | k | SMD (95% CI) | P-diff |
|---|---|---|---|---|
| Knowledge | ChatGPT family | 7 | 1.07 (0.32, 1.82) | 0.178 |
| DeepSeek family | 2 | 1.17 (− 8.64, 10.99) | ||
| Other GenAI | 2 | 0.54 (− 0.49, 1.57) | ||
| Knowledge | Undergraduate | 5 | 1.28 (0.20, 2.37) | 0.327 |
| Clinical trainee | 5 | 0.80 (− 0.05, 1.65) | ||
| Knowledge | Multimodal | 4 | 0.95 (− 0.58, 2.48) | 0.888 |
| Pure text | 6 | 1.03 (0.28, 1.79) | ||
| Knowledge | Not high risk | 8 | 0.85 (0.27, 1.43) | 0.253 |
| High risk | 3 | 1.47 (− 0.60, 3.53) | ||
| Skills | ChatGPT family | 4 | 1.28 (0.47, 2.08) | 0.221 |
| Other GenAI | 3 | 0.94 (0.48, 1.40) | ||
| Skills | Undergraduate | 3 | 1.06 (0.22, 1.90) | 0.754 |
| Clinical trainee | 4 | 1.15 (0.43, 1.87) | ||
| Skills | Not high risk | 6 | 1.05 (0.69, 1.42) | 0.315 |
| High risk | 2 | 1.31 (− 1.48, 4.10) |
P-diff P value for between-subgroup differences
†Within-subgroup P values omitted for brevity. k = number of studies
Sensitivity analyses
Dropping each study in turn from the knowledge meta-analysis (k = 11) did not change the significance of the pooled SMD; estimates fell between roughly 0.86 and 1.10. No single study was driving the result. The same was true for the skills analysis (k = 8), where pooled SMDs when omitting each study in turn ranged from approximately 0.97 to 1.17, confirming that no individual study had a disproportionate influence on the skills estimate.
When the analysis was restricted to studies with overall “some concerns” (excluding all high-risk studies), the pooled SMD for knowledge was 0.85 (95% CI: 0.27 to 1.43, P = 0.011; k = 8) and for skills was 1.05 (95% CI: 0.69 to 1.42, P < 0.001; k = 6). Both estimates were modestly attenuated but remained large and statistically significant. For the two crossover trials (Nissen et al. [5] and Hirosawa et al. [6]), the primary analysis assumed a within-person correlation of r = 0.5. Sensitivity analyses at r = 0.25 and r = 0.75 confirmed that the direction of effect was not sensitive to the assumed correlation. For Nissen et al. [5] (knowledge; n = 154 pairs), the mean difference remained − 4 points across all r values; the 95% CI ranged from (− 8.33 to 0.33) at r = 0.25 to (− 6.50 to − 1.50) at r = 0.75, directionally consistent with a null or marginally negative effect of GenAI feedback compared with expert feedback. For Hirosawa et al. [6] (skills; n = 20 pairs), the mean difference was − 0.58 across all r values; the 95% CI ranged from (− 0.745 to − 0.415) at r = 0.25 to (− 0.679 to − 0.481) at r = 0.75, with no change in direction. Both crossover studies were analysed as individual study results rather than pooled, given that neither domain contributed more than one crossover study.
Meta-regression
Univariate meta-regression for post-intervention knowledge scores examined total sample size as a study-level covariate. Sample size was not a statistically significant moderator (β = − 0.25, SE = 0.23, P = 0.290), indicating that effect size did not decrease systematically with increasing sample size. Individual study risk-of-bias category (coded as a binary study-level variable: high risk vs. some concerns) was also examined; this covariate was likewise non-significant (β = 0.61, SE = 0.52, P = 0.241). The observed heterogeneity is therefore unlikely to be explained by study size or risk-of-bias category alone, and residual unexplained heterogeneity remains substantial.
Publication bias
We ran Egger’s test on the knowledge funnel plot: intercept = 1.81, P = 0.104 (not significant). However, with only k = 11 studies, Egger’s test has limited statistical power to detect publication bias (Sterne et al., 2011); a non-significant result should not be interpreted as evidence that bias is absent. Visually, the funnel plot showed several large positive outliers and few negative results, consistent with small-study effects. Trim-and-fill added one imputed study and gave an adjusted SMD of 0.90 (95% CI: 0.46 to 1.35), suggesting modest downward bias may be present and that the unadjusted estimate of 0.99 may be a slight overestimate (Fig. 5).
Fig. 5.
Contour-enhanced funnel plot with trim-and-fill analysis for post-intervention knowledge scores. Open circle = one imputed study. Adjusted pooled SMD = 0.90 (95% CI: 0.46–1.35). Egger’s test P = 0.104
Certainty of evidence (GRADE)
GRADE ratings (Table 4): for post-intervention knowledge (k = 11, n = 796), certainty was rated low, downgraded once for risk of bias, once for inconsistency (I2 of 86.7%, prediction interval crossing zero). Post-intervention clinical skills (k = 8, n = 552) got moderate certainty, downgraded once for risk of bias. Follow-up knowledge retention (k = 4, n = 212) was also moderate, same reason. Everything else (follow-up attitudes, follow-up skills, efficiency during the intervention) ended up at very low certainty because of problems in multiple GRADE domains.
Table 4.
GRADE summary of findings
| Outcome (Timing) | SMD (95% CI) | k | n | Downgrade reasons | Certainty | Direction | Interpretation |
|---|---|---|---|---|---|---|---|
| Knowledge (Post) | 0.99 (0.49, 1.49) | 11 | 796 | RoB − 1; Inconsistency − 1 | Low | ↑ GenAI | True effect may differ substantially |
| Skills (Post) | 1.09 (0.81, 1.38) | 8 | 552 | RoB − 1 | Moderate | ↑ GenAI | Moderately confident in estimate |
| Knowledge (F/U) | 0.51 (0.24, 0.78) | 4 | 212 | RoB − 1 | Moderate | ↑ GenAI | Moderately confident in estimate |
| Attitudes (F/U) | —† | 2 | 60 | RoB − 2; Incon − 1; Imprec − 2 | Very low | Uncertain | Very uncertain evidence |
| Efficiency (During) | —† | 2 | 161 | RoB − 1; Incon − 1; Imprec − 2 | Very low | Uncertain | Very uncertain evidence |
| Skills (F/U) | —† | 2 | 61 | RoB − 1; Incon − 1; Imprec − 2 | Very low | Uncertain | Very uncertain evidence |
Certainty starts at High for RCTs and is downgraded per GRADE criteria
RoB Risk of bias, Incon Inconsistency, Imprec Imprecision, F/U Follow-up
Discussion
This review synthesized 20 RCTs and provides an updated estimate of the effects of GenAI-assisted teaching in medical education. The pooled SMD was 0.99 for knowledge and 1.09 for clinical skills; follow-up knowledge retention was also significant at 0.51. Among all outcomes, clinical skills had the firmest evidence base (moderate GRADE certainty, prediction interval 0.54 to 1.65 entirely above zero). The knowledge result, by contrast, came with high heterogeneity (I2 = 86.7%) and a prediction interval that crossed zero (− 0.57 to 2.56), so the real-world benefit is likely to differ from one setting to another. Exam scores remain the standard metric for learning assessment in most trials [2, 3], but they are an imperfect measure. Part of the between-study spread in our knowledge analysis can probably be traced to differences in how GenAI was used. Some trials ran for only a few weeks or grafted AI onto existing courses without much adaptation; longer, more deeply integrated implementations may be needed for consistent gains to emerge [4]. Students’ prior familiarity with AI tools varied across studies, and no trial described a standardized approach to training instructors on how to use GenAI in teaching. These unmeasured sources of variation likely account for at least some of the observed heterogeneity.
Two crossover trials ran counter to the overall trend. Nissen et al. [5] gave 154 students either GPT-4 personalized feedback or generic expert-written feedback; the AI group scored slightly (non-significantly) lower. On inspection, the AI feedback turned out to be uneven: some explanations for wrong answers were too brief, while trivial spelling errors were sometimes penalized. Hirosawa et al. [6] compared AI-mediated and face-to-face standardized patient encounters and again found lower scores in the AI condition. Text-based communication may simply work differently from spoken interaction, affecting performance in ways that have little to do with content quality. The lesson from both studies is the same: GenAI does not guarantee improvement. How the tool is designed, implemented, and fitted into the teaching program matters just as much as whether it is used.
The clinical skills findings were more stable. I2 was 41.0%, and the prediction interval did not cross zero. Skill acquisition, unlike factual recall, depends on repeated practice: working through cases, making decisions, receiving feedback, then trying again. GenAI suits this pattern because it gives students an active role; they have to search for information, frame questions, and reason their way through problems [4]. Luo et al. [14] found that their LLM-based digital patient raised history-taking scores across every subdomain they measured (demographics, present illness, past history, etc.), with voice-based feedback that correlated well with human raters (r = 0.88). Tong et al. [10] and Wu et al. [15] both saw greater gains in case analysis and diagnostic reasoning in GenAI groups, though basic factual recall was similar. Liu et al. [12] reported the biggest improvements in deep-comprehension items in targeted lung cancer therapy, not in basic knowledge questions, which aligns with the idea that GenAI does more for the upper tiers of Bloom’s taxonomy than for memorization. On the practical side, GenAI also addresses some constraints that have long limited clinical teaching: restricted instructor availability, difficulty personalizing standardized curricula, and the narrow range of cases students encounter during a single rotation [4, 9]. Several of the included studies noted that GenAI groups shifted their study time toward pre-class preparation and away from post-class review [10, 13], while total hours stayed about the same, a redistribution consistent with the flipped-classroom model.
None of the subgroup comparisons we ran (AI tool type, learner level, modality, risk of bias) reached statistical significance. The subgroups were small (k = 2 to 8), though, so the power to detect real differences was limited. Effect sizes for ChatGPT and for other GenAI tools were numerically close, which is at least consistent with the idea that the benefit comes from LLM technology in general rather than from any one commercial product. If future studies with larger samples confirm this, it would have practical relevance for institutions facing restrictions on platform access due to cost, regional policy, or privacy regulations. We also saw a numerical gap by training stage: SMD 1.28 for undergraduates versus 0.80 for clinical trainees. Earlier-stage students may respond more to the structured, step-by-step interactions that GenAI provides, while more advanced trainees may need different challenges (complex diagnostic reasoning, procedural simulations) to see comparable improvement [4]. It is also worth recognizing that GenAI, at present, works as a supplement to traditional teaching rather than an alternative to it [2, 3]. Trials that embedded GenAI within structured formats such as PBL, case-based modules, or supervised rotations reported bigger and more consistent effects than those offering open-ended AI access [10, 11, 13]. Wang et al. [13] built a custom GPT (LearnGuide) into a 12-week PBL course and saw gains in self-directed learning and critical thinking that lasted through 14-week follow-up, one of the longest tracking periods in this evidence base.
AI literacy is another area that needs attention. Medical education currently lacks a shared framework for deciding how GenAI should be woven into courses, what resources are needed, and what students should learn about AI itself [14, 16]. The known weaknesses of these tools make the gap more urgent. GenAI can hallucinate, producing inaccurate clinical recommendations or unsupported references, which poses risks for evidence-based medicine training [9].Output quality can also fluctuate from one prompt to the next. There is a further concern that students who rely heavily on AI assistance may lose initiative in their own learning [14, 16]. Wu et al. [11] provided a striking example from critical care: in the critical care trial, DeepSeek-R1 alone achieved 60% diagnostic accuracy, and physicians assisted by the model reached 58%, substantially higher than physicians without AI support, while the time to diagnosis was also reduced. These findings suggest that GenAI may improve performance in complex diagnostic tasks when used as an adjunct rather than a replacement for clinician judgment. Structured training in how to evaluate and use AI output is clearly needed for both students and faculty.
Methodological quality is a weakness across this evidence base. Not one of the 20 studies was rated as overall low risk of bias; 75% had “some concerns” and 25% were high risk. The two domains with the most problems were deviations from intended interventions (D2) and selection of the reported result (D5) [33]. Blinding is structurally impossible when participants use a recognisable AI tool, and none of the 20 included studies implemented a sham AI interface or attention-matched control. Eleven studies (55%) received “some concerns” on RoB 2 domain D4 (measurement of the outcome), reflecting non-blinded subjective assessment. The magnitude and direction of the resulting performance bias cannot be quantified; some portion of the observed benefit likely reflects the Hawthorne effect or novelty enthusiasm rather than genuine learning gains, and this uncertainty is not fully resolved by the sensitivity analyses. The interventions themselves were highly heterogeneous: they ranged from a single ChatGPT interaction session to 12-week custom-GPT-integrated PBL courses, using tools that included ChatGPT 3.5, ChatGPT 4.0, DeepSeek-R1, Coze-based platforms, and bespoke LLMDP systems, with intervention durations spanning one session to 12 weeks. The assumption that these diverse strategies share a common mechanism of action has not been established; the pooled SMD therefore represents an average effect across pedagogically diverse implementations rather than the effect of any particular tool or delivery mode. Several included studies used proprietary AI systems (nine using OpenAI’s ChatGPT, three using DeepSeek-R1) whose training data are opaque, model versions evolve without public announcement, and outputs cannot be guaranteed to be reproducible. Student interaction data submitted to third-party commercial servers may also raise institutional data-privacy concerns. These ethical dimensions—reproducibility, data sovereignty, equitable access for institutions without API funding—are not addressed by the included trials and should be incorporated into the design of future research and into institutional governance frameworks. Protocol pre-registration was uncommon, which raises concerns about selective outcome reporting. GRADE certainty (low for knowledge, moderate for skills) reflects these issues. Restricting the analysis to studies rated “some concerns” attenuated effects modestly (knowledge SMD = 0.85; skills SMD = 1.05) but did not eliminate them, and univariate meta-regression found that neither sample size nor individual study risk-of-bias category explained a significant portion of the observed heterogeneity. Other limitations include the small evidence base for some outcome domains (k = 2–4), modest sample sizes overall (median n = 60), and participants’ limited prior AI experience, which may conflate novelty effects with genuine learning gains. Geographic concentration is a further concern: 16 of 20 studies were conducted in China, and educational systems, assessment cultures, and regulatory environments for AI tools differ internationally in ways that may affect generalisability. An associated limitation is database language asymmetry: while CNKI and Wanfang were included specifically to capture Chinese-language evidence, this broadening also biases the identified literature toward Chinese institutions, which dominated the results. English-language databases (PubMed, Embase, Cochrane, Web of Science) returned relatively few eligible trials from European, North American, or low- and middle-income country settings. This is likely to reflect both the genuine preponderance of Chinese GenAI educational research in 2024–2025 and a possible regional publication bias (positive findings may be more publishable in settings where national AI adoption policies are active). Databases covering Japanese, Korean, Persian, Arabic, or Spanish literature were not searched, which may have introduced language restriction bias. A comparator-type subgroup analysis (e.g., traditional lecture vs. active digital control) was not pre-specified in the registered protocol and was not conducted, because the heterogeneity of control conditions across the 20 trials was not coded as a categorical moderator variable prior to data extraction; post-hoc subgroup construction would risk inflating type I error. Descriptively, control conditions included traditional lectures (n = 9 studies), internet-based self-study (n = 3), case-based learning without AI (n = 3), bedside teaching (n = 2), standardised patient (n = 1), expert feedback (n = 1), and mixed formats (n = 1). Future reviews should pre-specify comparator type as a moderator to address whether GenAI confers additional benefit over active digital controls or only over passive instruction. Ethics approval and informed consent were not extracted as a pre-specified data item in the extraction protocol; consequently, a study-level table of ethics declarations cannot be presented here. Of the 20 included studies, those that explicitly mentioned ethics approval or institutional review board exemption in their published reports include Kalam et al. [18] (IRB exempt, Study ID STUDY00007141), Nissen et al. [5] (Goettingen Medical School, application 20/3/24), Tong et al. [10] (Ethics approval 2024–015), Wu et al. [15] (Guizhou Medical University affiliated hospital ethics review), and Luo et al. [14] (prospectively registered NCT06229379). Thirteen studies provided no explicit ethics statement in the published article. Readers requiring ethics details should consult the primary publications [17, 18]. The rapid evolution of GenAI means some tools tested in these trials are arguably already outdated, and the inability to fully blind participants leaves open the possibility that part of the measured benefit reflects enthusiasm rather than learning [5, 6]. A further limitation concerns statistical power. None of the 20 included trials reported a formal a priori sample size calculation, and the median sample size of 60 participants per study is insufficient to reliably detect effects smaller than approximately SMD = 0.5 with 80% power in a two-arm parallel-group design. Meta-analyses that aggregate underpowered primary studies are susceptible to the “winner’s curse”: because underpowered studies reach statistical significance only when they happen to observe larger-than-true effects, the studies that enter the meta-analysis are a biased sample and the pooled estimate may overstate the true benefit even after trim-and-fill adjustment. The modest attenuation observed in the some-concerns-only sensitivity analysis (knowledge SMD = 0.85; skills SMD = 1.05) is consistent with this possibility. Future trials should be pre-registered with a protocol-specified sample size calculation adequate to detect a minimally important clinical difference.
This review has methodological strengths that add to its credibility: six databases (English and Chinese), RCT-only inclusion, RoB 2 [33] and GRADE [32] applied, prediction intervals reported, and a broad set of sensitivity analyses (leave-one-out, some-concerns-only, trim-and-fill, crossover-correlation). Subgroup analyses and meta-regression were used to explore heterogeneity. For future work, the geographic imbalance is the most obvious gap: 84% of trials came from one country, so studies in other educational and cultural settings are needed [17, 18]. Trial design should also improve: protocols should be pre-registered, sample sizes should be calculated from expected effect sizes [34], outcome assessors should be blinded, and primary outcomes should be locked in before data collection [8, 33]. Follow-up data are thin (two to four studies per outcome), so longer tracking is a priority. Head-to-head comparisons of different GenAI modes (tutor vs. virtual patient vs. feedback tool) would be more useful than the AI-versus-no-AI designs that have dominated the literature [5, 9]. Cost-effectiveness data and faculty-workload assessments are nearly absent. Given how fast GenAI technology moves, future trials will need flexible designs that can accommodate new model versions without losing rigor [1, 2].
Conclusion
Across 20 RCTs (n = 1,413 participants), GenAI-integrated teaching was associated with moderate-to-large pooled SMDs for post-intervention clinical skills (SMD = 1.09, 95% CI: 0.81–1.38; GRADE: moderate certainty) and knowledge acquisition (SMD = 0.99, 95% CI: 0.49–1.49; GRADE: low certainty). The prediction interval for knowledge crossed zero (− 0.57 to 2.56), indicating that positive effects cannot be assumed in all future contexts. Knowledge gains at follow-up were observed in four studies (SMD = 0.51; moderate certainty), but this finding rests on a small evidence base. Evidence certainty is constrained by high heterogeneity, pervasive risk of bias, the structural impossibility of blinding, and geographic concentration (80% of studies from China). These findings are best interpreted as preliminary signals rather than definitive evidence of efficacy. Large-scale, pre-registered, multi-centre RCTs with blinded outcome assessment, active comparators, and sustained follow-up are essential before GenAI can be recommended as a routine component of medical curricula.
Practical implementation requires deliberate curricular integration, adequate infrastructure support, AI literacy training for both educators and learners, and realistic acknowledgement of the technology’s limitations, including variable output accuracy and susceptibility to hallucination. With 16 of 20 studies from a single country, no study rated as low risk of bias overall, and substantial unexplained heterogeneity in knowledge outcomes, these findings warrant cautious interpretation. Future research should prioritise well-designed, pre-registered, multi-centre RCTs with adequate sample sizes, blinded outcome assessment, extended follow-up, and geographically diverse populations. Governance and evaluation frameworks for GenAI in medical education will need to evolve in step with the technology to ensure that educational gains are genuine, equitable, and durable.
Supplementary Information
Acknowledgements
Not applicable.
Authors’ contributions
CW,XjP and NnS designed the project; CW, WyW and YmJ performed the literature search and data acquisition; NnS and XjP performed data extraction; CW, NnS and XjP performed the statistical analyses for heterogeneity investigation; CW, NnS and ZjM supported the writing of the paper. All authors read and approved the final manuscript.
Funding
This study was funded by 2024 Research and Practice Project on Higher Education Teaching Reform in Henan Province(No.2024SJGLX0323);Henan Province Medical Education Research Project in 2024(No.WJLX2023111);Research and Practice Project on Higher Education Teaching Reform at Henan University of Science and Technology in 2024(No.2024BK133).
Data availability
All data generated or analyzed during this study are included in this article and its supplementary files. The complete dataset is available from the corresponding author upon reasonable request.
Declarations
Ethics approval and consent to participate
Not applicable.
Consent for publication
Not applicable.
Competing interests
The authors declare no competing interests.
Footnotes
Publisher’s Note
Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
References
- 1.Brown TB, Mann B, Ryder N, Subbiah M, Kaplan JD, Dhariwal P, et al. Language Models are Few-Shot Learners. In: Larochelle H, Ranzato M, Hadsell R, Balcan MF, Lin H, editors., et al., Advances in Neural Information Processing Systems, vol. 33. Red Hook (NY): Curran Associates, Inc.; 2020. p. 1877–901. [Google Scholar]
- 2.Qadir J. Engineering Education in the Era of ChatGPT: Promise and Pitfalls of Generative AI for Education. In: 2023 IEEE Global Engineering Education Conference (EDUCON). Piscataway (NJ): IEEE; 2023. pp. 1-9. 10.1109/EDUCON54358.2023.10125121.
- 3.Lee J, Wu AS, Li D, Kulasegaram KM, Diplock LJ, Alikhan N, et al. Artificial intelligence in undergraduate medical education: a scoping review. Acad Med. 2021;96(11 Suppl):S62–70. 10.1097/ACM.0000000000004291. [DOI] [PubMed] [Google Scholar]
- 4.Masters K. Artificial intelligence in medical education. Med Teach. 2019;41(9):976–80. 10.1080/0142159X.2019.1595557. [DOI] [PubMed] [Google Scholar]
- 5.Nissen L, Rother JF, Heinemann M, Röpke S, Mothes T, Dietze A, et al. A randomised cross-over trial assessing the impact of AI-generated individual feedback on written online assignments for medical students. Med Teach. 2025;47(9):1544–50. 10.1080/0142159X.2025.2451870. [DOI] [PubMed] [Google Scholar]
- 6.Hirosawa T, Yokose M, Sakamoto T, Tanaka Y, Uchida T, Tomiki Y, et al. Utility of generative artificial intelligence for Japanese medical interview training: randomized crossover pilot study. JMIR Med Educ. 2025;11:e77332. 10.2196/77332. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 7.Page MJ, McKenzie JE, Bossuyt PM, Boutron I, Hoffmann TC, Mulrow CD, et al. The PRISMA 2020 statement: an updated guideline for reporting systematic reviews. BMJ. 2021;372:n71. 10.1136/bmj.n71. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 8.Higgins JPT, Thomas J, Chandler J, Cumpston M, Li T, Page MJ, et al., editors. Cochrane Handbook for Systematic Reviews of Interventions. Version 6.5. Cochrane; 2024. https://training.cochrane.org/handbook.
- 9.Wang D, Li D, Luo Z, Fan J, Yang L. Application of generative artificial intelligence in orthopedic clinical education [in Chinese]. Orthopaedics. 2025a;16(5):431–5. 10.3969/j.issn.1674-8573.2025.05.008. [Google Scholar]
- 10.Tong X, Hu Y, Long Y, Zhang Y, Wang Y, Li Z, et al. The application of problem-based learning guided by ChatGPT in clinical education in the Department of Nephrology. BMC Med Educ. 2025;25:1048. 10.1186/s12909-025-07427-w. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 11.Wu X, Huang Y, He Q. A large language model improves clinicians’ diagnostic performance in complex critical illness cases. Crit Care. 2025a;29(1):230. 10.1186/s13054-025-05468-7. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 12.Liu S, Xu Z, Li Y, et al. Application of an AI-assisted intelligent teaching model for targeted lung cancer therapy in clinical internship teaching for medical students [in Chinese]. Northwest Pharmaceut J. 2025;40(6):265–71. [Google Scholar]
- 13.Wang S, Zuo Y, Zou B, Zhao Y, Li L, Zhang X, et al. Enhancing self-directed learning with custom GPT AI facilitation among medical students: a randomized controlled trial. Med Teach. 2025b;47(7):1126–33. 10.1080/0142159X.2024.2413023. [DOI] [PubMed] [Google Scholar]
- 14.Luo MJ, Bi S, Pang J, Xie M, Yang Y, Zhang X, et al. A large language model digital patient system enhances ophthalmology history taking skills. NPJ Digit Med. 2025;8(1):502. 10.1038/s41746-025-01841-6. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 15.Wu C, Chen L, Han M, Li Z, Yang N, Yu C. Application of ChatGPT-based blended medical teaching in clinical education of hepatobiliary surgery. Med Teach. 2025b;47(3):445–9. 10.1080/0142159X.2024.2339412. [DOI] [PubMed] [Google Scholar]
- 16.Gan W, Ouyang J, Li H, Wang Y, Zhang X, Liu Y, et al. Integrating ChatGPT in Orthopedic Education for Medical Undergraduates: Randomized Controlled Trial. J Med Internet Res. 2024;26:e57037. 10.2196/57037. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 17.Digiacomo A, Orsini A, Cicchetti R, Bassi PF, Bove P, Lucarelli G, et al. ChatGPT vs traditional pedagogy: a comparative study in urological learning. World J Urol. 2025;43(1):286. 10.1007/s00345-025-05654-w. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 18.Kalam KA, Masoud FD, Muntaser A, et al. ChatGPT as a learning tool for medical students: results from a randomized controlled trial. Cureus. 2025;17(6):e85767. 10.7759/cureus.85767. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 19.Ba H, Zhang L, Yi Z. Enhancing clinical skills in pediatric trainees: a comparative study of ChatGPT-assisted and traditional teaching methods. BMC Med Educ. 2024;24(1):558. 10.1186/s12909-024-05565-1. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 20.Chen Y. Evaluation of the impact of AI-driven personalized learning platform on medical students’ learning performance. Front Med. 2025;12:1610012. 10.3389/fmed.2025.1610012. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 21.Hui Z, Zhu Z, Hu J, Cui Y. Application of ChatGPT-assisted problem-based learning teaching method in clinical medical education. BMC Med Educ. 2025;25:50. 10.1186/s12909-024-06321-1. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 22.Liang Z, Gao W. Construction and practice of a large language model-based simulation teaching resource system for abnormal labor in obstetrics. China Mod Doctor. 2025;63(30):70–1 in Chinese. [Google Scholar]
- 23.Tan S, Deng Q, Wei Q, Zhu X, Li S. The Application of Flipped Classroom Integrated with ChatGPT in Improving Graduate Education on Choroidal Melanoma. J Cancer Educ. 2025. 10.1007/s13187-025-02801-0. Epub ahead of print. [DOI] [PubMed]
- 24.Wang Z, Ji Y, Guo F. Application of virtual patients based on artificial intelligence models in surgical resident training. Jiangsu Health Syst Manag. 2025c;36(10):1500–3 in Chinese. [Google Scholar]
- 25.Zhang M, Guo C, Zhu C, Xu R, Zhou Y. Short-term efficacy analysis of applying DeepSeek R1 large model software in clinical teaching for medical undergraduates. J Bengbu Med Univ. 2025a;50(7):1008–12. 10.13898/j.cnki.issn.2097-5252.2025.07.032. in Chinese. [Google Scholar]
- 26.Zhang H, Lou S. Application of case-based learning combined with generative artificial intelligence in gynecologic tumor immunotherapy teaching. J Basic Clin Oncol. 2025b;38(3):428–30 in Chinese. [Google Scholar]
- 27.Wan X, Wang W, Liu J, Tong T. Estimating the sample mean and standard deviation from the sample size, median, range and/or interquartile range. BMC Med Res Methodol. 2014;14:135. 10.1186/1471-2288-14-135. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 28.Kirkpatrick DL. Evaluating Training Programs: The Four Levels. San Francisco (CA): Berrett-Koehler; 1994. [Google Scholar]
- 29.Viechtbauer W. Conducting meta-analyses in R with the metafor package. J Stat Softw. 2010;36(3):1–48. 10.18637/jss.v036.i03. [Google Scholar]
- 30.Egger M, Davey Smith G, Schneider M, Minder C. Bias in meta-analysis detected by a simple, graphical test. BMJ. 1997;315(7109):629–34. 10.1136/bmj.315.7109.629. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 31.Duval S, Tweedie R. Trim and fill: a simple funnel-plot-based method of testing and adjusting for publication bias in meta-analysis. Biometrics. 2000;56(2):455–63. 10.1111/j.0006-341X.2000.00455.x. [DOI] [PubMed] [Google Scholar]
- 32.Guyatt GH, Oxman AD, Vist GE, Kunz R, Falck-Ytter Y, Alonso-Coello P, et al. GRADE: an emerging consensus on rating quality of evidence and strength of recommendations. BMJ. 2008;336(7650):924–6. 10.1136/bmj.39489.470347.AD. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 33.Sterne JAC, Savovic J, Page MJ, Elbers RG, Blencowe NS, Boutron I, et al. RoB 2: a revised tool for assessing risk of bias in randomised trials. BMJ. 2019;366:l4898. 10.1136/bmj.l4898. [DOI] [PubMed] [Google Scholar]
- 34.Cohen J. Statistical Power Analysis for the Behavioral Sciences. 2nd ed. Hillsdale (NJ): Lawrence Erlbaum Associates; 1988. [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Data Availability Statement
All data generated or analyzed during this study are included in this article and its supplementary files. The complete dataset is available from the corresponding author upon reasonable request.





