Abstract
Purpose
Large language models (LLMs) have demonstrated remarkable potential in medical education, yet their performance on discipline-specific medical school course examinations remains incompletely characterized. This study evaluated six contemporary LLMs on final examinations for four core medical school courses: Surgery, Musculoskeletal System Diseases, Digestive System Diseases, and Respiratory System Diseases.
Methods
A total of 400 multiple-choice questions (100 per course) were administered to each model. Responses were scored against official answer keys and compared to the performance of 312 medical students. Each question was tested three times per model to assess response consistency and reproducibility.
Results
All six LLMs achieved mean scores exceeding 93% across all four courses, substantially outperforming the student cohort mean of 71.8%, with all models scoring above the 99.5th percentile of the student score distribution. ChatGPT 5.5 achieved the highest overall accuracy (97.0%), followed by Claude Opus 4.7 (96.3%) and DeepSeek V4 (95.8%). Qwen 3.7 (95.3%) and GLM 5.2 (94.3%) demonstrated comparable performance, while Doubao 2.1 (93.8%) showed the lowest but still exceptional accuracy. A sensitivity analysis revealed that accuracy on the 12 visual-containing questions was substantially lower (50.0%) than on text-only questions (96.8%). All models demonstrated high response consistency (κ = 0.82–0.95) and reproducibility (>95%).
Conclusions
Contemporary LLMs can achieve near-perfect performance on medical school course examinations, substantially exceeding average student performance. While this capability suggests significant potential for LLMs as supplementary educational tools, it also raises important concerns regarding assessment integrity and the appropriate role of AI in medical training.
Keywords: large language models, medical education, artificial intelligence, medical school examinations, assessment integrity
Introduction
The integration of artificial intelligence into medical education has accelerated dramatically over the past three years, with large language models emerging as transformative tools with the potential to reshape how medical knowledge is acquired, assessed, and applied.1–5 Since the introduction of GPT-3.5 and subsequent generations, LLMs have been systematically evaluated on high-stakes medical licensing examinations across multiple jurisdictions.6–9 On the United States Medical Licensing Examination (USMLE), GPT-4o has achieved an overall accuracy of 90.4% across 750 multiple-choice questions, significantly outperforming the medical student average of 59.3%.10,11 GPT-4 has demonstrated accuracy rates of 80–90% on the USMLE, with some studies reporting as high as 86% on Step 1 questions.12 On the Chinese National Medical Licensing Examination (NMLE), recent studies have revealed similarly impressive results.13,14 A systematic evaluation of GPT-4.0, ERNIE Bot 4.0, and GPT-4o on the 2023 Chinese Medical Licensing Examination found that both ERNIE Bot 4.0 and GPT-4o significantly outperformed GPT-4.0, achieving accuracies above the national pass mark.15 Longitudinal analysis has shown that DeepSeek-V3.2 achieved 94.7% accuracy on the 2025 Chinese NMLE, significantly surpassing ChatGPT-5.2 (89.3%).16 These findings underscore the rapid advancement of LLMs capabilities in medical knowledge assessment across both English and Chinese language contexts. Concurrently, global-scale informatics studies have systematically mapped the research landscape of ChatGPT in healthcare, highlighting its transformative potential while identifying critical challenges such as accuracy, safety, and ethical governance.2 Furthermore, bibliometric analyses have elucidated current concerns and future directions, emphasizing the urgent need for standardized evaluation.17
However, the majority of existing research has focused on licensing examinations rather than the discipline-specific course examinations that form the foundation of medical school curricula. Licensing examinations assess cumulative, end-of-training competency, whereas curriculum-embedded course examinations target granular organ-system knowledge directly aligned with ongoing pre-clinical and clinical teaching objectives. Evaluating LLMs on course-level assessments enables detection of fine-grained domain-specific strengths and weaknesses, provides direct evidence for AI-tutoring applicability, and exposes threats to the validity of routine in-school summative assessment in the generative-AI era.
A few recent studies have benchmark multiple contemporary LLMs across final examinations of medical school courses. A study comparing ChatGPT to medical students found that ChatGPT achieved a 90.2% correct response rate, outperforming 293 other students, with no significant difference in performance across surgical, internal, and fundamental medical sciences.18 Another investigation found that LLM mean course scores ranged from 7.46 to 9.88 (on a 10-point scale), vs student means of 4.28–7.32, with OpenAI o1 achieving the highest mean in three courses.19 System-based comparisons of AI chatbots on human anatomy demonstrated that GPT-4.1 achieved 95.7% accuracy, with all tested bots scoring 100% in respiratory and circulatory systems.20 A comparative analysis across four clinical courses found that LLMs consistently outperformed students, with performance gaps widening as cognitive complexity increased.19 Recent work by Khosravi et al comparing ChatGPT-5, Gemini 3, and medical students on neurology questions further confirmed this trend.21
Despite this growing body of evidence, few studies have systematically compared the performance of multiple contemporary LLMs across a comprehensive set of organ-system-based medical school course examinations. This study aims to address this gap by evaluating these six models on final examinations for four core medical school courses, comparing their performance to that of medical students, and analyzing performance patterns across courses and question types.
Materials and Methods
This was a cross-sectional comparative study evaluating the performance of six large language models on final examinations for four core medical school organ-system-based courses. The methodology consisted of three main input components feeding into a standardized testing protocol (Figure 1).
Figure 1.

Schematic illustration of the cross-sectional comparative study design.
Examination Courses
We collected final examination questions from four core medical school courses at a single medical institution, and uploaded them to the Mendeley Data repository.22 All questions were drawn from a closed, non-public institutional item bank, developed internally and used exclusively for on-campus summative assessments. The courses and their respective question counts were: Surgery (100), Musculoskeletal System Diseases (100), Digestive System Diseases (100), and Respiratory System Diseases (100). All questions were single-answer multiple-choice items with four to five options, scored dichotomously against official answer keys.
Questions were categorized by cognitive level according to Bloom’s taxonomy:23 recall (A1), comprehension (A2), application and analysis (A3/A4). Two experienced medical educators independently classified all 400 questions using a detailed rubric. Inter-rater reliability was excellent (Cohen’s κ = 0.87, 95% CI: 0.83–0.91), and disagreements were resolved through consensus with a third senior medical educator. Of the 400 questions, 12 (3%) originally contained visual elements such as radiological images, pathological slides, anatomical diagrams, or ECGs. A sensitivity analysis was subsequently performed to compare model performance on questions with and without these visual elements.
A1 Type (Single-sentence Multiple Choice): These questions contain no clinical case descriptions and directly assess fundamental medical theories, definitions, parameters and mechanisms, representing recall-oriented cognitive tasks. Each item includes one question stem and five options with a single correct answer.
A2 Type (Short Case-based Multiple Choice): Every question provides an independent brief clinical case of an individual patient. It evaluates elementary clinical reasoning, straightforward diagnosis and preliminary management, corresponding to the cognitive levels of comprehension and application.
A3 Type (Case Set Multiple Choice): One complete fixed clinical case is paired with 2 to 3 independent subquestions. All items draw on identical case information without updates of disease progression, focusing on comprehensive clinical application capacity.
A4 Type (Progressive Case Sequence Multiple Choice): This is the most challenging format. Follow-up subquestions sequentially add new manifestations, complications or lab findings to the initial case. All items are logically interrelated and demand sophisticated integrated clinical reasoning, classified as analysis-level cognitive tasks.
Large Language Models
Six contemporary LLMs were evaluated in this study: ChatGPT 5.5 (OpenAI, model version gpt-5.5-2026-01), Claude Opus 4.7 (Anthropic, model version claude-4-opus-2026-01), DeepSeek V4 (DeepSeek AI, model version Deepseek-v4-202512), Qwen 3.7 (Alibaba, model version qwen-max-2026-01), GLM 5.2 (Zhipu AI, model version glm-5.2-2026-01), and Doubao 2.1 (ByteDance, model version doubao-pro-2026-01). All models were accessed via their official web interfaces between June 30 and July 02, 2026. For every test session, built-in web-search and browsing functions of web interfaces were manually disabled to prevent real-time online retrieval. Questions were manually copied and pasted into the chat dialog using a standardized prompt, and questions were submitted three times to assess response consistency. All six models supported image upload through their web interfaces. Therefore, the 12 visual-containing questions were presented as original images.
Student Comparison Cohort
Student performance data were obtained from the same medical institution for the 2025–2026 academic year. A total of 312 medical students (second-year clinical medicine program) completed the same four course examinations. Student scores were de-identified, with only the average score per course used for comparative analysis.
Testing Protocol
Questions were administered to each model individually using a standardized prompt structure, consistent with prior LLM evaluation methodologies. The prompt was designed to simulate the examination format: “You are a medical student taking a [course name] final examination. Please select the single best answer from the options provided. Provide only the letter of your answer (A, B, C, D, or E)”. For each question, the full question stem and all options were provided in the language in which the question was originally written. For the 12 visual-containing questions, the original images were uploaded directly through each model’s web interface.
Each course was tested three times per model to assess response consistency and reproducibility, following established protocols. The three responses were generated in separate sessions with identical prompts. A response was considered consistent if the model provided the same answer across all three trials. The modal answer (the answer provided in at least two of three trials) was used as the final response for scoring purposes.
Statistical Analysis
Accuracy rates were calculated as percentages with 95% confidence intervals using the Wilson score interval. Between-model comparisons were performed using the chi-square test with Bonferroni correction for multiple comparisons. Model-student comparisons were performed by calculating z-scores and corresponding percentile ranks within the student score distribution, as each model produced a deterministic accuracy rate rather than a sample mean with sampling variability. Response consistency was assessed using Cohen’s κ coefficient, with values interpreted as: <0.20 = poor, 0.21–0.40 = fair, 0.41–0.60 = moderate, 0.61–0.80 = good, and 0.81–1.00 = excellent. Subgroup analyses were performed by course and cognitive level. All statistical tests were two-sided, with statistical significance defined as P < 0.05. Statistical analyses were conducted using Python version 3.13.1 and R version 4.3.2 (R Foundation for Statistical Computing, Vienna, Austria).
Results
Student Cohort Performance
The 312 medical students achieved a mean overall score of 71.8% (SD = 8.4) across the four courses. Course-specific mean scores were: Surgery 72.3% (SD = 9.1), Musculoskeletal System 69.7% (SD = 8.9), Digestive System 72.1% (SD = 8.2), and Respiratory System 73.1% (SD = 7.6). The student score distribution showed a normal pattern, with scores ranging from 48% to 91%.
Overall Model Performance
All six LLMs achieved mean scores exceeding 93% across all four courses, substantially outperforming the student cohort mean of 71.8% (Figure 2). ChatGPT 5.5 achieved the highest overall accuracy (97.0%), followed by Claude Opus 4.7 (96.3%) and DeepSeek V4 (95.8%). Qwen 3.7 (95.3%) and GLM 5.2 (94.3%) demonstrated comparable performance, while Doubao 2.1 (93.8%) showed the lowest but still exceptional accuracy. The mean accuracy of the six LLMs was 95.4%.
Figure 2.

Overall accuracy of each model on four courses vs students.
A sensitivity analysis comparing the full 400 questions (including 12 visual elements) against the 388 text-only questions revealed a substantial performance gap. Across all models, the mean accuracy on text-only questions was 96.8%, which was 1.4% points higher than the full 400 questions (95.4%). The mean accuracy on the 12 visual-containing questions was 50.0% (range across models: 33.3–66.7%), compared with 96.8% on text-only questions. Given the small number of visual items, this estimate has wide confidence intervals and should be interpreted cautiously.
We compared each model’s performance against the distribution of 312 medical students’ scores (mean:
= 71.8%, SD = 8.4%). Because each model produced a deterministic accuracy rate, we treated each model as a single individual in the student population and calculated its z-score and percentile rank within that distribution. This approach is statistically more appropriate than one-sample t-test, which would incorrectly treat a fixed constant as a random variable. The z-score is defined as:
![]() |
(1) |
Using this formula, we obtained the following results (Table 1).
Table 1.
Overall Accuracy of Six LLMs Across Four Medical School Courses
| Model | Mean Accuracy (%) | z-score | Approximate % of Students Outperformed |
|---|---|---|---|
| ChatGPT 5.5 | 97.0 | 3.02 | > 99.9% |
| Claude Opus 4.7 | 96.3 | 2.93 | > 99.8% |
| DeepSeek V4 | 95.8 | 2.86 | > 99.7% |
| Qwen 3.7 | 95.3 | 2.77 | > 99.7% |
| GLM 5.2 | 94.3 | 2.68 | > 99.6% |
| Doubao 2.1 | 93.8 | 2.62 | > 99.5% |
All six models achieved z-scores above 2.6, placing them well beyond the 99th percentile of the student score distribution. Even the lowest-performing model (Doubao 2.1, 93.8%) outperformed approximately 99.5% of the student cohort. The highest-performing model (ChatGPT 5.5, 97.0%) exceeded 99.9% of students. These percentile ranks provide a clear, intuitive interpretation of the performance gap: all tested LLMs scored higher than virtually all students in the reference cohort, confirming their exceptional knowledge retrieval and reasoning capabilities on these discipline-specific examinations without relying on questionable inferential statistics.
These results are consistent with recent findings on the Chinese NMLE, where DeepSeek-R1 achieved 96% accuracy and DeepSeek-V3.2 achieved 94.7% on the 2025 Chinese NMLE,13,24 as well as GPT-4o’s 90.4% accuracy on USMLE questions.25,26 Similarly, prior studies have shown that ChatGPT significantly outperforms medical students,7 and that GPT-4o and GPT-4 significantly outperformed the medical student average accuracy of 59.3% on USMLE questions.26–28
Course-Specific Performance
All models demonstrated the highest accuracy in Respiratory System questions, with ChatGPT 5.5 achieving 99.0% and all models exceeding 97.0% in this domain. The Musculoskeletal System course presented the greatest challenge across all models, with accuracies ranging from 90.0% (Doubao 2.1) to 95.0% (ChatGPT 5.5). Surgery and Digestive System courses showed intermediate performance levels, with all models exceeding 94.0% in Surgery and 94.0% in Digestive System (Table 2).
Table 2.
Course-Specific Accuracy of Six LLMs (%)
| Model | Surgery | Musculoskeletal | Digestive | Respiratory |
|---|---|---|---|---|
| DeepSeek V4 | 96.0 | 93.0 | 96.0 | 98.0 |
| Doubao 2.1 | 94.0 | 90.0 | 94.0 | 97.0 |
| ChatGPT 5.5 | 97.0 | 95.0 | 97.0 | 99.0 |
| Qwen 3.7 | 96.0 | 92.0 | 95.0 | 98.0 |
| Claude Opus 4.7 | 96.0 | 94.0 | 97.0 | 98.0 |
| GLM 5.2 | 95.0 | 91.0 | 94.0 | 97.0 |
The performance gap between models and students was most pronounced in Respiratory System (model mean 97.8% vs student mean 73.1%, difference 24.7% points) and least pronounced in Musculoskeletal System (model mean 92.5% vs student mean 69.7%), The performance gap ranged from 22.8 to 24.7% points across the four courses, with models consistently outperforming students in all course domains.
Cognitive-Level Performance
Questions were categorized by cognitive level according to Bloom’s taxonomy: recall (A1), comprehension (A2), application and analysis (A3/A4). All models demonstrated declining accuracy with increasing cognitive complexity. Performance was highest on recall questions (96.8–99.2%) and lowest on analysis questions (92.8–96.8%). ChatGPT 5.5 showed the strongest performance across all cognitive levels, particularly in analysis questions (96.8%). The performance decrement from recall to analysis was smallest for ChatGPT 5.5 (2.2% points) and largest for Doubao 2.1 (Table 3).
Table 3.
Cognitive-Level Accuracy of Six LLMs (%)
| Model | Recall (A1) | Comprehension (A2) | Application and Analysis (A3/A4) |
|---|---|---|---|
| DeepSeek V4 | 98.5 | 97.2 | 95.0 |
| Doubao 2.1 | 97.0 | 95.4 | 92.8 |
| ChatGPT 5.5 | 99.2 | 98.4 | 96.8 |
| Qwen 3.7 | 98.0 | 96.6 | 94.2 |
| Claude Opus 4.7 | 98.8 | 97.8 | 95.6 |
| GLM 5.2 | 96.8 | 95.0 | 93.0 |
Response Consistency and Reproducibility
ChatGPT 5.5 demonstrated the highest consistency (97.8%) and reproducibility (99.2%), with a κ coefficient of 0.946, indicating excellent inter-trial agreement. Claude Opus 4.7 (κ = 0.918) and DeepSeek V4 (κ = 0.892) also demonstrated excellent consistency. Doubao 2.1 (κ = 0.824) and GLM 5.2 (κ = 0.815) showed excellent agreement. The variability rate (percentage of questions with discordant responses across trials) was lowest for ChatGPT 5.5 (2.2%) and highest for GLM 5.2 (5.8%) (Table 4). These findings align with recent reports demonstrating high reproducibility of LLM responses on large-scale medical assessments.15,26
Table 4.
The Response Consistency and Reproducibility Metrics for Each Model
| Model | Consistency (%) | κ Coefficient | Reproducibility (%) |
|---|---|---|---|
| DeepSeek V4 | 96.2 | 0.892 | 98.1 |
| Doubao 2.1 | 94.5 | 0.824 | 96.8 |
| ChatGPT 5.5 | 97.8 | 0.946 | 99.2 |
| Qwen 3.7 | 95.8 | 0.863 | 97.5 |
| Claude Opus 4.7 | 96.9 | 0.918 | 98.6 |
| GLM 5.2 | 94.2 | 0.815 | 96.4 |
Discussion
This study provides the first comprehensive evaluation of six contemporary large language models—DeepSeek V4, Doubao 2.1, ChatGPT 5.5, Qwen 3.7, Claude Opus 4.7, and GLM 5.2—on four core medical school course examinations. Our findings demonstrate that all six models achieved mean scores exceeding 93% across all courses, substantially outperforming the student cohort mean of 71.8%. ChatGPT 5.5 achieved the highest overall performance (97.0%), followed by Claude Opus 4.7 (96.3%) and DeepSeek V4 (95.8%). These results have significant implications for medical education, assessment integrity, and the appropriate integration of AI into medical training.
Comparison with Prior Studies
Our findings are broadly consistent with and extend previous evaluations of LLMs on medical examinations. The 90.4% accuracy reported for GPT-4o on USMLE questions and the 80–90% accuracy range for GPT-4 are consistent with our finding that ChatGPT 5.5 achieved 97.0% on course-specific medical school examinations.7,12,15 The 96% accuracy of DeepSeek-R1 on the Chinese NMLE aligns with our finding that DeepSeek V4 achieved 95.8% overall accuracy, suggesting consistent performance across examination types.16,24
The performance of Chinese-developed models in our study merits particular attention. DeepSeek V4 (95.8%) significantly outperformed the previously reported performance of DeepSeek-V3 and DeepSeek-R1 on the Chinese NMLE,13 reflecting substantial advancement in Chinese LLM capabilities. Qwen 3.7 (95.3%) similarly exceeded the 92.5% accuracy previously reported for Qwen 3 on the Chinese NMLE.25 These findings suggest that Chinese LLMs have achieved parity with—and in some cases exceeded—international models in Chinese medical knowledge assessment, consistent with recent studies showing that Chinese corpus–based models perform similarly well in NMLE.16
Course-Specific Performance Patterns
The observation that all models performed best in Respiratory System questions (97.0–99.0%) and worst in Musculoskeletal System questions (90.0–95.0%) warrants further discussion. The near-perfect performance in Respiratory System may reflect several factors: (1) respiratory physiology and pathophysiology are relatively well-represented in training data, (2) respiratory disease presentations follow predictable patterns, and (3) the question types in this domain may rely more heavily on recall of established knowledge rather than complex clinical reasoning.
Conversely, the relatively lower performance in Musculoskeletal System may reflect the complexity of this domain, which requires integration of anatomy, biomechanics, rheumatology, and orthopedic surgery. Musculoskeletal questions often involve spatial reasoning, interpretation of imaging studies, and nuanced clinical decision-making that may challenge even advanced LLMs. This pattern is consistent with prior observations that LLMs show variable performance across medical subspecialties.
The performance decrement observed from recall to analysis questions (4.7–7.2% points) across all models suggests that while LLMs excel at knowledge retrieval, they face greater challenges with higher-order cognitive tasks requiring synthesis and evaluation. This finding is consistent with the broader literature on LLM capabilities and underscores the importance of designing assessments that test not just factual recall but clinical reasoning.
Response Consistency and Reliability
The high consistency (94.2–97.8%) and reproducibility (96.4–99.2%) observed across all models demonstrate that these LLMs produce reliable outputs under identical testing conditions. ChatGPT 5.5 showed the highest consistency (κ = 0.946), indicating excellent inter-trial agreement. This reliability is important for considering LLMs as potential educational tools, as it suggests that their performance is stable and predictable rather than stochastic.
Implications for Medical Education
These results should be interpreted with appropriate caution. The near-perfect examination scores primarily reflect exceptional medical knowledge retrieval and text-based clinical reasoning within the constrained format of multiple-choice questions. They should not be equated with real-world clinical competence, which requires multimodal reasoning, communication, and procedural skills.5
However, the finding that all LLMs scored above the 99.5th percentile of students raises a profound question regarding the purpose and validity of traditional multiple-choice examinations, but the implications of this finding are critically contingent upon the assessment setting.
In unproctored or remote settings, where AI is readily accessible, the finding indeed challenges the foundational premise that exam performance is a valid proxy for clinical knowledge, as AI renders such assessments invalid indicators of competency. However, in secure, supervised environments where AI access is restricted, traditional examinations retain their utility as a measure of baseline knowledge.
Therefore, this does not render all traditional assessments obsolete, but rather compels educators to move beyond simply restricting AI access and towards fundamentally redesigning assessment for scenarios where AI cannot be effectively blocked. The solution is not just to “secure” the exam, but to ensure the exam measures what is uniquely human about medical practice. Future assessments should:
Emphasize process over product: Evaluate the reasoning pathway (eg, through oral justifications, written clinical reasoning narratives) rather than just the final answer.
Focus on authentic, multimodal tasks: Use observed clinical encounters (OSCEs), simulation, and interpretation of real images/video where visual skills are paramount.
Integrate AI as a tool: Assess a student’s ability to collaborate with and critically evaluate AI outputs, preparing them for a future where AI is a clinical partner.
Implement longitudinal, portfolio-based assessment: Capture performance over time on a wider range of complex tasks that AI cannot easily replicate.
Moving forward, medical education must pivot towards assessments that emphasize human-centric and complex cognitive skills. This includes real-world clinical performance evaluations (eg, OSCEs), longitudinal workplace-based assessments, structured oral examinations (vivas), and evaluations of clinical reasoning processes rather than mere final answers. Furthermore, the distinction between unproctored and proctored environments must be carefully managed when integrating digital tools into medical curricula.
Limitations
This study has several limitations. First, a small subset (n=12, 3%) of examination questions originally contained visual elements in our study. This very small number substantially limits the statistical power and generalizability of our sensitivity analysis comparing visual vs non-visual questions. This demonstrates that although LLMs achieve high scores on text-heavy medical course examinations, their performance remains limited for tasks relying on genuine visual diagnostic reasoning. Multimodal LLM performance on larger sets of visual-rich medical assessment items warrants further dedicated investigation. Second, we used only single-answer multiple-choice questions, which represent a subset of medical school assessment formats. The performance of LLMs on open-ended questions, essay responses, and clinical simulation tasks requires further evaluation. Third, our testing protocol used a single standardized prompt for all models. Prior research has shown that prompt engineering can significantly affect LLM performance, and different prompting strategies might yield different results. Fourth, all models were tested via web interfaces rather than programmable APIs. This precluded precise control over inference parameters (eg, temperature) and precluded verification of hidden system prompts, which may have introduced unintended cross-model variability. Future investigations using API-based controlled inference are needed to confirm the generalizability of our findings.
Finally, we did not assess the potential for test-set contamination—the possibility that examination questions were included in model training data. While we used a closed, non-public, current course bank to minimize this risk, contamination cannot be definitively ruled out.
Conclusions
This comprehensive evaluation of six contemporary large language models on four core medical school course examinations demonstrates that all models achieved mean scores exceeding 93%, substantially outperforming the student cohort mean of 71.8%. ChatGPT 5.5 achieved the highest overall accuracy (97.0%), followed by Claude Opus 4.7 (96.3%) and DeepSeek V4 (95.8%). Notably, accuracy on questions containing visual elements was substantially lower (50.0%) than on text-only questions (96.8%), highlighting a critical limitation in visuospatial reasoning. All models demonstrated high response consistency and reproducibility, with ChatGPT 5.5 showing the most reliable performance. Performance was highest in Respiratory System questions and lowest in Musculoskeletal System questions, and all models showed declining accuracy with increasing cognitive complexity from recall to analysis.
These findings demonstrate that LLMs have achieved a ceiling effect on traditional text-based multiple-choice assessments in medical education. While this capability suggests significant potential as supplementary educational tools, it simultaneously poses a fundamental challenge to the validity of conventional assessment formats. Medical schools must urgently evolve their assessment strategies beyond simple MCQ formats toward more authentic, multimodal, and process-oriented evaluations that capture the unique cognitive and interpersonal skills of human clinicians. By thoughtfully redesigning assessments and thoughtfully integrating AI as a collaborative educational tool, we can harness the power of these technologies to enhance, rather than undermine, the development of competent, compassionate physicians.
Funding Statement
This research was funded by Science Foundation (Key Project) of Chongqing Medical and Pharmaceutical College, grant number ygzrc2023108.
Institutional Review Board Statement
The study was conducted in accordance with the Declaration of Helsinki, and approved by the Ethics Committee of Chongqing Medical and Pharmaceutical College (protocol code KYLLSC20260630024).
Abbreviations
LLM, Large Language Models; NMLE, National Medical Licensing Examination; USMLE, United States Medical Licensing Examination; CI, Confidence interval; κ, Cohen’s κ coefficient.
Data Sharing Statement
The data presented in this study are openly available in the Mendeley Data database at https://data.mendeley.com/datasets/cfwkcsxjc5/2.
Disclosure
The authors report no conflicts of interest in this work.
References
- 1.Nouri H, Mahdavi A, Abedi A, Mohammadnia A, Hamedan M, Amanzadeh M. Performance of large language models in medical licensing examinations: a systematic review and meta-analysis. J Educ Eval Health Prof. 2025;22:36. doi: 10.3352/jeehp.2025.22.36 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 2.Guo SB, Shen XM, Yao XW, Jiang L, Meng Y, Cai XY. Large language model ChatGPT in healthcare scenarios: a global-scale, cross-sectional, machine learning-based informatics study. Int J Surg. 2025;112:8961–10. doi: 10.1097/JS9.0000000000004413 [DOI] [PubMed] [Google Scholar]
- 3.Li R, Wu T. Delving into the practical applications and pitfalls of large language models in medical education: narrative Review. Adv Med Educ Pract. 2025;16:625–636. doi: 10.2147/AMEP.S497020 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 4.Abdelhafiz AS, Farghly MI, Sultan EA, Abouelmagd ME, Ashmawy Y, Elsebaie EH. Medical students and ChatGPT: analyzing attitudes, practices, and academic perceptions. BMC Med Educ. 2025;25(1):187. doi: 10.1186/s12909-025-06731-9 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 5.Alhur AA, Al-Kahtani NK, Mahgoub EA. Effects of Artificial intelligence-supported education on critical thinking, problem-solving, clinical reasoning, and decision-making in Health Science education: a systematic Review of interventional studies. Adv Med Educ Pract. 2026;17:1–21. doi: 10.2147/AMEP.S641965 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 6.Shaikh Y, Jeelani-Shaikh ZA, Jeelani MM, et al. Collaborative intelligence in AI: evaluating the performance of a council of AIs on the USMLE. PLoS Digit Health. 2025;4(10):e0000787. doi: 10.1371/journal.pdig.0000787 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 7.Casals-Farre O, Baskaran R, Singh A, et al. Assessing ChatGPT 4.0’s capabilities in the United Kingdom medical Licensing Examination (UKMLA): a robust categorical analysis. Sci Rep. 2025;15(1):13031. doi: 10.1038/s41598-025-97327-2 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 8.Madrid J, Diehl P, Selig M, et al. Performance of plug-in augmented ChatGPT and its ability to quantify uncertainty: simulation study on the German medical Board Examination. JMIR Med Educ. 2025;11:e58375. doi: 10.2196/58375 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 9.Kim HJ, Jung K, Shin S, et al. Performance evaluation of large language models on Korean medical licensing examination: a three-year comparative analysis. Sci Rep. 2025;15(1):36082. doi: 10.1038/s41598-025-20066-x [DOI] [PMC free article] [PubMed] [Google Scholar]
- 10.Gilson A, Safranek CW, Huang T, et al. How does ChatGPT perform on the United States medical Licensing Examination (USMLE)? The implications of large language models for medical education and knowledge assessment. JMIR Med Educ. 2023;9:e45312. doi: 10.2196/45312 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 11.Kung TH, Cheatham M, Medenilla A, et al. Performance of ChatGPT on USMLE: potential for AI-assisted medical education using large language models. PLoS Digit Health. 2023;2(2):e0000198. doi: 10.1371/journal.pdig.0000198 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 12.Bicknell BT, Butler D, Whalen S, et al. ChatGPT-4 omni performance in USMLE disciplines and clinical skills: comparative analysis. JMIR Med Educ. 2024;10:e63430. doi: 10.2196/63430 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 13.Wang W, Zhou Y, Fu J, Hu K. Evaluating the performance of DeepSeek-R1 and DeepSeek-V3 versus OpenAI models in the Chinese National medical Licensing Examination: cross-sectional comparative study. JMIR Med Educ. 2025;11:e73469. doi: 10.2196/73469 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 14.Zong H, Li J, Wu E, Wu R, Lu J, Shen B. Performance of ChatGPT on Chinese national medical licensing examinations: a five-year examination evaluation study for physicians, pharmacists and nurses. BMC Med Educ. 2024;24(1):143. doi: 10.1186/s12909-024-05125-7 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 15.Lian L, Luo X, Chipusu K, Ashraf MA, Wong KKL, Zhang W. Large language models evaluation of medical Licensing Examination using GPT-4.0, ERNIE Bot 4.0, and GPT-4o. Bioengineering. 2026;13(1):113. doi: 10.3390/bioengineering13010113 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 16.Wang Z, Qin Y, Wu J. Performance stability despite iteration: evaluating DeepSeek and ChatGPT on Chinese medical licensing examinations. Front Med. 2026;13:1874194. doi: 10.3389/fmed.2026.1874194 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 17.Guo SB, Liu DY, Fang XJ, et al. Current concerns and future directions of large language model ChatGPT in medicine: a machine-learning-driven global-scale bibliometric analysis. Int J Surg. 2026;112(2):2805–2822. doi: 10.1097/JS9.0000000000003668 [DOI] [PubMed] [Google Scholar]
- 18.Bharatha A, Ojeh N, Fazle Rabbi AM, et al. Comparing the performance of ChatGPT-4 and medical students on MCQs at varied levels of Bloom’s Taxonomy. Adv Med Educ Pract. 2024;15:393–400. doi: 10.2147/AMEP.S457408 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 19.Ros-Arlanzón P, Gutarra-Ávila R, Arrarte-Esteban V, et al. When AI models take the exam: large language models vs medical students on multiple-choice course exams. Med Educ Online. 2025;30(1):2592430. doi: 10.1080/10872981.2025.2592430 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 20.Nahir M, Kasap A, Sahin B. System-based comparison of the knowledge level of popular AI chatbots on human anatomy: a multiple-choice exam analysis of GPT-4.1, DeepSeek, Co-pilot, and Gemini models. Surg Radiol Anat. 2025;48(1):1. doi: 10.1007/s00276-025-03769-8 [DOI] [PubMed] [Google Scholar]
- 21.Khosravi M, Yousefi-Roobiyat M, Asghari Z, Nakhaei M, Khosravi F. Comparison of the performance of ChatGPT-5, Gemini 3, copilot, perplexity, and medical students in answering neurology questions: a cross-sectional study. Sci Rep. 2026;16(1):16070. doi: 10.1038/s41598-026-47666-5 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 22.Zheng W. Original data for: comprehensive evaluation of six large language models on four core medical school courses. 2026. Mendeley Data. doi: 10.17632/CFWKCSXJC5.2 [DOI]
- 23.Anderson LW, Krathwohl DR, Bloom BS. A Taxonomy for Learning, Teaching, and Assessing: A Revision of Bloom’s Taxonomy of Educational Objectives. New York: Longman; 2001. [Google Scholar]
- 24.Luo D, Liu M, Yu R, et al. Evaluating the performance of GPT-3.5, GPT-4, and GPT-4o in the Chinese National medical Licensing Examination. Sci Rep. 2025;15(1):14119. doi: 10.1038/s41598-025-98949-2 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 25.Kasagga A, Sapkota A, Changaramkumarath G, et al. Performance of ChatGPT and large language models on medical Licensing exams worldwide: a systematic Review and network meta-analysis with meta-regression. Cureus. 2025;17(10):e94300. doi: 10.7759/cureus.94300 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 26.El Natour D, Abou Alfa M, Chaaban A, Assi R, Dally T, Bou Dargham B. Performance of 5 AI models on United States medical Licensing Examination Step 1 questions: comparative Observational study. JMIR AI. 2026;5:e76928. doi: 10.2196/76928 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 27.Elkin PL, Mehta G, LeHouillier F, et al. Semantic clinical Artificial intelligence vs native large language model performance on the USMLE. JAMA Netw Open. 2025;8(4):e256359. doi: 10.1001/jamanetworkopen.2025.6359 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 28.Yang Z, Yao Z, Tasmin M, et al. Unveiling GPT-4V’s hidden challenges behind high accuracy on USMLE questions: observational study. J Med Internet Res. 2025;27:e65146. doi: 10.2196/65146 [DOI] [PMC free article] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Data Availability Statement
The data presented in this study are openly available in the Mendeley Data database at https://data.mendeley.com/datasets/cfwkcsxjc5/2.

