Skip to main content
Scientific Reports logoLink to Scientific Reports
. 2025 Sep 26;15:33083. doi: 10.1038/s41598-025-17030-0

Comparing the performance of ChatGPT, Gemini, and Claude in English and Polish on medical examinations

Dorota Wójcik 1,, Ola Adamiak 2, Gabriela Czerepak 2, Oskar Tokarczuk 2, Leszek Szalewski 3
PMCID: PMC12474969  PMID: 41006462

Abstract

In the realm of medical education, the utility of chatbots is being explored with growing interest. One pertinent area of investigation is the performance of these models on standardized medical examinations, which are crucial for certifying the knowledge and readiness of healthcare professionals. In Poland, dental and medical students have to pass crucial exams known as LDEK (Medical-Dental Final Examination) and LEK (Medical Final Examination) exams respectively. The primary objective of this study was to conduct a comparative analysis of chatbots: ChatGPT-4, Gemini and Claude to evaluate their accuracy in answering exam questions of the LDEK and the Medical-Dental Verification Examination (LDEW), using queries in both English and Polish. The analysis of Generalized Linear Mixed-Effects Model, which compared chatbots within question groups, showed that the chatbot Claude achieved the highest probability of accuracy for all question groups except the area of prosthetic dentistry compared to ChatGPT-4 and Gemini. In addition, the probability of a correct answer to questions in the field of integrated medicine was higher than in the field of dentistry for all chatbots in both prompt languages. Our results demonstrated that Claude achieved the highest accuracy in all areas analysed and outperformed other chatbots. This suggests that Claude has significant potential to support the medical education of dental students. This study showed that the performance of chatbots varied depending on the prompt language and the specific field. This highlights the importance of considering language and specialty when selecting a chatbot for educational purposes.

Supplementary Information

The online version contains supplementary material available at 10.1038/s41598-025-17030-0.

Keywords: Medical education, Dentistry, Large language models (LLMs), Artificial intelligence

Subject terms: Medical research, Dental education

Introduction

Artificial intelligence (AI) is a transformative technology with the potential to revolutionize various aspects of our lives by enhancing efficiency, enabling new capabilities, and providing personalized experiences. Various applications are available, such as LLM, generative AI. Generative AI refers to a subset of artificial intelligence focused on creating new content or data that is similar to existing data. It uses machine learning models to generate text, images, music, and other media. A Large Language Model is a type of AI model specifically designed to understand, generate, and manipulate human language on a large scale13.

In the realm of medical education, the utility of chatbots is being explored with growing interest.

In medical education, chatbots powered by artificial intelligence (AI) are employed to enhance the teaching and learning process by generating instructional tasks, evaluating prior knowledge, fostering student engagement in educational activities, and supporting the development of higher-order cognitive skills4. They are also used to generate examination questions and provide personalized feedback5. Chatbots are increasingly utilized in medical education and practice for tasks such as responding to queries from medical and dental students, supporting clinicians in diagnostic decision-making, and providing patients with information about their conditions. Their high effectiveness in addressing both undergraduate and specialist medical examination questions, such as those from the United States Medical Licensing Examination (USMLE), has been consistently demonstrated in the literature. Recently, numerous studies have evaluated ChatGPT’s overall performance in passing medical examinations, highlighting its potential in educational settings3,69. Findings from these studies indicate that chatbots can enhance student learning; however, their role should remain complementary to traditional teaching methods, such as lectures, textbooks, and consultations with instructors10. Despite growing interest in deploying AI chatbots in medical education, their implementation remains limited, as these tools may generate inaccurate or inconsistent responses due to challenges in understanding complex medical contexts, with their performance varying depending on the medical specialty, type of examination questions, and language11,12. To date, empirical findings on the performance of AI chatbots in medical education have been mixed, preventing definitive conclusions. This study addresses this gap by evaluating the performance of ChatGPT-4, Claude, and Gemini within the Polish context of the Medical-Dental Final Examination (LDEK) and the Medical-Dental Verification Examination (LDEW). One pertinent area of investigation is the performance of these models on standardized medical examinations, which are crucial for certifying the knowledge and readiness of healthcare professionals.

In Poland, dental and medical students have to pass crucial exams known as LDEK (Medical-Dental Final Examination) and LEK (Medical Final Examination) exams respectively. Evaluating the performance of AI chatbots on such examinations could provide valuable insights into how artificial intelligence might support medical education. The LEK, which is comparable to the United States Medical Licensing Examination (USMLE), is particularly significant. After passing this exam, candidates are eligible to apply for a medical licence in Poland. In addition, this qualification is recognised in all EU Member States due to European Union regulations outlined in Directive 2005/36/EC. This qualification is recognized across EU member states, allowing for potential practice throughout the Union13.

This paper aimed to assess the performance of GPT-4 Legacy (ChatGPT4), Claude and Gemini (chatbots) on the LDEK, examining their accuracy, strengths, and limitations in this context. Through this investigation, we sought to understand the potential role of chatbots in medical education and their implications for the future of healthcare training and practice.

The primary objective of this study was to conduct a comparative analysis of chatbots: ChatGPT-4, Gemini and Claude to evaluate their accuracy in answering exam questions of the LDEK and the Medical-Dental Verification Examination (LDEW), using queries in both English and Polish.

Several secondary objectives were also formulated. The first secondary objective was to determine the relationship between the self-assessed confidence of the chatbots and the percentage of correct answers in the population as reported by the Medical Examination Centre (CEM). The second secondary objective was to determine whether the self-assessed confidence was related to the actual accuracy of the chatbots. The third secondary objective was to investigate and evaluate the agreement of the responses generated by each chatbot between prompts in English and Polish languages. The fourth subsidiary objective involved evaluating the agreement of responses across three measurements within each prompt language and among the chatbots. The fifth subsidiary objective was to ascertain if significant differences in accuracy existed among the chatbots within identical prompt languages. The sixth subsidiary objective was to rigorously assess the existence of disparities in accuracy between the prompt languages employed by the same chatbot. The final, seventh subsidiary objective was to conduct a detailed evaluation of the variations in accuracy among chatbots across different categories of questions (specifically, those pertaining to the seven distinct disciplines within the dental faculty: conservative dentistry, paediatric dentistry, dental surgery, prosthetics, periodontology, orthodontics, and integrated medicine) and between various prompt languages.

Materials and methods

Study design

Study design is presented in the flowchart (Fig. 1).

Fig. 1.

Fig. 1

Flowchart illustrating the study design.

Questions

The study was conducted from 27 March to 13 April. Questions from the LDEK and LDEW exams conducted in February 2024 were used for the study. These exams, provided by the Central Examination Board (CEM), comprised 198 questions each. The LDEK exam was held in Polish, while the LDEW exam was in English. Both exams contained identical questions, differing only in the language of presentation. Each question was preceded by a pre-standardized prompt in Polish or English, depending on the exam version.

Question group

The study examined responses to questions from seven areas of dentistry and medicine (see Table 1 for a detailed description of the variables). For the purpose of this study, the question groups emergency medicine, bioethics and medical law, medical certification and public health were combined into one category – integrated medicine. Each question was a multiple-choice question with five possible answers from A to E.

Table 1.

Mean accuracy values of the chatbots between Language categories and question specialties (N = 1188).

Question group Language Mean accuracy
ChatGPT4
M (SD)
Gemini
M(SD)
Claude
M (SD)

Conservative dentistry

n = 46 (23%)

English 57 (3.5) 54 (3.1) 72 (3.4)
Polish 61 (3.5) 44 (3.8) 72 (3.6)

Paediatric dentistry

n = 29 (15%)

English 49 (4.4) 55 (4.4) 68 (4.6)
Polish 63 (4.8) 53 (4.9) 71 (4.8)

Dental surgery

n = 25 (13%)

English 56 (4.5) 51 (4.5) 69 (4.9)
Polish 40 (4.9) 23 (3.9) 55 (5.1)

Prosthetic dentistry

n = 25 (13%)

English 44 (4.9) 51 (4.4) 45 (5.2)
Polish 52 (4.9) 51 (4.3) 39 (5.1)

Periodontology

n = 19 (10%)

English 49 (5.8) 54 (5.3) 60 (5.9)
Polish 54 (6.3) 30 (4.5) 54 (6.2)

Orthodontics

n = 19 (10%)

English 42 (5.6) 56 (5.2) 54 (5.6)
Polish 61 (5.5) 44 (5.2) 56 (5.4)

Integrated medicine

n = 35 (18%)

English 90 (2.4) 75 (3.2) 95 (1.9)
Polish 82 (2.7) 81 (3.2) 92 (2.4)

Languages and measurements

For the data collection phase, the official websites of the chatbot providers were utilized: responses from ChatGPT were sourced from https://chatgpt.com, those from Gemini were obtained from https://deepmind.google/technologies/gemini/, and responses from Claude were accessed via https://console.anthropic.com. A single prompt was used to elicit responses to one question, employing a total of 198 prompts per examination. To ensure the internal consistency of the responses, the examination was repeated three times with prompts in Polish and three times with prompts in English, with intervals of a few hours between each session. In total, information from 1,188 prompts per chatbot was collected.

Answer selection and self-assessed confidence

The chatbot was asked to select a correct answer and provide a self-assessed confidence, which represents the chatbot’s own assessment of the probability that its answer to the test question was correct. The self-assessed confidence was to be expressed on a scale of 0–100%, with 0% representing complete uncertainty about the accuracy of the answer and 100% representing absolute confidence in the accuracy of the answer. The chatbots’ responses were evaluated against an official answer key provided by the Polish Centre for Medical Examinations.

Statistical analysis was employed to analyse the data and evaluate the answers through statistical methods. Risk assessment enabled the chatbots to evaluate the risk of an incorrect answer and its potential consequences.

Phase of answer verification

A trained human rater verified the correctness of the response provided by the chatbot to the examination question, which involved selecting the letter corresponding to the answer option, by comparing it with the official answer key provided by the Centre for Evaluation and Measurement (CEM). If the response matched the key, it was considered correct; otherwise, it was deemed incorrect.

Hypotheses

The investigation was segmented into several key hypotheses, each aiming to dissect distinct aspects of chatbot performance. This included examining the correlation between self-assessed and actual accuracy, the consistency of responses across different languages, and the comparative performance of different chatbot models. Additionally, the research evaluated the performance of chatbots on specific domains of knowledge and across different types of queries, thereby providing a comprehensive overview of their capabilities and limitations. The full list of hypotheses was included in Appendix.

Statistical analysis

The significance level of the statistical tests was set at α = 0.05. The distribution of variables was described using descriptive statistics. Measures for the central tendency, in particular the mean (M) and the variance in the form of the standard deviation (SD), were used. Categorical variables were described using frequencies across individual categories (n) and percentages.

Correlation analysis

The association between a dichotomous variable and a numeric variable was assessed using the rank biserial correlation coefficient (r̂ biserial rank). Spearman’s rank correlation coefficient (rho) was used to measure the strength and direction of the association between two numerical variables, especially for non-normally distributed variables. The rank biserial coefficient was interpreted on the basis of Funder convention (Funder, 2019)14. The rho coefficient was interpreted on the basis of the Cohen convention (Cohen, 1988)15.

Analysis of the agreement

For each of the three chatbots under investigation, the linguistic consistency between their Polish and English versions of answer was evaluated using Cohen’s Kappa coefficient. The results of Cohen’s Kappa were interpreted in accordance with McHugh’s guidelines (McHugh, 2012)16.

The agreement between the three responses of the same chatbot to the same exam question, evaluated separately in both Polish and English, was calculated using Fleiss’ Kappa coefficient. Each response was treated as an independent rater for the purpose of this evaluation. The results were interpreted on the basis of Fleiss’ convention (Fleiss, 2003)17.

Multivariate analysis

Generalised linear mixed model (GLMM)

Generalised linear mixed models (GLMM) were used to assess the probability of correct responses and to enable comparisons between specific chatbots as well as to estimate effect sizes.

The choice of this methodology was motivated by several factors. A multivariate approach was required due to the complexity of the data, which included multiple variables affecting the accuracy of chatbot responses. A generalised model was chosen due to the dichotomous nature of the outcome variables (accuracy), which violates the assumptions of classical linear regression. The implementation of a mixed model was essential to account for repeated measurements, as each chatbot provided three responses to the same exam question.

Models

Two GLMM models were implemented:

Model 1 aimed to predict and determine the effects of chatbot type and query language on response accuracy.

Model 2 was developed to determine the influence of question group, chatbot type and language on the probability of a correct answer.

In both models, the dichotomous outcome variable was fitted using a binomial distribution with a logit link function.

Detailed specifications of Model 1 and Model 2 can be found in Appendix.

Specification of fixed and random effects in GLMER

Two Generalized Linear Mixed-Effects Models (GLMER) were implemented to account for dependencies between observations and to analyze data with repeated measurements.

Model 1: accuracy_chat ~ chatbot * language + (1|ID_Q).

This model examined the interaction between chatbot type and prompt language on response accuracy. Fixed effects included chatbot type, language, and their interaction. The random effect was the random intercept for questions (ID_Q).

Model 2: accuracy_chat ~ chatbot * Q_group * language + (language|ID_Q).

This model investigated the interaction between question group, chatbot type, and prompt language on response accuracy. Fixed effects included chatbot type, question group, language, and their interactions. Random effects included random intercepts for questions and random slopes for language by question.

Both models used a binomial distribution with a logit link function. The models were fitted using maximum likelihood estimation, specifically the Laplace approximation.

Model 2 target effect: The three-way interaction between question group, chatbot type (ChatGPT-4, Gemini, Claude), and prompt language (English, Polish), with language as a moderator.

Goodness-of-fit for both models was assessed using conditional and marginal coefficients of determination (Marginal R2 for fixed effects, Conditional R2 for both fixed and random effects).

For Model 2, optimisation was performed using the Bound Optimisation BY Quadratic Approximation (BOBYQA) algorithm, developed by M.J.D. Powell (2009). This algorithm was chosen to handle the complexity and size of the data involved. Model 1 did not require additional optimisation procedures beyond the standard fitting process18.

Marginal means and contrast analysis

To test hypotheses about specific differences between groups, an estimation of marginal means (EMM) and a contrast analysis were performed. The Holm correction was applied to adjust significance levels when comparing three or more groups, controlling the family-wise error rate in cases of multiple pairwise comparisons.

In the EMM analysis, values of the outcome variable in subgroups were presented as probabilities of obtaining a correct response. For the contrast analysis, these values were transformed into odds ratios (OR). Effect sizes were interpreted based on Cohen’s (1988) convention. P-values were calculated as approximations to Wald Z distribution statistics15.

The target effect was calculated in two ways: first, analyzing differences between chatbots within language categories, and second, examining differences between languages for each chatbot. Detailed model specifications can be found in the Appendix.

Statistical environment

Analyses were conducted using the R Statistical language (version 4.4.0; R Core Team, 2024) on Windows 11 pro 64 bit, using the packages: readxl19 and dplyr20 effectsize21 lme422, sjPlot23 irr24 tidyr25 ggstatsplot26 ggplot227.

Characteristics of the study sample

The existing dataset comprised 3564 observations and included the following variables: a unique identifier for each exam question, the category of the question, the language of the prompt, which was either Polish or English, the model of the chatbot used, including ChatGPT-4, Gemini and Claude, the identifier of the rater, labelled as A, B or C, the repetition number, which ranged from 1 to 3, the accuracy of the chatbot response, coded as 1 for correct and 0 for incorrect, self-assessed confidence and the percentage of correct answers as reported by the Medical Examination Centre.

Missing data were observed in one variable: self-assessed confidence (n = 1).

Results

Descriptive statistics

Claude showed the highest mean accuracy in most specialties (Table 1), both in English and Polish. The best results were achieved in the specialty of integrated medicine, where accuracy reached 95% for English and 92% for Polish. ChatGPT-4 and Gemini showed more varied results. (Table 2).

Table 2.

Mean self-assessed confidence values of the chatbots for different language categories and question specialties (N = 1188).

Question group Language Self-assessed confidence
ChatGPT4
M (SD)
Gemini
M(SD)
Claude
M (SD)

Conservative dentistry

n = 46 (23%)

English 62 (2.3) 66 (2.1) 78 (1.9)
Polish 69 (2.4) 70 (2.3) 71 (1.6)

Paediatric dentistry

n = 29 (15%)

English 71 (2.4) 66 (2.3) 84 (2.1)
Polish 76 (2.5) 89 (1.6) 71 (1.6)

Dental surgery

n = 25 (13%)

English 64 (2.4) 65 (2.8) 76 (3.0)
Polish 71 (3.2) 69 (3.3) 81 (2.5)

Prosthetic dentistry

n = 25 (13%)

English 59 (2.6) 67 (3.0) 72 (2.8)
Polish 72 (3.4) 67 (3.1) 73 (1.9)

Periodontology

n = 19 (10%)

English 68 (3.7) 61 (3.9) 84 (3.0)
Polish 85 (3.7) 63 (4.2) 90 (1.8)

Orthodontics

n = 19 (10%)

English 60 (2.7) 78 (3.7) 71 (3.4)
Polish 70 (3.4) 62 (4.9) 72 (3.2)

Integrated medicine

n = 35 (18%)

English 78 (2.0) 75 (1.6) 90 (1.4)
Polish 85 (1.9) 72 (2.1) 84 (1.4)

Analysis of the percentage of correct answers by Language group

The accuracy metric was determined by calculating the arithmetic mean of the three measurements. Calculations were performed separately for each language group to account for possible differences in the quality of the answers given by the chatbots depending on the language of the prompt (Fig. 2).

Fig. 2.

Fig. 2

Accuracy by chatbot and language*. *The line at 56% indicates the minimum score for passing the LDEK and LDEW exams. Based on our results, considering the average of three measurements of chatbots answering exam questions, only Claude would pass the exam in each attempt. On the other hand, only Gemini, when taking the exam in Polish, would certainly fail.

Correlation analysis between self-assessed confidence and aggregate CEM system test scores for the population

The results of the Spearman correlation analysis (Table 3) showed a statistically significant, small, positive correlation between the aggregated CEM system exam results for the population and the self-assessed confidence for ChatGPT-4 and Gemini. The effect size for ChatGPT-4 and Gemini is small.

Table 3.

Spearman correlation results: self-assessed confidence vs. aggregated CEM test results (N = 1188).

Chatbot rho p
ChatGPT4 0.16 < 0.001
Gemini 0.10 < 0.001
Claude 0.04 0.192

There was a weak, positive correlation between the self-assessed confidence of ChatGPT-4 and Gemini and the effectiveness of the students’ answers to the exam questions. The more difficult the question was for the students, the lower the chatbot rated its probability of giving a correct answer. The small effect size suggested that the chatbots’ self-assessment may not be a completely reliable indicator of their actual effectiveness. Based on the results of these analyses, the first hypothesis was accepted.

Correlation between self-assessed confidence and actual accuracy when answering exam questions

There was a statistically significant correlation between self-assessed confidence and the actual accuracy of ChatGPT-4 and Claude in providing correct answers to exam questions (Table 4). The higher the self-assessed confidence, the higher the actual accuracy in answering the questions correctly. The rank biserial correlation coefficients for ChatGPT-4 and Claude indicated a medium and very small effect sizes, respectively.

Table 4.

Rank biserial correlation: self-assessed confidence vs. actual accuracy.

Chatbot
(N = 1188)
Inline graphic p
ChatGPT4 0.22 < 0.001
Gemini 0.02 0.590
Claude 0.09 < 0.001

There was a positive correlation between the self-assessed confidence of ChatGPT-4 and Claude and their actual effectiveness in answering exam questions. Based on the results of these analyses, the second hypothesis was accepted.

Analysis of the agreement of chatbot answers to exam questions based on the language of the prompt

In the study, the agreement of answers given by chatbots to exam questions in Polish and English was assessed using Cohen’s Kappa coefficient as a measure of reliability.

The Cohen’s Kappa coefficient (Table 5) shows that ChatGPT-4 and Claude indicated moderate agreement between the responses in English and Polish. Gemini showed weak agreement between the responses in English and Polish. All values determined are statistically significant (p < 0.001).

Table 5.

Analysing the agreement of answers to exam questions within the chatbots between the query languages.

Chatbot (N = 1188) Cohen’s κ z p
Chat GPT4 0.45 10.9 < 0.001
Gemini 0.39 9.47 < 0.001
Claude 0.59 14.3 < 0.001

Cohen’s κ cohen’s Kappa coefficient, z z-score.

p p value.

All chatbots analysed differed in the consistency of responses between Polish and English prompts. While there was some agreement between languages for all chatbots, the level of consistency differs. As a result of these analyses, the fourth hypothesis was accepted.

Analysis of the agreement of the three chatbot answers to the same questions using Fleiss’ kappa coefficient (Table 6) revealed differences depending on the chatbot model and the prompt language. The best results, indicating excellent agreement between consecutive measurements, were obtained for Claude chatbot in Polish. ChatGPT-4 also achieved good agreement between measurements, regardless of the language of the responses. Gemini showed the weakest response consistency in English.

Table 6.

Analysis of the agreement of answers to exam questions across three measurements, depending on chatbot and prompt language.

Chatbot
(N = 594)
Prompt language Fleiss’ κ z p
Chat GPT4 Polish 0.58 14.1 < 0 0.001
English 0.59 14.3 < 0 0.001
Gemini Gemini Polish 0.60 14.7 < 0 0.001
English 0.40 9.82 < 0 0.001
Claude Claude Polish 0.78 19.1 < 0 0.001
English 0.72 17.6 < 0 0.001

Fleiss’ κ Fleiss’ Kappa coefficient, z z-score, p p value

All analysed chatbots differ in the degree of response consistency across three measurements in both prompt languages. Chatbot responses to the same questions may vary depending on the subsequent measurement for both languages. The findings from these analyses led to the acceptance of the fourth hypothesis.

Model 1 fit results

Model fit statistics for Model 1 and 2, including marginal and conditional R2 values as well as random effects parameters, were obtained using the sjPlot package in R (via the tab_model() function) based on mixed-effects models fitted with the lme4 package.

graphic file with name d33e1331.gif
graphic file with name d33e1337.gif

EMM

In the range for English, the probability ranged from 0.63 to 0.80, in the range for Polish from 0.48 to 0.77 (Table 7).

Table 7.

EMM - Estimated marginal means for the analysed variables (N = 594).

Prompt language Chatbot Probability SE CI 95%
English ChatGPT4 0.64 0.05 0.55–0.73
Gemini 0.63 0.05 0.54–0.72
Claude 0.80 0.03 0.73–0.86
Polish ChatGPT4 0.65 0.05 0.56–0.74
Gemini 0.48 0.05 0.39–0.58
Claude 0.77 0.04 0.70–0.84

SE standard error, CI 95% confidence interval 95%

Analysis of contrasts between chatbots across languages

The analysis of contrasts between chatbots across languages (Table 8; Fig. 3) revealed significant differences in accuracy for both prompt languages (p < 0.001). The only exception was the comparison between ChatGPT-4 and Gemini in English, which showed no significant difference (p = 0.808).

Table 8.

Contrast analysis to investigate differences between chatbots across Language categories.

Prompt language Contrast OR SE CI 95% p adjust
English ChatGPT4 / Gemini 1.04 0.16 0.72–1.50 0.808
ChatGPT4 / Claude 0.44 0.07 0.30–0.64 < 0.001
Gemini / Claude 0.43 0.07 0.29–0.62 < 0.001
Polish ChatGPT4 / Gemini 2.02 0.31 1.40–2.91 < 0.001
ChatGPT4 / Claude 0.55 0.09 0.38–0.80 < 0.001
Gemini / Claude 0.27 0.04 0.19–0.40 < 0.001

OR odds ratio, SE standard error, p adjust p-value adjusted using the Holm correction

Fig. 3.

Fig. 3

Differences in the probability of accuracy between chatbots in different language categories.

Most differences showed a small effect size, with the exception of the comparison between Gemini and Claude in Polish, where a medium effect size was observed.

For English, ChatGPT-4 and Gemini showed similar effectiveness in predicting correct answers. The contrast analysis showed that both ChatGPT-4 and Gemini were less effective compared to Claude (p < 0.001). In the case of Polish, ChatGPT-4 achieved significantly better results than Gemini (p < 0.001); however, both chatbots were less effective than Claude (p < 0.001). The findings from these analyses led to the acceptance of the fifth hypothesis.

Analysis of contrasts between chatbots within the prompt Language

The Gemini model achieved significantly higher accuracy in English compared to Polish (Table 9; Fig. 4). The odds ratio was 1.85, indicating that the odds of Gemini providing a correct answer were 84.5% higher in English than in Polish. This difference was statistically significant, and corresponded to a small effect size.

Table 9.

Intergroup differences between the languages for individual chatbots.

chatbot Contrast OR SE CI 95% p
ChatGPT4 English / Polish 0.95 0.15 0.71–1.28 0.746
Gemini English / Polish 1.85 0.28 1.37–2.50 < 0.001
Claude English / Polish 1.20 0.19 0.87–1.62 0.278

OR odds ratio, SE standard error, p p value

Fig. 4.

Fig. 4

Differences in the probability of correct answers between languages within chatbots.

For the ChatGPT-4 and Claude models, no statistically significant differences were observed in the accuracy of providing correct answers between English and Polish. The findings from this analysis support the acceptance of the sixth hypothesis.

Model 2

The reference categories for the chatbots were ChatGPT-4 responding to questions about conservative dentistry in English. The marginal R2 was 0.208. The conditional R2 was 0.701. The random effects were as follows: σ2 = 3.29, τ₀₀ for ID_Q = 5.19, τ11 ID_Q.languagepl = 1.24, ρ01 ID_Q = -0.15, ICC = 0.62.

EMM

In the range for English, the probability ranged from 0.37 to 1.0, in the range for Polish from 0.13 to 0.99 (Table 10).

Table 10.

EMM—estimated marginal means for the analysed variables.

Question group Chatbot English Polish
Probability SE CI 95% Probability SE CI 95%
Conservative dentistry ChatGPT4 0.63 0.10 0.43–0.79 0.63 0.10 0.42–0.80
Gemini 0.58 0.10 0.38–0.75 0.39 0.10 0.21–0.60
Claude 0.83 0.06 0.68–0.92 0.85 0.06 0.71–0.93
Paediatric dentistry ChatGPT4 0.44 0.13 0.21–0.69 0.59 0.14 0.32–0.81
Gemini 0.53 0.13 0.28– 0.76 0.51 0.14 0.26–0.76
Claude 0.77 0.10 0.53–0.91 0.87 0.07 0.69–0.96
Dental surgery ChatGPT4 0.59 0.13 0.33–0.81 0.32 0.13 0.13–0.59
Gemini 0.53 0.14 0.28–0.77 0.13 0.07 0.04–0.32
Claude 0.79 0.09 0.55–0.92 0.60 0.14 0.33–0.82
Prosthetic dentistry ChatGPT4 0.41 0.14 0.19–0.68 0.56 0.14 0.29–0.80
Gemini 0.53 0.14 0.27–0.77 0.54 0.14 0.27–0.78
Claude 0.43 0.14 0.20–0.70 0.33 0.13 0.14–0.61
Periodontology ChatGPT4 0.50 0.16 0.21–0.78 0.58 0.16 0.27–0.84
Gemini 0.60 0.16 0.29–0.84 0.17 0.10 0.05–0.44
Claude 0.69 0.14 0.38–0.89 0.57 0.16 0.27–0.84
Orthodontics ChatGPT4 0.37 0.15 0.14–0.67 0.66 0.15 0.35–0.88
Gemini 0.61 0.15 0.31–0.85 0.39 0.16 0.15–0.70
Claude 0.58 0.16 0.28–0.83 0.61 0.16 0.30–0.85
Integrated medicine ChatGPT4 0.98 0.01 0.94–1.00 0.94 0.03 0.85–0.98
Gemini 0.88 0.06 0.72–0.95 0.94 0.03 0.83–0.98
Claude 1.00 0.01 0.98–1.00 0.99 0.01 0.96–1.0

Analysis of contrasts between chatbots across question specialties and languages

English language

No statistically significant differences were detected among the chatbots in the specialties of prosthetic dentistry, periodontology, and orthodontics regarding the probability of accuracy to exam questions in English (Table 11; Fig. 5).

Table 11.

Analysis of accuracy differences across question specialties and prompt language categories between chatbots.

Question group Contrast Prompt language
English Polish
OR SE CI 95% p adjust OR SE CI 95% p adjust
Conservative dentistry ChatGPT4 / Gemini 1.23 0.40 0.57–2.67 0.518 2.64 0.88 1.18–5.87 0.004
ChatGPT4 / Claude 0.35 0.12 0.15–0.79 0.004 0.29 0.10 0.12–0.67 < 0.001
Gemini / Claude 0.28 0.10 0.12–0.64 < 0.001 0.11 0.04 0.04–0.26 < 0.001
Paediatric dentistry ChatGPT4 / Gemini 0.69 0.30 0.24–1.94 0.39 1.37 0.62 0.46–4.07 0.495
ChatGPT4 / Claude 0.24 0.11 0.08–0.70 < 0.005 0.21 0.10 0.06–0.68 0.003
Gemini / Claude 0.34 0.15 0.12–1.01 0.035 0.15 0.08 0.05–0.51 < 0.001
Dental surgery ChatGPT4 / Gemini 1.28 0.51 0.49–3.35 0.546 3.28 1.57 1.04–10.36 0.015
ChatGPT4 / Claude 0.39 0.16 0.14–1.07 0.051 0.32 0.14 0.11–0.89 0.015
Gemini / Claude 0.31 0.13 0.11–0.84 0.015 0.10 0.05 0.03–0.31 < 0.001
Prosthetic dentistry ChatGPT4 / Gemini 0.61 0.27 0.21–1.78 0.813 1.10 0.47 0.39–3.06 0.830
ChatGPT4 / Claude 0.91 0.40 0.31–2.62 0.824 2.58 1.14 0.89–7.43 0.097
Gemini / Claude 1.48 0.65 0.51–4.26 0.813 2.35 1.04 0.82–6.77 0.106
Periodontology ChatGPT4 / Gemini 0.67 0.35 0.19–2.32 0.874 6.76 3.77 1.78–25.65 0.002
ChatGPT4 / Claude 0.45 0.23 0.13–1.57 0.371 1.00 0.52 0.29–3.43 1.000
Gemini / Claude 0.67 0.35 0.19–2.32 0.874 0.15 0.08 0.04–0.56 0.002
Orthodontics ChatGPT4 / Gemini 0.37 0.19 0.11–1.25 0.151 3.06 1.56 0.90–10.36 0.085
ChatGPT4 / Claude 0.42 0.21 0.13–1.41 0.171 1.29 0.64 0.39–4.27 0.617
Gemini / Claude 1.13 0.56 0.34–3.72 0.803 0.42 0.21 0.13–1.41 0.171
Integrated medicine ChatGPT4 / Gemini 7.44 4.15 1.96–28.28 < 0.001 1.12 0.53 0.36–3.46 0.814
ChatGPT4 / Claude 0.28 0.21 0.05–1.70 0.092 0.20 0.11 0.05–0.79 0.011
Gemini / Claude 0.04 0.03 0.01–0.22 < 0.001 0.18 0.10 0.04–0.71 0.008

OR odds ratio, SE standard error, CI confidence interval, p adjust – the adjusted p-value for the given comparison, accounting for Holm’s correction for multiple comparisons

Fig. 5.

Fig. 5

Analysis of differences between group within question specialties based on prompt language for chatbots.

In the specialty of integrated medicine, ChatGPT-4 exhibited over 7 times higher probability of accuracy compared to Gemini (p < 0.001). Gemini showed 96% lower probability of accuracy compared to Claude (p < 0.001). Both effect sizes are large.

Polish language

In the specialty of conservative dentistry, ChatGPT-4 showed over 7 times higher probability of accuracy compared to Gemini (p = 0.004, medium effect size). Gemini showed almost 90% lower probability of accuracy compared to Claude (p < 0.001, large effect size).

No statistically significant differences were detected among the chatbots in the specialties of prosthetic dentistry and orthodontics regarding the probability of accuracy to exam questions.

In the specialty of periodontology, ChatGPT-4 showed 6.76 times higher probability of providing a correct answer compared to Gemini (p = 0.002). Gemini showed 85% lower probability of providing a correct answer compared to Claude (p = 0.002). Both effect sizes are large.

Based on these results, the seventh hypothesis is supported.

Discussion

In this study, chatbots showed significant consistency in their responses between English and Polish. The results of Cohen’s Kappa analysis indicate that Claude demonstrated moderate accuracy consistency (κ = 0.59), regardless of the prompt language. While this suggests a reasonable level of reliability, the moderate agreement also underscores the need to further investigate factors that may contribute to performance variability. These results indicate that Claude gave the most consistent answers to the same dental examination questions regardless of the number of measurements. All chatbots showed satisfactory inter-rater reliability, emphasising their potential to generate consistent answers in dental and medical contexts. Remarkably, Gemini showed the lowest agreement in answers among the analysed chatbots, with a surprisingly lower consistency in English prompts compared to Polish ones. While response accuracy was assessed by comparison with the official answer key, the interrater agreement (kappa) reflects consistency between chatbot outputs, not correctness. These results formed the basis for the inclusion of the language variable in both generalised linear mixed models.

Model 1, which includes both fixed and random effects, explained 62,4% of the variance in responses, while fixed effects alone explain only 2.7% of the variance. This suggested that the characteristics of the questions significantly influence the differences in chatbot responses.

Model 2 showed that the probability of a correct answer to questions in the field of integrated medicine was higher for all chatbots than in the field of dentistry, regardless of the language of the prompt. These results suggesed that, in this sample, Claude responded more accurately to questions from integrated medicine than to those from dentistry; however, this does not imply general superiority in the field.

In recent years, there has been growing interest in the ability of chatbots to pass medical exams. Increasingly, researchers are investigating the potential of chatbots to enhance the education of medical students and professionals. Analysing the accuracy of chatbots can help improve these tools and enable their more effective use in the educational process. It can also help users to select the chatbot that works best in a particular area, which could lead to better support in medical education and practise28.

A recent study of Chau et al. examined the performance of ChatGPT 4.0, in dental licensing exams. Researchers posed 1,461 multiple-choice questions from dental licensing exams in the US and UK to two versions of chatGPT - chatGPT 3.5 and chatGPT 4.0. ChatGPT 4.0 passed both tests, answering 80.7% of the questions correctly in the US dental licensing exam and 62.7% correctly in the UK dental licensing exam5. In this comparative analysis, in the Polish dental licensing exam, ChatGPT 4.0 correctly answered 58% of the questions in English and 61% of the questions in Polish, thereby achieving the required score.

A study of Ahmad et al. compared the performance of ChatGPT-4o, Claude 3 OPUS, and Gemini Advanced on the 2024 periodontology in-service examination. ChatGPT-4o achieved the highest accuracy in all domains, ranging from 85.7 to 100%, significantly outperforming Claude 3 OPUS and Gemini Advanced29. Kung et al. demonstrated in their research that ChatGPT exhibited moderate performance on the USMLE exam, with accuracy rates varying between 45.4% and 75.0% based on question type, approaching but not consistently achieving passing scores30. Weng et al. evaluated the performance of ChatGPT on Taiwan’s Family Medicine Board Exam. Surprisingly, ChatGPT-3.5 answered 52 out of 125 questions correctly, which corresponds to an accuracy rate of 41.6%, which was insufficient to pass the exam31. A study by D’Anna et al. evaluated the performance of ChatGPT-4 and Google Bard on European Society of Neuroradiology (ESNR) course exams. ChatGPT-4 showed an overall accuracy of 70%, clearly outperforming Google Bard. ChatGPT-4 passed all four multiple-choice exams (MCEs) from the ESNR courses, while Google Bard did not32.

Several studies have evaluated the performance of chatbots, particularly ChatGPT4, in Polish medical exams. ChatGPT failed Polish Board Certification Examination in Internal Medicine. ChatGPT’s-4 performance on the PES in internal medicine was challenging, with an average success rate of 49.4%, falling short of the 60% passing threshold1.

In another study, ChatGPT-3.5 was reported to have passed the final medical examination in Poland. The performance was satisfactory, but still worse than that of human doctors, as they answered 53.4–64.9% of the questions correctly. In 8 out of 11 exam sessions, ChatGPT-4 achieved the score required to pass the examination (60%)13.

Our study showed statistically significant differences between chatbots in periodontology in terms of the probability of accuracy only for exam questions in Polish. ChatGPT-4 showed 6.76 times higher probability of providing a correct answer compared to Gemini. Gemini showed 85% lower probability of providing a correct answer compared to Claude.

Collectively, prior research highlighted that chatbot performance varied depending on exam format, domain specificity, and language. The current study built on this foundation by demonstrating that language-related factors could significantly influence chatbot accuracy, particularly in specialized domains like periodontology. This underscored the importance of linguistic considerations in the application of AI tools in medical education.

Limitations of the study

This study focussed on investigating the accuracy of three chatbots: ChatGPT-4, Gemini and Claude. The analysis focused on specific versions of these chatbots, which may affect the generalisability of the results to other models. Moreover, the results may vary depending on the prompt language. Our study focused on Polish and English, which may limit the generalisability of the results to other languages. The investigation was limited to a single, recent examination paper. Whilst this approach provides current insights, it may not fully represent the performance of chatbots on a wider range of dental topics or question types.

This study focused on evaluating the performance of three specific versions of chatbot models using a standardized set of medical examination questions in dentistry. Although the use of a single question set enabled controlled comparison, it may limit the generalizability of findings to other medical subfields, question types, or complexity levels. Moreover, as the questions originated from a single source, their currency does not necessarily ensure content diversity or representativeness across the broader scope of dental education.

Furthermore, the analysis was restricted to prompts in Polish and English, which narrows the linguistic scope of the conclusions. Language-specific features, such as grammatical complexity, semantic structure, and cultural framing, may differentially affect model performance, underscoring the need for multilingual assessment in future research.

Another limitation lies in the static nature of the AI model versions used. Given the rapid pace of development and frequent updates to large language models, the performance observed in this study may not reflect future iterations of the same systems. Consequently, longitudinal evaluations are needed to monitor performance stability over time.

In addition, this study did not systematically account for prompt engineering strategies, which are known to influence chatbot output. The prompts used were intentionally neutral, yet variations in structure, tone, or specificity could yield different results. Future studies should explore the role of prompt design as a moderator of model accuracy. Taken together, these limitations underscore the exploratory nature of this study and indicate that its findings should be interpreted within the specific constraints of language, content scope, and chatbot version used.

Further research is essential to more fully assess the educational potential of chatbots in medicine. The findings reported here carry crucial implications for the national medical-education system. First, the study shows that ChatGPT-4, Gemini, and Claude can meet or even surpass the pass threshold on both the Polish Medical-Dental Final Examination (LDEK) and the Medical-Dental Verification Examination (LDEW).

Chatbots not only identify the correct answers but also generate detailed, step-by-step justifications. Such annotated solutions can expedite the diagnosis of knowledge gaps, foster example-based learning, and enable students to develop their own examination strategies. Such benefits are unattainable when relying exclusively on a traditional answer key, as the high reliability and accuracy of responses are crucial to ensure that answers to medical questions are both trustworthy and aligned with the current state of medical knowledge33,34.

To effectively support medical education, the use of large language models (LLMs) requires careful consideration, as inaccurate responses can have potentially severe consequences for both the medical education of students and future patient care outcomes. Moreover, LLMs may inherit biases embedded in their training data and generate responses containing hallucinations, such as fabricated or incorrect information. Students must therefore be mindful of the limitations of LLMs and exercise rigorous critical thinking when utilizing these tools35,36.

This study highlighted the potential utility of chatbots as educational tools for dental students preparing for high-stakes licensing examinations, such as the LDEK. Evaluating the accuracy and reliability of their responses serves as a foundational step toward in-depth analyses of the correct answers generated by chatbots, representing a promising direction for future research.

Conclusions

  1. Self-assessment and effectiveness:

The results suggested that chatbots had some ability to recognise the relative difficulty of questions, but the accuracy of this assessment is limited.

  • 2.

    Self-evaluation and actual performance:

To some extent chatbots could estimate their answers reasonably correctly. However, caution is required when interpreting the results. The small effect size indicates that the chatbots’ self-assessment may not be a completely reliable indicator of their actual effectiveness.

  • 3.

    Linguistic consistency:

Chatbot responses to the same questions may vary depending on the language of the prompt. This linguistic variability should be considered when using chatbots in multilingual contexts or for tasks requiring consistency between languages.

  • 4.

    Measurement consistency:

The results suggested that chatbot responses to the same questions were not completely predictable or repeatable. Differences in consistency could result from various factors: model updates between measurements, built-in randomness in generating responses, differences in interpreting the context of questions at different points in time.

  • 5.

    Accuracy in different prompt languages:

Chatbots differed in the accuracy of providing correct answers for individual prompt languages. The performance differences were more pronounced in Polish than in English, especially when you compare Gemini with other chatbots.

  • 6.

    Accuracy across all languages:

Only Gemini differed significantly in accuracy of correct answers between Polish and English prompts. The accuracy of Gemini was significantly (by 31.25%) higher in English than in Polish.

  • 7.

    Specialty-specific accuracy:

Chatbot accuracy was highest in the integrated medicine domain, indicating greater accuracy for general medicine questions than dental questions Our results demonstrated that Claude achieved the highest accuracy in all areas analysed and outperformed other chatbots. This suggests that Claude has significant potential to support the medical education of dental students.

This study showed that the performance of chatbots varied depending on the prompt language and the specific field. For instance, ChatGPT-4 showed better performance than Gemini in some areas, especially in Polish. This highlights the importance of considering language and specialty when selecting a chatbot for educational purposes.

Supplementary Information

Below is the link to the electronic supplementary material.

Supplementary Material 1 (31.5KB, docx)

Author contributions

Investigation and research: O.A., O.T., G.C. Writing - original draft: D.W.Figures and tables preparation: D.W. Data analysis: D.W. Conceptualization: D.W., L.S., O.A., O.T., G.C. Methodology: D.W. Writing - review & editing: All authorsSupervision: L.S., D.W.

Funding

This study did not receive any funding.

Data availability

The datasets generated and analysed during the current study are available in the Mendeley Data repository doi: https://data.mendeley.com/datasets/2f59x25mvm/1 doi: 10.17632/2f59x25mvm.1.

Declarations

Competing interests

The authors declare no competing interests.

Ethics approval

The Ethics Committee determined that the evaluation of research involving Large Language Models (LLMs) does not require their approval and falls outside their purview. An application for ethical approval had been submitted to the Bioethics Committee at the Medical University of Lublin.

Consent to participate

Not applicable.

Clinical trial number

Not applicable.

Patient consent statement

Not applicable.

Permission to reproduce material from other sources

Not applicable.

Footnotes

The original online version of this Article was revised: The original version of this Article contained an error in the title of the paper. The title now reads: “Comparing the performance of ChatGPT, Gemini, and Claude in English and Polish on medical examinations.”

Publisher’s note

Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.

Change history

12/15/2025

A Correction to this paper has been published: 10.1038/s41598-025-31400-8

References

  • 1.Wójcik, S. et al. Reshaping medical education: Performance of ChatGPT on a PES medical examination. Cardiol. J. (2023).
  • 2.Lewandowski, M., Łukowicz, P., Świetlik, D. & Barańska-Rybak, W. An original study of ChatGPT-3.5 and ChatGPT-4 dermatological knowledge level based on the dermatology specialty certificate examinations. Clin Exp. Dermatol llad255 (2023).
  • 3.Gilson, A. et al. How does ChatGPT perform on the united States medical licensing examination (USMLE)? The implications of large Language models for medical education and knowledge assessment. JMIR Med. Educ.9, e45312 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 4.Stathakarou, N. et al. Students’ perceptions on chatbots’ potential and design characteristics in healthcare education. In The Importance of Health Informatics in Public Health During a Pandemic 209–212 (IOS Press, 2020).
  • 5.Chau, R. C. W. et al. Performance of generative artificial intelligence in dental licensing examinations. Int. Dent. J.74, 616–621 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6.Botross, M., Mohammadi, S. O., Montgomery, K., Crawford, C. & Montgomery, K. E. Performance of google’s artificial intelligence chatbot bard(now Gemini) on ophthalmology board exam practice questions. Cureus16 (2024).
  • 7.Al Kahf, S. et al. Chatbot-based serious games: A useful tool for training medical students? A randomized controlled trial. PloS One. 18, e0278673 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 8.Şahin, M. F. et al. Which current chatbot is more competent in urological theoretical knowledge? A comparative analysis by the European board of urology in-service assessment. World J. Urol.43, 116 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 9.Shieh, A. et al. Assessing ChatGPT 4.0’s test performance and clinical diagnostic accuracy on USMLE STEP 2 CK and clinical case reports. Sci. Rep.14, 9330 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 10.Mohammad-Rahimi, H. et al. Deep learning in periodontology and oral implantology: A scoping review. J. Periodontal Res.57, 942–951 (2022). [DOI] [PubMed] [Google Scholar]
  • 11.Azamfirei, R., Kudchadkar, S. R. & Fackler, J. Large Language models and the perils of their hallucinations. Crit. Care. 27, 120 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12.Alfertshofer, M. et al. Sailing the seven seas: A multinational comparison of chatgpt’s performance on medical licensing examinations. Ann. Biomed. Eng.52, 1542–1545 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 13.Suwała, S. et al. ChatGPT-3.5 passes poland’s medical final examination—Is it possible for ChatGPT to become a Doctor in poland?? SAGE Open. Med.12, 20503121241257777 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 14.Funder, D. C. & Ozer, D. J. Evaluating effect size in psychological research: Sense and nonsense. Adv. Methods Pract. Psychol. Sci.2, 156–168 (2019). [Google Scholar]
  • 15.Cohen, J. Statistical Power Analysis for the Behavioral Sciences (Routledge, 1988).
  • 16.McHugh, M. L. Interrater reliability: The kappa statistic. Biochem. Med.22, 276–282 (2012). [Google Scholar]
  • 17.Fleiss, J. L., Levin, B. & Paik, M. C. Statistical methods for rates and proportions (2003).
  • 18.Powell, M. J. The BOBYQA algorithm for bound constrained optimization without derivatives. Cambridge NA Report NA2009/06, University of Cambridge, Cambridge 26 26–46 (2009).
  • 19.Wickham, H. & Bryan, J. Read Excel Files. R package version 1.3. 1. (2019).
  • 20.Wickham, H., François, R., Henry, L., Müller, K. & Vaughan, D. dplyr: A Grammar of Data Manipulation. R package version 1.1.4 (2023).
  • 21.Ben-Shachar, M. S., Lüdecke, D. & Makowski, D. Estimation of effect size indices and standardized parameters. J. Open. Source Softw.5, 2815 (2020). [Google Scholar]
  • 22.Bates, D., Mächler, M., Bolker, B. & Walker, S. Fitting linear mixed-effects models using lme4. J. Stat. Softw.67, 1–48 (2015). [Google Scholar]
  • 23.Lüdecke, M. D. Package ‘sjPlot’ (2024).
  • 24.Gamer, M., Lemon, J., Gamer, M. M. & Robinson, A. Kendall’s, W. Package ‘irr’. Var. Coeff Interrater Reliab. Agreem.22, 1–32 (2012). [Google Scholar]
  • 25.Wickham, H. Tidyr Tidy Messy Data: R Package Version 1.3.1 2024. https://cran.r-proj.org/package/tidyr (2024).
  • 26.Lüdecke, D. et al. See: An R package for visualizing statistical models. J. Open. Source Softw.6, 3393 (2021). [Google Scholar]
  • 27.Wickham, H. & Wickham, H. Data Analysis (Springer, 2016).
  • 28.Al-Rawas, M. et al. Identification of dental related ChatGPT generated abstracts by senior and young academicians versus artificial intelligence detectors and a similarity detector. Sci. Rep.15, 11275 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 29.Ahmad, B., Saleh, K., Alharbi, S., Alqaderi, H. & Jeong, Y. N. Artificial intelligence in periodontology: performance evaluation of ChatGPT, claude, and gemini on the in-service examination. medRxiv 2024.05.29.24308155 (2024). 10.1101/2024.05.29.24308155
  • 30.Kung, T. H. et al. Performance of ChatGPT on USMLE: Potential for AI-assisted medical education using large Language models. PLoS Digit. Health. 2, e0000198 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 31.Weng, T. L., Wang, Y. M., Chang, S., Chen, T. J. & Hwang, S. J. ChatGPT failed taiwan’s family medicine board exam. J. Chin. Med. Assoc.86, 762–766 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 32.D’Anna, G., Van Cauter, S., Thurnher, M., Van Goethem, J. & Haller, S. Can large language models pass official high-grade exams of the European Society of Neuroradiology courses? A direct comparison between OpenAI chatGPT 3.5, OpenAI GPT4 and Google Bard. Neuroradiology 1–6 (2024).
  • 33.Abd-Alrazaq, A. et al. Large Language models in medical education: Opportunities, challenges, and future directions. JMIR Med. Educ.9, e48291 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 34.Labadze, L., Grigolia, M. & Machaidze, L. Role of AI chatbots in education: Systematic literature review. Int. J. Educ. Technol. High. Educ.20, 56 (2023). [Google Scholar]
  • 35.Li, R. & Wu, T. Delving into the practical applications and pitfalls of large language models in medical education: Narrative review. Adv. Med. Educ. Pract. 625–636 (2025).
  • 36.Zhui, L. et al. Impact of large Language models on medical education and teaching adaptations. JMIR Med. Inf.12, e55933 (2024). [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Supplementary Material 1 (31.5KB, docx)

Data Availability Statement

The datasets generated and analysed during the current study are available in the Mendeley Data repository doi: https://data.mendeley.com/datasets/2f59x25mvm/1 doi: 10.17632/2f59x25mvm.1.


Articles from Scientific Reports are provided here courtesy of Nature Publishing Group

RESOURCES