Abstract
Objective
This study aimed to evaluate the accuracy and consistency of responses provided by three large language models (LLMs), ChatGPT-5.2, Gemini-3, and DeepSeek-V3.2, to multiple-choice questions based on undergraduate endodontic education, asked on different days and at different times of the day.
Materials and methods
A total of 60 text-based multiple-choice questions were developed across six undergraduate endodontic topics: dental caries, pulpitis, apical periodontitis, periapical abscess, root fracture, and root resorption. Each question was presented to ChatGPT-5.2, Gemini-3, and DeepSeek-V3.2 at three time points per day (morning, afternoon, and evening) over four consecutive days. Accuracy and response consistency were analyzed using SPSS and R software, with statistical significance set at p < 0.05 and a 95% confidence interval.
Results
ChatGPT-5.2 and Gemini-3 demonstrated significantly higher accuracy and consistency than DeepSeek-V3.2 (p < 0.001 and p = 0.004, respectively). Model performance varied according to question category. Accuracy differed significantly across categories for ChatGPT-5.2 and Gemini-3, whereas consistency was influenced by question category only in ChatGPT-5.2. Model performance remained largely stable across different assessment times.
Conclusions
Advanced LLMs demonstrated promising performance in answering undergraduate endodontic multiple-choice questions and may serve as useful adjunctive tools in dental education. However, differences among models and variations in performance across topics highlight the need for critical evaluation of AI-generated responses before their educational use.
Supplementary Information
The online version contains supplementary material available at https://doi.org/10.1186/s12903-026-09174-w.
Keywords: Accuracy, Artificial intelligence, ChatGPT, Consistency, DeepSeek, Education, Endodontics, Gemini, Large language models
Introduction
Artificial intelligence (AI) refers to systems capable of performing tasks that typically require human cognitive abilities, such as learning, reasoning, and problem solving [1]. Through learning from data and experience, AI systems can improve their performance and adapt to changing environments. In dentistry and healthcare, AI technologies are increasingly being used to support diagnosis, predict treatment outcomes, enhance telemedicine services, optimize clinical workflows, and facilitate research and education [2]. Natural language processing (NLP), a subfield of AI, enables computers to interpret and process human language, thereby improving human–computer interaction. Recent advances in large language models (LLMs), which utilize NLP techniques to generate responses to text-based inputs, have attracted considerable attention in healthcare and education.
A wide range of LLM-based chatbots with varying capabilities have been developed in recent years. Interest in LLM-based conversational systems has grown markedly since the public release of ChatGPT in November 2022, followed by the introduction of Google’s conversational model, initially known as Bard and later renamed Gemini, in February 2023 [3, 4]. OpenAI introduced ChatGPT-5.2 on 11 December 2025, a highly capable model designed for advanced text-based interaction [5]. Similarly, Google launched Gemini-3 with enhanced multimodal capabilities [6], while DeepSeek-V3.2 has emerged as a widely used open-source alternative [7]. As these systems continue to evolve and become increasingly accessible, both patients and healthcare professionals are relying on them more frequently to obtain information on a wide range of topics, including healthcare-related inquiries.
The ability to apply clinical techniques and the level of medical knowledge possessed by clinicians have long been key determinants of accurate diagnosis and effective treatment in medicine and dentistry. Consequently, medical and dental education require students to acquire extensive theoretical knowledge, which forms a fundamental component of professional examinations [1]. Access to specialized medical knowledge has traditionally been limited to trained professionals, thereby restricting patients’ direct access to reliable health information. AI-driven systems such as ChatGPT have demonstrated potential educational benefits by supporting clinical reasoning, generating diagnostic and treatment suggestions through case-based scenarios, and enhancing communication, problem-solving, and critical-thinking skills [1, 8–10]. Accordingly, several studies have evaluated the performance of AI systems in medical and dental education, including their performance on major licensing examinations such as the United States Medical Licensing Examination and national dental examinations in different countries [11–13].
The use of AI for educational purposes is becoming increasingly popular among both educators and students, highlighting the growing need to systematically evaluate the performance of newly developed chatbot models. This study aimed to assess the accuracy and consistency of responses generated by ChatGPT-5.2, Gemini-3, and DeepSeek-V3.2 to multiple-choice questions covering key undergraduate endodontic topics, including dental caries, pulpitis, apical periodontitis, periapical abscess, root fracture, and root resorption. To evaluate the stability of model performance, responses were collected at different times of the day and across multiple days. The null hypothesis of this study was that there would be no significant differences among the LLMs in terms of accuracy and consistency.
Materials and methods
Ethical approval was not required for this study as no human or animal participants were involved. The ESE Undergraduate Curriculum Guidelines for Endodontology were used to define the scope and content of the topics [14]. These guidelines outline the minimum competencies expected of undergraduate dental students, integrating scientific knowledge with clinical skills and emphasizing evidence-based and patient-centered care. Accordingly, the questions were categorized into six thematic groups based on the underlying endodontic condition: Group A (dental caries), Group B (pulpitis), Group C (apical periodontitis), Group D (periapical abscess), Group E (root fracture), and Group F (root resorption), with 10 questions in each group.
A total of 60 multiple-choice, text-based questions were developed by three endodontists using two widely accepted reference textbooks: Endodontics: Principles and Practice (Fifth and Sixth edition) [15], and Cohen’s Pathways of the Pulp (Tenth edition) [16]. The questions were independently formulated by the authors and were not directly reproduced from the textbooks or derived from publicly indexed question banks. This approach was adopted to minimize the risk of data leakage while ensuring alignment with established undergraduate endodontic curricula.
The validation process was conducted in accordance with the Item-Level Content Validity Index (I-CVI) methodology as described by Yusoff et al. [17]. A 4-point Likert scale was employed, in which ratings of 3 (quite relevant) and 4 (highly relevant) were considered acceptable and scored as 1. Each item was evaluated independently by three endodontic experts, and only those items that received ratings of 3 or 4 from all three experts were considered to have achieved an I-CVI value of 1.00 and were therefore deemed acceptable for inclusion in the study. Based on the experts’ feedback, several items were revised to improve clarity, readability, and overall comprehensibility while maintaining their scientific integrity. Following the validation procedure, 60 content-validated questions was created. The questions were presented to the AI models, and the responses were evaluated by comparing them with the reference answers provided in the textbooks. The complete set of questions and corresponding answers are available in the Supplemental File.
Querying procedure
Three LLMs were evaluated: ChatGPT-5.2 (OpenAI, San Francisco, CA, USA; https://chat.openai.com), Gemini 3 (Google Inc., Mountain View, CA, USA; https://gemini.google.com ), and DeepSeek V3.2 (https://chat.deepseek.com). All models were accessed via their official web-based interfaces (non-API access). All models were used with default settings, and no modifications were made to parameters such as temperature, top-p, or maximum token limits. Browsing features and external tools were not enabled in this study.
Data collection was performed over four consecutive days (January 18–21, 2026). Each question was submitted in English at three time points each day (morning, afternoon, and evening). To minimize contextual carryover, each query was entered into a newly initiated, independent chat session with no retained conversation history. No user accounts were personalized, and no follow-up prompts, clarifications, corrections, or feedback were provided during the study. All models were accessed through their freely available web interfaces using standard user accounts. Model selection was performed manually through the respective web platforms. Memory or personalization features were disabled when available, and no reasoning/thinking modes or advanced model-specific options were activated. The default chat interfaces were used throughout the study.
A standardized prompting protocol was used. All questions were presented verbatim using a copy-and-paste approach without any modifications to the wording, punctuation, or syntax. The exact prompt used for all models was: “Please select the single best answer for the following multiple-choice question. Respond only with the letter corresponding to your chosen option. Do not provide explanations or additional comments.” The same prompt structure was applied consistently across all models and query sessions. No prompt engineering or optimization was performed. To assess response stability, each question was repeated under identical conditions.
The primary outcome of this study was the accuracy of LLM responses, defined as the proportion of correct answers across repeated queries for each question. Response consistency was evaluated as a secondary outcome and defined as the proportion of the most frequently selected response across all repeated queries for a given question. For each thematic group, both accuracy and consistency were summarized as mean ± standard deviation based on question-level ratios.
Statistical analysis
All statistical analyses were performed using R software (version 4.4.1) and Stata softare (version 14). Confidence intervals for success rates, both within and between the models, were calculated using the binomial Wald method. Comparisons of success rates were conducted using the Pearson chi-square test and the Fisher’s exact test with Monte Carlo correction when appropriate. Multiple comparisons were performed using the Bonferroni-adjusted Z test.
For comparisons of consistency values across three or more groups, the Kruskal–Wallis H test was applied, followed by Dunn’s test for post-hoc multiple comparisons. Quantitative variables were presented as median (minimum–maximum), whereas categorical variables were expressed as frequency and percentage. Statistical significance was set at p < 0.05.
Results
No patient or public involvement was included in the design, conduct, reporting, interpretation, or dissemination of this study. Across all question categories and time points, ChatGPT-5.2 and Gemini-3 demonstrated significantly higher accuracy than DeepSeek-V3.2 (p<0.001; Table 1). No significant difference in accuracy was observed between ChatGPT-5.2 and Gemini-3 across the question categories. However, the accuracy differed significantly among the question groups for ChatGPT-5.2 and Gemini-3 (both p<0.001), whereas no significant difference was observed across question groups for DeepSeek-V3.2 (p = 0.058) (Table 1).Specifically, ChatGPT-5.2 showed higher accuracy than DeepSeek-V3.2 in Group A; both ChatGPT-5.2 and Gemini-3 outperformed DeepSeek-V3.2 in Groups D and F; and Gemini-3 demonstrated higher accuracy than ChatGPT-5.2 in Group E.
Table 1.
Comparison of correct response rates within and between models
| ChatGPT 5.2 | Gemini 3 | DeepSeek V3.2 | Test statistics | p | |||||||
|---|---|---|---|---|---|---|---|---|---|---|---|
| Value (%) p | CI 95%-Wald binominal |
Standard error |
Value (%) p | CI 95%-Wald binominal |
Standard error |
Value (%) p | CI 95%-Wald binominal |
Standard error |
|||
| Group A | 108 (90) aA | 84,632: 95,368 | 0,027 | 106 (88,3) abA | 82,590: 94,077 | 0,029 | 93 (77,5) b | 70,029: 84,971 | 0,038 | 8,253 | 0,016x |
| Group B | 107 (89,2) A | 83,606: 94,728 | 0,028 | 98 (81,7) A | 74,744: 88,590 | 0,035 | 95 (79,2) | 71,900: 86,433 | 0,037 | 4,833 | 0,109x |
| Group C | 104 (86,7) A | 80,585: 92,749 | 0,031 | 107 (89,2) A | 83,606: 94,728 | 0,028 | 106 (88,3) | 82,590: 94,077 | 0,029 | 0,387 | 0,880x |
| Group D | 119 (99,2) aB | 97,540: 100,000 | 0,008 | 120 (100) aB | 00,000: 100,000 | 0,000 | 103 (85,8) b | 79,594: 92,072 | 0,032 | 29,225 | < 0,001x |
| Group E | 94 (78,3) aA | 70,962: 85,704 | 0,038 | 108 (90) bA | 84,632: 95,368 | 0,027 | 103 (85,8) ab | 79,594: 92,072 | 0,032 | 6,299 | 0,048x |
| Group F | 108 (90) aA | 84,632: 95,368 | 0,027 | 108 (90) aA | 84,632: 95,368 | 0,027 | 91 (75,8) b | 68,174: 83,493 | 0,039 | 11,845 | 0,001x |
| Total | 640 (88,9) a | 86,593: 91,184 | 0,012 | 647 (89,9) a | 87,656: 92,066 | 0,011 | 591 (82,1) b | 79,282: 84,884 | 0,014 | 22,783 | < 0,001y |
| Test statistics | 30,936 | 29,824 | 11,023 | ||||||||
| p | < 0,001x | < 0,001x | 0,058x | ||||||||
x Fisher’s Exact Test with Monte Carlo correction
y Pearson’s chi-square test; frequency (percentage)
a–b No significant difference between language models sharing the same letter
A–B No significant difference between groups sharing the same letter. Bold values indicate statistically significant differences (p < 0.05).
Across all question categories and time points, consistency differed significantly among the LLMs (p = 0.004; Table 2). ChatGPT-5.2 and Gemini-3 demonstrated significantly higher consistency than DeepSeek-V3.2, whereas the consistency levels of ChatGPT-5.2 and Gemini-3 were comparable. Question categories significantly affected consistency in ChatGPT-5.2 (p = 0.023), whereas no significant effect was observed for Gemini-3 or DeepSeek-V3.2. At the group level, ChatGPT-5.2 showed higher consistency than DeepSeek-V3.2 in Groups A and F, and Gemini-3 demonstrated higher consistency in Groups D and F (p<0.05).
Table 2.
Comparison of consistency values within and between models
| Group | ChatGPT 5.2 | Gemini | DeepSeek | Test statistics | p x |
|---|---|---|---|---|---|
| A | 100 (100: 100) / 19,5 a | 100 (83,33: 100) / 15,55 ab | 95,83 (58,33: 100) / 11,45 b | 6,912 | 0,032 |
| B | 100 (66,67: 100) / 15,75 | 100 (66,67: 100) / 16,35 | 100 (58,33: 100) / 14,4 | 0,367 | 0,832 |
| C | 100 (58,33: 100) / 13,65 | 100 (91,67: 100) / 18,35 | 100 (58,33: 100) / 14,5 | 2,676 | 0,262 |
| D | 100 (91,67: 100) / 16,65 ab | 100 (100: 100) / 18 a | 100 (83,33: 100) / 11,85 b | 6,412 | 0,041 |
| E | 95,83 (75: 100) / 14,45 | 100 (50: 100) / 15,75 | 100 (83,33: 100) / 16,3 | 0,3 | 0,861 |
| F | 100 (100: 100) / 18 a | 100 (100: 100) / 18 a | 95,83 (66,67: 100) / 10,5 b | 11,484 | 0,003 |
| Total | 100 (58,33: 100) / 94,94 a | 100 (50: 100) / 99,93 a | 100 (58,33: 100) / 76,63 b | 11,015 | 0,004 |
| Test statistics | 13,028 | 8,260 | 1,348 | ||
| p x | 0,023 | 0,143 | 0,930 |
x Kruskal–Wallis H test; median (minimum–maximum)
a–b No significant difference between language models sharing the same letter. Bold values indicate statistically significant differences (p < 0.05).
No statistically significant differences in accuracy rates were observed between the AI models within the daytime interaction groups (p > 0.05). Similarly, within each language model, accuracy rates did not differ significantly across the daytime interaction groups (p > 0.05). Statistically significant differences in accuracy rates among the AI models were observed during the morning sessions of Day 1 and Day 3 (p = 0.026 and p = 0.012, respectively), with Gemini-3 demonstrating significantly higher accuracy than DeepSeek-V3.2 in both comparisons.
Discussion
This study evaluated the performance of ChatGPT-5.2, Gemini-3, and DeepSeek-V3.2 in answering undergraduate endodontic multiple-choice questions in terms of accuracy and response consistency. ChatGPT-5.2 and Gemini-3 demonstrated significantly higher accuracy and consistency than DeepSeek-V3.2, whereas no significant differences were observed between the two leading models. In addition, model performance varied according to question category, indicating that the ability of LLMs to answer endodontic questions may depend on the specific topic being assessed.
Hallquist et al. [18] conducted a systematic review and reported that AI has diverse applications in medical education, including generating exam questions, supporting clinical training, and assisting student learning. The authors highlighted the considerable potential of AI in medical education. However, understanding the capabilities and limitations of AI is essential before its effective integration into educational practice. Therefore, the present study evaluated the accuracy and consistency of several contemporary LLMs in answering multiple-choice questions on endodontics.
Similar findings have been reported in previous studies evaluating the performance of LLMs in medical and dental education. Jaworski et al. [19] reported that GPT-4o and GPT-5 achieved comparable accuracy on a gastroenterology specialty examination, while GPT-5 demonstrated better alignment between confidence and correctness. Luitel et al. [20] reported that ChatGPT-4 achieved an accuracy rate of 87.8% on the Nepal Medical Council Licensing Examination, demonstrating performance comparable to, or even exceeding, that of medical graduates. Balu et al. [21] reported that ChatGPT-3.5 could generate USMLE Step 1–style multiple-choice questions with reasonable factual accuracy, suggesting its potential utility as a supplementary tool for medical exam preparation. Li et al. [22] reported that the integration of generative AI tools in case-based learning significantly improved students’ learning efficiency and examination performance in medical biochemistry, suggesting that AI can serve as a valuable supplement to traditional medical education. Ros-Arlanzón et al. [23] reported that several LLMs achieved higher mean scores than medical students on multiple-choice clinical course examinations, demonstrating strong performance in institution-authored assessments. Lee et al. [24] reported that AI chatbot–based simulation and peer role-play showed complementary benefits for OSCE preparation, with chatbots supporting self-directed clinical reasoning practice and peer role-play enhancing realistic communication and examination skills. These findings support the growing evidence that advanced LLMs can demonstrate high levels of accuracy in medical knowledge assessments.
Jalali et al. [25] reported that the accuracy of AI chatbots in answering endodontic specialty examination questions ranged from 48% to 71%, highlighting both their potential and their current limitations in handling specialized endodontic knowledge. Among the evaluated models, Gemini Advanced, GPT-3.5, GPT-4o, and Claude 3.5 Sonnet demonstrated the highest performance, correctly answering more than 70% of the questions. Similarly, Abdulrab et al. evaluated chatbot performance using 100 multiple-choice questions derived from two standard endodontic textbooks, administered in two rounds separated by a one-week interval. ChatGPT-4o demonstrated the highest performance stability, followed by Gemini Advanced. Zhang et al. [26] reported that Gemini-3-Flash achieved the highest accuracy in answering ophthalmology multiple-choice questions. Şirani [27] reported that although the overall accuracy of ChatGPT-4o, DeepSeek-R1, and Gemini-2 Pro on fixed prosthodontics questions was moderate, ChatGPT demonstrated the highest consistency for multiple-choice questions, whereas Gemini showed improved performanve over time on short-answer questions. Gokkurt Yilmaz et al. [28] showed that ChatGPTo1 demonstrated the highest accuracy among several LLMs when answering head and neck anatomy questions from the Dental Specialization Examination. Yılmaz et al. [29] reported that among several LLMs evaluated on oral pathology multiple-choice questions (ChatGPT-o1, Claude 3.5, Gemini-2, DeepSeek, ChatGPT-4o, ChatGPT-4, Gemini-1.5, and Copilot), ChatGPT-o1 demonstrated the highest accuracy. Azhari et al. [30] reported that Gemini achieved significantly higher accuracy and passing rates than ChatGPT-3.5 when answering dental caries–related multiple-choice questions across multiple simulated examination formats. Haylaz et al. [31] reported that ChatGPT-4 demonstrated the highest accuracy in answering oral and maxillofacial radiology multiple-choice questions, whereas DeepSeek-R1 showed the lowest performance. These findings are consistent with the results of the present study and further suggest that differences among LLMs may depend on their architecture, training data, and domain-specific optimizations. Such variations highlight the importance of tailoring LLMs for specific subject areas and underscore the need to consider the strengths and limitations of each model when interpreting its performance.
The findings of the present study show that the different question groups affected accuracy, similar to the results of various studies in the medical field using LLMs. Several factors may explain the observed differences in performance across question groups. The availability and accessibility of training data may play an important role, as some medical knowledge sources are not easily accessible through publicly available Internet databases [32]. In addition, the continuous updating of certain datasets may influence the performance of LLMs over time [33]. Another possible explanation is the uneven distribution of user interactions across topics. Topics that receive greater attention from users, including more frequent queries and feedback, may be better represented within model training or refinement processes, potentially leading to improved performance in those areas [34]. Consequently, variations in data accessibility, dataset updates, and user interaction patterns may contribute to differences in model performance across specific domains.
Inconsistent accuracy in AI-generated responses may increase the risk of users receiving inaccurate information [35]. In this study, ChatGPT-5.2 and Gemini 3 demonstrated comparable levels of response consistency. Similar observations have been reported in previous medical research, where Builoff et al. [36] found that ChatGPT-4, ChatGPT-4 Turbo, and Gemini produced responses with comparable consistency. Alshammari et al. [37] evaluated the performance of several AI language models in answering standardized dental multiple-choice questions and reported that ChatGPT and Gemini demonstrated higher accuracy and consistency, whereas DeepSeek showed comparatively poorer performance. These findings are consistent with those of the present study. Variations in study outcomes may be influenced by several methodological factors, including differences in question formats and the intervals between repeated queries [38]. Moreover, discrepancies in how consistency is defined and calculated across studies may further contribute to variability in the reported results. It should also be noted that high consistency does not necessarily indicate high accuracy. Because consistency was defined as the frequency of the most commonly repeated response, a model may achieve high consistency by repeatedly providing the same incorrect answer. Therefore, consistency should be interpreted as a measure of response stability rather than correctness or reliability.
The present findings indicate that ChatGPT-5.2 and Gemini-3 outperform DeepSeek-V3.2 in both accuracy and response consistency when answering undergraduate endodontic multiple-choice questions. Although the overall performances of ChatGPT-5.2 and Gemini-3 were comparable, both models demonstrated topic-dependent variations in accuracy, suggesting that the complexity or representation of specific endodontic concepts may influence LLM performance. In contrast, DeepSeek-V3.2 exhibited more uniform performance across question categories, albeit at a lower overall accuracy level. The absence of substantial time-related effects further suggests that model performance remained stable throughout the study period.
Differences in performance among the evaluated LLMs may partly be attributable to variations in their underlying training data, data curation strategies, and model architectures. ChatGPT-5.2 and Gemini-3 are proprietary models trained on large-scale, heterogeneous datasets that may include licensed content, web-based resources, and human feedback [39]. In contrast, DeepSeek-V3.2 relies more heavily on publicly available data sources [40]. However, because the exact composition and weighting of these datasets are not fully disclosed, the extent to which training data contributed to the observed differences remains uncertain.
Multiple-choice questions are widely used for objective assessment of knowledge [41]. In the present study, this format was selected to avoid ambiguity related to partially correct or partially incorrect responses. Accordingly, the responses of three LLMs, ChatGPT-5.2, Gemini-3, and DeepSeek-V3.2, to multiple-choice questions on endodontics were evaluated. Previous studies have suggested that question format may not substantially influence the performance of some AI models [42]. However, AI systems may occasionally generate inaccurate or fabricated information when an appropriate response cannot be produced [43]. To minimize this risk, the models were instructed to provide only the selected answer without additional explanations. Furthermore, each query was submitted in a new conversation session to reduce contextual bias. Model performance was assessed at three different time points each day over multiple days to evaluate the temporal stability and consistency of responses while minimizing the potential influence of time-dependent variations, such as stochastic response generation, fluctuations in server load, or platform-level optimizations.
The findings demonstrated no significant differences in accuracy across most time points, suggesting that the performance of the evaluated LLMs remained largely stable over time. Although Gemini-3 achieved significantly higher accuracy than DeepSeek-V3.2 during specific morning sessions, these isolated differences did not affect the overall temporal consistency of the models. Given that stochastic variability is an inherent characteristic of LLMs operating under default settings, repeated queries were intentionally performed to evaluate the stability of model outputs under real-world conditions. Therefore, response variability was considered an inherent characteristic of LLMs and was directly evaluated through the consistency analysis.
This study has several notable limitations. First, all questions were asked exclusively in English, which restricted the ability to evaluate the performance of AI language models in understanding and responding to queries in other languages or dialects. Additionally, the use of a stringent I-CVI threshold of 1.00 to ensure strong content validity may have resulted in the exclusion of some potentially relevant items. Although the questions were developed de novo, they were based on widely used endodontic reference textbooks. Therefore, prior exposure of the evaluated LLMs to similar concepts during pretraining cannot be completely excluded. The findings are specific to the versions of the LLMs available at the time of data collection; therefore, subsequent updates, retraining processes, or architectural modifications may lead to different performance characteristics in future studies.
Conclusion
Within the limitations of this study, ChatGPT-5.2 and Gemini-3 demonstrated higher accuracy and consistency than DeepSeek-V3.2 when answering undergraduate endodontic multiple-choice questions. Although the overall performance of the two leading models was comparable, accuracy varied across different question groups, suggesting that topic-related factors may influence LLM performance in dental knowledge assessments. Model performance remained largely stable across different assessment times. These findings indicate that advanced LLMs may serve as supportive educational tools in endodontic training. Nevertheless, their performance was not uniform across all question groups, highlighting the importance of critically evaluating AI-generated responses in dental education. Further studies are needed to evaluate the performance of LLMs across a wider range of dental disciplines and clinically oriented scenarios.
Supplementary Information
Acknowledgements
Not applicable.
Authors’ contributions
Meltem Sümbüllü: Conceptualization, Data curation, Investigation, Methodology, Writing−original draft, Writing−review and editing. Oğuzhan Ünal: Data curation, Investigation, Methodology, Writing−review and editing. İlke Menteş and Muzaffer Enes Kayahan: Methodology, Supervision, and Writing−review and editing. All authors contributed significantly to the study and approved the final version of the manuscript.
Funding
The authors declare that no funds, grants, or other support were received during the preparation of this manuscript.
Data availability
Data is provided within the supplementary information files.
Declarations
Ethics approval and consent to participate
Not applicable.
Competing interests
The authors declare no competing interests.
Footnotes
Publisher’s note
Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
References
- 1.Künzle P, Paris S. Performance of large language artificial intelligence models on solving restorative dentistry and endodontics student assessments. Clin Oral Investig. 2024;28. 10.1007/s00784-024-05968-w. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 2.Roldan-Vasquez E, Mitri S, Bhasin S, et al. Reliability of artificial intelligence chatbot responses to frequently asked questions in breast surgical oncology. J Surg Oncol. 2024;130:188–203. 10.1002/jso.27715. [DOI] [PubMed] [Google Scholar]
- 3.OpenAI. Introducing ChatGPT. 2022. https://openai.com/blog/chatgpt. Accessed May 1st 2024.
- 4.Google. An important next step on our AI journey. 2023. https://blog.google/technology/ai/bard-google-ai-search-updates/. Accessed May 1st 2024.
- 5.ChatGPT. https://chatgpt.com/. Accessed 15 Jan 2026.
- 6.Google Gemini. https://gemini.google.com/app. Accessed 15 Jan 2026.
- 7.Temsah A, Alhasan K, Altamimi I, et al. DeepSeek in Healthcare: Revealing Opportunities and Steering Challenges of a New Open-Source Artificial Intelligence Frontier. Cureus. 2025;17. 10.7759/cureus.79221. [DOI] [PMC free article] [PubMed]
- 8.Arılı Öztürk E, Turan Gökduman C, Çanakçi BC. Evaluation of the performance of ChatGPT-4 and ChatGPT-4o as a learning tool in endodontics. Int Endod J. 2025. 10.1111/iej.14217. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 9.Ali R, Tang OY, Connolly ID, et al. Performance of ChatGPT, GPT-4, and Google Bard on a Neurosurgery Oral Boards Preparation Question Bank. Neurosurgery. 2023;93:1090–8. 10.1227/neu.0000000000002551. [DOI] [PubMed] [Google Scholar]
- 10.Suárez A, Adanero A, Díaz-Flores García V, et al. Using a Virtual Patient via an Artificial Intelligence Chatbot to Develop Dental Students’ Diagnostic Skills. Int J Environ Res Public Health. 2022;19. 10.3390/ijerph19148735. [DOI] [PMC free article] [PubMed]
- 11.Chau RCW, Thu KM, Yu OY, et al. Performance of Generative Artificial Intelligence in Dental Licensing Examinations. Int Dent J. 2024;74:616–21. 10.1016/j.identj.2023.12.007. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 12.Liu M, Okuhara T, Dai Z, et al. Evaluating the Effectiveness of advanced large language models in medical Knowledge: A Comparative study using Japanese national medical examination. Int J Med Inf. 2025;193. 10.1016/j.ijmedinf.2024.105673. [DOI] [PubMed]
- 13.Kung TH, Cheatham M, Medenilla A et al. Performance of ChatGPT on USMLE: Potential for AI-assisted medical education using large language models. PLOS Digit Heal. 2023;2. 10.1371/journal.pdig.0000198. [DOI] [PMC free article] [PubMed]
- 14.De Moor R, Hülsmann M, Kirkevang LL, et al. Undergraduate curriculum guidelines for endodontology. Int Endod J. 2013;46:1105–14. 10.1111/IEJ.12186. [DOI] [PubMed] [Google Scholar]
- 15.Torabinejad M. Walton RE Endodontics: principles and practice.Elsevier Health Sciences, Philadelphia, PA. 2009.
- 16.Hargreaves KM. Cohen S Cohen’s pathways of the pulp. Elsevier, Philadelphia, PA. 2010.
- 17.Yusoff MSB. ABC of Content Validation and Content Validity Index Calculation. Educ Med J. 2019;11:49–54. 10.21315/eimj2019.11.2.6. [DOI]
- 18.Hallquist E, Gupta I, Montalbano M, Loukas M. Applications of Artificial Intelligence in Medical Education: A Systematic Review. Cureus. 2025;17. 10.7759/cureus.79878. [DOI] [PMC free article] [PubMed]
- 19.Jaworski W, Dolata T, Sawina P, et al. Comparison of GPT-5 and GPT-4o in Solving the Polish Centre for Medical Examinations (CEM) Gastroenterology Examination. Cureus. 2026;18. 10.7759/cureus.102497. [DOI] [PMC free article] [PubMed]
- 20.Luitel P, Paudel S, Upadhya D, et al. Performance of ChatGPT-4 on the Nepalese Undergraduate Medical Licensing Examination: A Cross-Sectional Study. J Med Educ Curric Dev. 2025;12. 10.1177/23821205251384836. [DOI] [PMC free article] [PubMed]
- 21.Balu A, Prvulovic ST, Fernandez Perez C, et al. Evaluating the value of AI-generated questions for USMLE step 1 preparation: A study using ChatGPT-3.5. Med Teach. 2025;47:1645–53. 10.1080/0142159X.2025.2478872. [DOI] [PubMed] [Google Scholar]
- 22.Li L, Zhang W, Zhang K, et al. The role of generative AI tools in case-based learning and teaching evaluation of medical biochemistry. BMC Med Educ. 2025;25. 10.1186/s12909-025-07567-z. [DOI] [PMC free article] [PubMed]
- 23.Ros-Arlanzón P, Gutarra-Ávila R, Arrarte-Esteban V, et al. When AI models take the exam: large language models vs medical students on multiple-choice course exams. Med Educ Online. 2025;30. 10.1080/10872981.2025.2592430. [DOI] [PMC free article] [PubMed]
- 24.Lee HY, Kim J, Choi H, et al. Comparing AI chatbot simulation and peer role-play for OSCE preparation: a pilot randomized controlled trial. BMC Med Educ. 2025;25. 10.1186/s12909-025-08308-y. [DOI] [PMC free article] [PubMed]
- 25.Jalali P, Mohammad-Rahimi H, Wang FM, et al. Performance of 7 Artificial Intelligence Chatbots on Board-style Endodontic Questions. J Endod. 2025;51:1413–9. 10.1016/j.joen.2025.06.014. [DOI] [PubMed] [Google Scholar]
- 26.Zhang P, Wang J, Hu X, et al. Comparative performance of GPT-4, GPT-o3, GPT-5, Gemini-3-Flash, and DeepSeek-R1 in ophthalmology question answering. Front cell Dev Biol. 2026;14. 10.3389/fcell.2026.1744389. [DOI] [PMC free article] [PubMed]
- 27.Shirani M. Comparing the performance of ChatGPT 4o, DeepSeek R1, and Gemini 2 Pro in answering fixed prosthodontics questions over time. J Prosthet Dent. 2025;135. 10.1016/j.prosdent.2025.04.038. [DOI] [PubMed]
- 28.Gokkurt Yilmaz BN, Ozbey F, Yilmaz BE. Evaluation of the performance of different large language models on head and neck anatomy questions in the dentistry specialization exam in Turkey. Surg Radiol Anat. 2025;47. 10.1007/S00276-025-03723-8. [DOI] [PubMed]
- 29.Yilmaz BE, Gokkurt Yilmaz BN, Ozbey F. Artificial intelligence performance in answering multiple-choice oral pathology questions: a comparative analysis. BMC Oral Health. 2025;25. 10.1186/s12903-025-05926-2. [DOI] [PMC free article] [PubMed]
- 30.Azhari AA, Ahmed WM, Alhamadani A, et al. Assessing the Efficacy of Artificial Intelligence Platforms in Answering Dental Caries Multiple-Choice Questions: A Comparative Study of ChatGPT and Google Gemini Language Models. Dent J. 2026;14. 10.3390/dj14020072. [DOI] [PMC free article] [PubMed]
- 31.Haylaz E, Gumussoy I, Kalabalik F, et al. Evaluation of the performance of four different large language models (ChatGPT, DeepSeek, Copilot, and Gemini) in answering oral, and maxillofacial radiology questions: pilot study. BMC Oral Health. 2026;26. 10.1186/s12903-026-07790-0. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 32.Builoff V, Shanbhag A, Miller RJ, et al. Evaluating AI proficiency in nuclear cardiology: Large language models take on the board preparation exam. J Nucl Cardiol. 2025;45:102089. 10.1016/j.nuclcard.2024.102089. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 33.Tsai CY, Hsieh SJ, Huang HH et al. Performance of ChatGPT on the Taiwan urology board examination: insights into current strengths and shortcomings. World J Urol. 2024;421:42:250. 10.1007/s00345-024-04957-8. [DOI] [PubMed]
- 34.Ali R, Tang OY, Connolly ID, et al. Performance of ChatGPT and GPT-4 on Neurosurgery Written Board Examinations. Neurosurgery. 2023;93:1353–65. 10.1227/neu.0000000000002632. [DOI] [PubMed] [Google Scholar]
- 35.Kochanek K, Skarzynski H, Jedrzejczak WW. Accuracy and Repeatability of ChatGPT Based on a Set of Multiple-Choice Questions on Objective Tests of Hearing. Cureus. 2024;16. 10.7759/cureus.59857. [DOI] [PMC free article] [PubMed]
- 36.Builoff V, Shanbhag A, Miller RJ et al. Evaluating AI Proficiency in Nuclear Cardiology: Large Language Models take on the Board Preparation Exam. medRxiv 2024;07.16.24310297. 10.1101/2024.07.16.24310297. [DOI] [PMC free article] [PubMed]
- 37.Alshammari AF, Madfa AA, Anazi BA et al. Comparison of accuracy and consistency of AI Language models when answering standardised dental MCQs. BMC Med Educ 2025;251(25):1507. 10.1186/s12909-025-07624-7. [DOI] [PMC free article] [PubMed]
- 38.Freire Y, Santamaría Laorden A, Orejas Pérez J, et al. ChatGPT performance in prosthodontics: Assessment of accuracy and repeatability in answer generation. J Prosthet Dent. 2024;131:659.e1-659.e6 10.1016/j.prosdent.2024.01.018. [DOI] [PubMed]
- 39.Shen OY, Pratap JS, Li X, et al. How Does ChatGPT Use Source Information Compared With Google? A Text Network Analysis of Online Health Information. Clin Orthop Relat Res. 2024;482:578. 10.1097/CORR.0000000000002995. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 40.Sandmann S, Hegselmann S, Fujarski M, et al. Benchmark evaluation of DeepSeek large language models in clinical decision-making. Nat Med. 2025;31:2546. 10.1038/S41591-025-03727-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 41.Gonsalves C. On ChatGPT: what promise remains for multiple choice assessment? J Learn Dev High Educ. 2023. 10.47408/jldhe.vi27.1009. [DOI] [Google Scholar]
- 42.Sumbal A, Sumbal R, Amir A. Can ChatGPT-3.5 Pass a Medical Exam? A Systematic Review of ChatGPT’s Performance in Academic Testing. J Med Educ Curric Dev. 2024;11. 10.1177/23821205241238641. [DOI] [PMC free article] [PubMed]
- 43.Seth I, Sinkjær Kenney P, Bulloch G, et al. Artificial or Augmented Authorship? A Conversation with a Chatbot on Base of Thumb Arthritis. Plast Reconstr Surg Glob open. 2023;11:e4999. 10.1097/GOX.0000000000004999. [DOI] [PMC free article] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Data Availability Statement
Data is provided within the supplementary information files.
