Abstract
Introduction
Artificial intelligence (AI), particularly large language models (LLMs), is transforming healthcare education and clinical decision-making. While models like ChatGPT and Claude have demonstrated utility in medical contexts, their performance in dental diagnostics remains underexplored; additionally, the potential of emerging platforms, like Manus, is yet to be evaluated.
Objective
To compare the diagnostic accuracy and consistency of the ChatGPT, Claude, and Manus—using authentic, case-based dental scenarios.
Methods
A set of 117 multiple-choice questions based on validated clinical dental vignettes spanning various specialities was administered to each model under standardised conditions at two separate time points. Responses were scored against expert-validated answer keys. Inter-rater reliability was assessed using Cohen's kappa, and statistical comparisons were made using the chi-square, McNemar, and t-tests.
Results
Claude and Manus consistently outperformed ChatGPT across both testing phases. In the second round, Claude and Manus achieved a diagnostic accuracy of 92.3%, compared to ChatGPT's 76.9%. Claude and Manus also demonstrated higher intra-model consistency (Cohen's kappa = 0.714 and 0.782, respectively) than ChatGPT (kappa = 0.560). Although the numerical trends favoured Claude and Manus, pairwise differences in accuracy did not reach statistical significance.
Conclusion
Claude and Manus demonstrated numerically higher diagnostic performance and greater response stability compared with ChatGPT; however, these differences did not reach statistical significance and should therefore be interpreted cautiously. This variability across models highlights the need for larger-scale evaluations. These findings underscore the importance of considering both accuracy and consistency when selecting AI tools for integration into dental practice and curricula.
Keywords: artificial intelligence, large language models, clinical decision-making, manus AI, ChatGPT, claude, intra-model consistency, dental education
Introduction
The emergence of artificial intelligence (AI) in healthcare is transforming the way clinicians, educators, students, and stakeholders approach patient care, education, and policymaking (1). AI-driven tools have rapidly advanced in areas such as diagnostic imaging, personalised treatment planning, and clinical decision-making, allowing healthcare workers to make faster, more accurate decisions. Furthermore, AI offers new avenues for research and innovation, reducing workloads and enhancing the quality of care delivered across diverse specialities (2). This revolution also holds promise in dentistry, where oral health professionals can leverage AI to improve both clinical practice and educational experiences (3). The increasing clinical, scientific, and educational uses of ChatGPT in dentistry are highlighted in a recent thorough narrative review by Puleio et al., highlighting its potential to help academic writing, patient communication, diagnostic reasoning, and simulation-based learning (4).
In the context of oral health, AI is increasingly being used to support diagnostics, interpret imaging, predict treatment outcomes, and guide students through simulated learning experiences. Dental students and faculty can benefit immensely from AI-powered platforms that provide rapid, evidence-based feedback, allowing for more interactive, self-directed learning (3). Practising dentists may also use AI tools as chairside aids to enhance diagnostic accuracy for complex cases, improving efficiency and ensuring that patients receive high-quality care. Moreover, for health stakeholders and policymakers, the incorporation of AI into oral healthcare holds potential for improving resource allocation, quality assurance, and continuous professional development. These advances align with a broader global trend toward digital transformation across health disciplines (5). ChatGPT and other LLM-based systems already show significant value across various domains, as noted by Puleio et al., underscoring the necessity of systematic assessments of their effectiveness in dental contexts (4).
Within this evolving landscape, large language models (LLMs), such as ChatGPT (developed by OpenAI), Claude (developed by Anthropic), and Manus (developed by the startup Monica), have demonstrated a unique capacity to understand complex text-based queries, retrieve knowledge from vast datasets, and produce nuanced, context-aware responses (4). ChatGPT, a transformer-based LLM, is trained on a broad corpus and excels at conversational responses; Claude emphasises safety and alignment with user needs and is known for its concise and carefully reasoned output; and Manus, though less publicised, has been designed with robust integrative algorithms that synthesise data and potentially enhance accuracy in highly specialised domains like healthcare (6–8). Taken together, these models may significantly contribute to knowledge dissemination and skill development in dentistry.
Several preliminary studies have examined the effectiveness of AI language models in dental contexts. ChatGPT has demonstrated the ability to correctly answer a substantial proportion of dental knowledge-based multiple-choice questions (MCQs), suggesting its potential as a valuable educational support tool (9). Claude has shown competence in generating evidence-based diagnostic suggestions for oral pathologies, highlighting its usefulness in clinical training environments (10). However, to the best of our knowledge, no prior studies have evaluated the effectiveness of Manus in dental contexts, and its potential to support clinical decision-making remains largely unexplored.
Despite their promising outcomes, the need for direct, standardised comparisons among these models is underscored by variations in diagnostic accuracy, reasoning depth, and consistency over time. Direct comparative research examining how these models perform under identical clinical conditions remains lacking, especially when applied to realistic, case-based dental questions. Differences in training data, model architecture, and versioning may result in substantial variations in clinical utility. Thus, a direct, standardised comparison would generate valuable insight into which model most reliably supports dental professionals' and students' diagnostic processes.
Authentic case vignettes (structured to mimic actual patient presentations) offer an ideal means of assessing diagnostic reasoning and decision-making capabilities. Unlike purely theoretical or factual questions, case vignettes require the integration of diverse data points, including demographic information, chief complaints, history, clinical findings, and diagnostic imaging (11). This approach better simulates the complexity and nuance of real-world practice, allowing for a more robust evaluation of an AI model's capacity to provide accurate and clinically sound guidance. Furthermore, testing these models across different time points enables an assessment of consistency and knowledge retention, further approximating clinical reality, where stable performance over time is paramount.
The primary aim of this research was to evaluate and compare the diagnostic reasoning and clinical decision-making performance of the three AI models described above, ultimately providing evidence to inform their future integration into oral healthcare education and practice. Using a PICO framework, the present study aimed to address the following research question: “How do ChatGPT, Claude, and Manus compare in terms of diagnostic accuracy and consistency when responding to authentic clinical dental scenarios across two time points?”.
Methodology
Ethical considerations
The study did not include human participants or involve the use of identifiable personal data; therefore, informed consent was not required. Nonetheless, the research protocol was reviewed and approved by the Medical Ethical Committee of the College of Dentistry, University of Ha'il, Saudi Arabia (Approval No. H-2025-619). All methodological steps complied with established ethical principles for AI-related research, ensuring transparency, fairness, and adherence to academic integrity.
Study design
A structured, comparative, longitudinal study design was employed to evaluate the diagnostic reasoning and clinical decision-making abilities of three LLMs: ChatGPT (GPT-4-turbo, OpenAI, version released January 2024), Claude (Claude 3 Opus, Anthropic, version released March 2024), and Manus (latest version as of March 2025). The objective was to assess baseline accuracy and temporal consistency by testing the models on the same questions at two different time points.
Question development and validation
A total of 117 MCQs were constructed based on authentic clinical vignettes across all major dental specialities, including Oral and Maxillofacial Radiology, Oral Medicine and Oral Pathology, Periodontology, Endodontics, Operative and Restorative Dentistry, Prosthodontics, Paediatric Dentistry, and Orthodontics (Supplementary Material S1).
Each MCQ included patient demographics, chief complaints, history, clinical findings, diagnostic imaging findings, and any additional test results as required for accurate clinical decision-making.
All questions followed a single-best-answer format with 4 options. Correct answers were derived based on current, evidence-based clinical guidelines, such as those of the American Dental Association, European Federation of Periodontology, European Academy of Paediatric Dentistry, and American Association of Endodontists.
A panel of four experienced dental consultants (with ≥ 7 years of clinical and teaching experience) reviewed all questions and answer keys to ensure content validity and clinical relevance.
Each consultant independently reviewed all items, and any discrepancies in answer keys or question clarity were resolved through a consensus meeting. Questions were iteratively revised until full agreement was reached among all experts. To ensure adequate content coverage, consultants evaluated items both within and outside their respective specialties, confirming balanced representation across major dental domains. Although a formal inter-rater reliability coefficient was not calculated due to the consensus-based approach, unanimous agreement was required before finalizing each question and answer key.
AI interaction and data collection
In the initial evaluation phase (Phase 1), each of the 117 MCQs (12–14) was individually submitted to the interface of each AI model under identical, standardised conditions (Supplementary Material S1). To ensure fairness and reproducibility, identical prompts and phrasing were used across all models without providing any additional guidance, hints, or background context (Supplementary Material S2). The full responses generated by each AI model were then carefully documented as received, without modification, in a structured Microsoft Excel file. For each interaction, metadata, including the model's name and version, the query date, the question ID, the selected answer, and the model's complete textual explanation, were recorded for subsequent analysis.
Ten days after the first round of testing, the same 117 questions were re-administered to each AI model under the same strict protocol (Phase 2). This follow-up evaluation was conducted to assess whether the models' diagnostic reasoning and answer accuracy remained consistent over time, as well as to identify any variations in their responses or apparent knowledge retention. Again, all outputs were captured verbatim and logged in the data sheet alongside the corresponding metadata for comparison with the initial results.
Scoring and response evaluation
Two independent reviewers, both dental undergraduates, classified all responses as correct or incorrect by comparing them to the pre-validated answer keys. Undergraduate reviewers were selected because the task required only binary scoring of responses against consultant-validated answer keys, rather than clinical judgement or interpretation. The reviewers had completed the relevant clinical courses and were trained to ensure scoring accuracy and consistency. Inter-rater reliability was calculated using Cohen's kappa, with a target κ ≥ 0.80 indicating strong agreement.
Where provided, textual justifications were noted separately for subsequent qualitative analyses of clinical reasoning, although accuracy scoring was based solely on the selected multiple-choice option.
Data management and statistical analysis
All responses from both phases were systematically compiled into a structured Microsoft Excel file, with each row representing one question–model interaction. Relevant metadata, including the model's name and version, question ID, test phase, selected answer, correctness classification, and any explanatory text provided, were recorded. Prior to analysis, the data were cleaned and checked for completeness, and scoring accuracy was verified independently by two reviewers. Inter-rater reliability was calculated using Cohen's kappa, with κ ≥ 0.85 indicating strong agreement.
The finalised dataset was then exported to IBM SPSS Statistics (Version 28) for statistical analysis. Descriptive statistics, including mean accuracy scores, standard deviations, and frequency distributions, were calculated separately for each model across both time points. Differences in accuracy between models were assessed using chi-square tests, and McNemar's tests were performed to evaluate changes in performance across phases. Independent-samples t-tests were employed to explore variations in mean accuracy across dental specialities. Statistical significance was set at p < 0.05.
Results
Overall diagnostic accuracy (first and second measurements)
In the first measurement, Claude demonstrated the highest diagnostic accuracy, answering 107 out of 117 questions correctly (91.5%), followed closely by Manus with 106 correct responses (90.6%). ChatGPT exhibited comparatively poorer performance, correctly answering 87 questions (74.4%). Correspondingly, the number of incorrect answers showed the inverse, with Claude and Manus giving 10 (8.5%) and 11 (9.4%) incorrect responses, respectively, while ChatGPT provided 30 incorrect responses (25.6%); see Table 1.
Table 1.
Diagnostic accuracy of AI models in two measurements.
| First measurement | |||
|---|---|---|---|
| Variables | Manus | Chat GPT | Claude |
| Correct answer | 106 (90.6%) | 87 (74.4%) | 107 (91.5%) |
| Incorrect answer | 11 (9.4%) | 30 (25.6%) | 10 (8.5%) |
| Second measurement | |||
| Variables | Manus | Chat GPT | Claude |
| Correct answer | 108 (92.3%) | 90 (76.9%) | 108 (92.3%) |
| Incorrect answer | 9 (7.7%) | 27 (23.1%) | 9 (7.7%) |
In the second measurement, both Manus and Claude achieved the same performance, each correctly answering 108 questions (92.3%), while ChatGPT's performance improved slightly, with 90 correct responses (76.9%). The number of incorrect answers was reduced to nine (7.7%) for both Manus and Claude and 27 (23.1%) for ChatGPT. Figure 1 compares the diagnostic accuracy of Manus, ChatGPT, and Claude during the first and second rounds of testing.
Figure 1.
Diagnostic accuracy across two time points.
Consistency over time (agreement between first and second measurements)
To evaluate intra-model consistency, responses from the first and second measurements were cross-tabulated, and Cohen's kappa coefficient was calculated to quantify agreement levels. Manus showed high consistency between the two time points. Of the 106 initially correct responses, 105 (89.7%) remained correct in the second measurement, with only one response changing to incorrect (0.9%). Among the 11 initially incorrect responses, three became correct, and eight remained incorrect. The Cohen's kappa coefficient for Manus was 0.782, indicating substantial agreement (Table 2).
Table 2.
Intra-model consistency between first and second measurements.
| Manus | |||
|---|---|---|---|
| Measurement | Second Measurement | ||
| Correct | Incorrect | ||
| First measurement | Correct | 105 (89.7%) | 1 (0.9%) |
| Incorrect | 3 (2.6%) | 8 (6.8%) | |
| Measure of Agreement Kappa | .782 | ||
| ChatGPT | |||
| Measurement | Second Measurement | ||
| Correct | Incorrect | ||
| First measurement | Correct | 79 (67.5%) | 8 (6.8%) |
| Incorrect | 11 (9.4%) | 19 (16.2%) | |
| Measure of Agreement Kappa | .560 | ||
| Claude | |||
| Measurement | Second Measurement | ||
| Correct | Incorrect | ||
| First measurement | Correct | 105 (89.7%) | 2 (1.6%) |
| Incorrect | 3 (2.6%) | 7 (6.0%) | |
| Measure of Agreement Kappa | .714 | ||
Claude demonstrated similar stability. Of the 107 initially correct responses, 105 (89.7%) remained so, while two turned incorrect in the second round (1.6%). Of the 10 initially incorrect responses, three became correct, while seven remained incorrect. The kappa coefficient for Claude was 0.714, also reflecting substantial agreement. ChatGPT, in contrast, showed lower stability. Of the 87 correct responses in the first measurement, 79 (67.5%) were retained in the second, while eight turned incorrect (6.8%). Among the 30 initially incorrect responses, 11 were corrected and 19 remained incorrect. ChatGPT's kappa coefficient was 0.560, indicating moderate agreement and suggesting more variability in its responses across time. Figure 2 illustrates the consistency of each model between the two testing phases, as measured by Cohen's kappa.
Figure 2.
Inter-Measurement agreement (Cohen's Kappa).
Cross-model comparisons of accuracy
Pairwise comparisons between models were conducted to further examine differences in performance and their statistical significance. Comparing Manus to ChatGPT, Manus consistently outperformed ChatGPT in both measurements. In the second round, 85 of ChatGPT's correct responses matched those of Manus, while Manus provided only two incorrect responses where ChatGPT was correct. However, the difference in the correct response rate between the two was not statistically significant (p = 0.273), possibly due to the relatively small number of questions (Table 3).
Table 3.
Cross-comparison of accuracy between three AI language models, Manus, ChatGPT and Claude.
| AI Models | Answers | P Value | |||
|---|---|---|---|---|---|
| Both A&B correct | A correct & B incorrect | A incorrect & B correct | Both A&B incorrect | ||
| Manus (A) vs. ChatGPT (B) | 85 (78.7%) | 23 (19.7%) | 2 (1.7%) | 7 (6.0%) | .273 |
| Manus (A) vs. Claude (B) | 105 (89.7%) | 7 (6.0%) | 2 (1.7%) | 3 (2.6%) | .714 |
| ChatGPT (A) vs. Claude (B) | 88 (82.2%) | 2 (1.7%) | 19 (17.8%) | 8 (6.8%) | .352 |
In comparisons between Claude and ChatGPT, Claude showed higher accuracy. Of ChatGPT's 90 correct responses in the second round, 88 matched Claude's correct responses, while only two questions were answered correctly by ChatGPT but not by Claude. The difference between the two models was not statistically significant (p = 0.352), although Claude's overall accuracy and consistency trends were superior.
In contrast, Claude and Manus showed near-identical performance in the second measurement. Claude correctly answered 105 of the questions that Manus also answered correctly, with only minimal divergence between them (two and three incorrect responses, respectively). The p-value for this comparison was 0.714, supporting the absence of a statistically significant difference in diagnostic accuracy between the two models.
Discussion
The present study evaluated the diagnostic accuracy of three LLMs, Claude, Manus, and ChatGPT, in responding to dental examination questions. The results indicate that Claude and Manus consistently outperformed ChatGPT, reflecting both stronger reasoning mechanisms and potential domain-specific fine-tuning. These findings are consistent with prior evaluations of LLMs in medical and dental contexts (12–24). For example, ChatGPT (GPT-4) achieved 86.6% on the USMLE (21) and 88.6% on dental board-style questions (22), demonstrating strong general medical knowledge, yet it showed limitations in specialized dental domains compared with Claude and Manus. The superior performance of Claude and Manus may relate to their model architecture, structured clinical training datasets, or fine-tuning with exam-style questions, although the precise details of their training corpora remain proprietary (23, 24).
The differences in performance were most evident in the frequency of incorrect responses. Across both assessment rounds, ChatGPT produced nearly three times more errors than Claude or Manus, highlighting a notable discrepancy in reliability. This aligns with prior observations, such as those by Rao et al. (25), which noted that ChatGPT may present highly confident yet incorrect responses in clinical scenarios. ChatGPT exhibited a modest improvement in the second round (from 74.4% to 76.9%), suggesting some adaptability to question structures or increased familiarity with prompt formats, though this requires further evaluation. In contrast, Claude and Manus maintained relatively stable performance, with error rates ranging from 7.7%–9.4% across rounds, indicating more consistent accuracy.
Intra-model consistency was assessed using Cohen's kappa coefficient. Manus demonstrated the highest stability, retaining 89.7% of initially correct responses (κ = 0.782), followed by Claude (86.1%, κ = 0.714). ChatGPT showed moderate agreement (κ = 0.560), retaining only 67.5% of correct responses. These findings are consistent with prior studies reporting variability in repeated outputs from GPT-based models, particularly in complex or non-deterministic scenarios (26, 27). The higher stability of Claude and Manus may result from differences in architecture or targeted fine-tuning on domain-specific tasks, reducing stochastic output patterns. Minor fluctuations were observed even in these models (1.6%–1.9% of responses), indicating that small variations can occur despite high overall performance.
Pairwise comparisons were conducted to assess relative performance among the models. Although Claude and Manus consistently outperformed ChatGPT, the differences did not reach statistical significance, likely due to the modest sample size of test items. Manus and ChatGPT shared 85 correct responses in the second round, with ChatGPT outperforming Manus in only two instances. Claude and Manus exhibited nearly identical performance, correctly answering 105 of the same questions, with only minimal divergence (two questions unique to Claude and three unique to Manus; p = 0.714). These results suggest that both Claude and Manus achieved comparable proficiency in structured diagnostic tasks, supporting prior reports of parity or superiority of domain-optimized LLMs over general-purpose models in medical reasoning contexts (23, 24, 28, 29).
The intra-model stability also revealed patterns in how responses changed between assessment rounds. ChatGPT corrected only 11 of 30 initial errors while changing eight previously correct answers to incorrect, demonstrating greater variability than the other models. Claude and Manus showed far fewer such changes, further reflecting their more deterministic or fine-tuned response mechanisms. These observations are consistent with prior benchmarking studies and internal documentation suggesting that targeted fine-tuning and more deterministic architectures contribute to reduced output fluctuation (23, 24).
Overall, the study highlights consistent trends in performance and stability among the three LLMs. Claude and Manus consistently demonstrated higher accuracy, lower error rates, and greater intra-model agreement than ChatGPT, which exhibited moderate reliability and a higher incidence of incorrect responses. Pairwise analyses reinforce that the performance of Claude and Manus is closely aligned, while ChatGPT exhibits variability both in accuracy and response consistency across repeated assessments. These patterns are consistent with prior research examining large language models in medical and dental educational settings, reflecting the influence of model architecture, training data, and task-specific optimization on performance outcomes (21–31).
The lack of statistical significance in all pairwise comparisons underscores a key study limitation. The relatively small number of test questions (n = 117) may have reduced the power to detect meaningful differences. Although Manus and Claude demonstrated higher accuracy and consistency than ChatGPT across both testing phases, these differences were not statistically significant. As such, the observed performance trends should be interpreted as preliminary rather than conclusive. As such, while the observed numerical trends favour Claude and Manus over ChatGPT, larger-scale evaluations are warranted to substantiate these preliminary observations. Importantly, the convergence of Claude and Manus in diagnostic accuracy suggests a promising future for domain-specific LLMs tailored for healthcare. While ChatGPT performs respectably, it may be more prone to minor lapses in domain fidelity or contextual interpretation, particularly regarding specialised medical or dental knowledge.
Clinical implications
The findings of this study have several important clinical and educational implications. First, the high diagnostic accuracy and intra-model consistency demonstrated by Claude and Manus suggest that these AI models may serve as reliable adjunctive tools in dental education and clinical decision-making. Their consistent performance across repeated assessments indicates potential for use in formative assessments and board exam preparation and as real-time clinical support systems, particularly in settings with limited access to specialists. Second, while ChatGPT showed moderate accuracy and lower stability, it nevertheless performed at a level that could support general dental learning and preliminary case reviews. However, due to its greater response variability, it should be used with caution in contexts requiring high reliability. Third, the similarity in performance between Claude and Manus underscores that domain-specific fine-tuning and model design can substantially enhance clinical utility. The integration of such models into digital dental platforms, patient education materials, and case-based simulations could streamline diagnostic workflows and improve training outcomes for dental professionals. Finally, although none of the models should be relied upon as standalone diagnostic tools, their potential to support differential diagnosis, reinforce learning, and reduce clinician cognitive load makes them promising components in the future of AI-augmented dental care. Looking ahead, future developments of ChatGPT and similar LLMs may significantly expand their role in dentistry. As multimodal capabilities advance, these models are expected to integrate radiographic images, clinical photographs, and 3D scans to support more comprehensive diagnostic reasoning. Improved domain-specific fine-tuning, alignment with dental clinical guidelines, and incorporation of safety-focused guardrails may enhance accuracy and reliability in real patient scenarios. Additionally, integration with electronic dental records, personalised treatment planning systems, and chairside decision-support tools could position ChatGPT as an interactive educational and clinical companion. As these technologies evolve, establishing robust validation frameworks and regulatory standards will be essential to ensure safe, ethical, and effective deployment within dental education and practice.
This study has several limitations, including a small sample size and the exclusive use of MCQs, which may restrict generalisability and fail to capture the models' broader reasoning abilities. The static testing format also overlooks interactive elements crucial in clinical settings. Additionally, the questions may not fully represent all dental specialities or complexity levels, limiting applicability to broader educational and clinical contexts, as model performance might vary with tailored datasets or prompts. Lastly, the study did not assess the quality of the answer rationales, and the results reflect a single time point, highlighting the need for ongoing evaluations to track model development over time. Although a qualitative analysis of the models' reasoning was not conducted, the recorded justifications provide a foundation for future studies to explore AI reasoning and performance differences in depth. Additionally, the exclusive reliance on MCQs represents a limitation. MCQs were chosen because they offer a standardized, objective, and reproducible method for comparing accuracy across models and reflect the structure of most dental board examinations. However, this format does not fully capture multi-step clinical reasoning or the depth of diagnostic justification. Future research should incorporate open-ended clinical scenarios, interactive case simulations, or multimodal inputs to more comprehensively assess LLM reasoning.
Conclusions
This study compared the diagnostic accuracy and consistency of Manus, Claude, and ChatGPT in responding to case-based dental questions. Across both assessment rounds, Manus and Claude showed numerically higher accuracy and greater response stability compared with ChatGPT; however, these differences did not reach statistical significance and should therefore be interpreted cautiously. The observed trends indicate the potential of these models as supportive tools in dental education and clinical reasoning, but further research with larger datasets and more diverse assessment formats is required to confirm these preliminary findings. As LLMs continue to evolve, their integration into dental practice and education should be guided by careful evaluation of both accuracy and consistency.
Funding Statement
The author(s) declared that financial support was not received for this work and/or its publication.
Footnotes
Edited by: Jinyang Wu, Shanghai Jiao Tong University, China
Reviewed by: Hongyang Ma, Peking University Hospital of Stomatology, China
Francesco Puleio, University of Messina, Italy
Data availability statement
The datasets presented in this study can be found in online repositories. The names of the repository/repositories and accession number(s) can be found in the article/Supplementary Material.
Author contributions
AM: Methodology, Resources, Writing – original draft, Investigation, Software, Visualization, Conceptualization, Funding acquisition, Data curation, Validation, Formal analysis, Project administration, Writing – review & editing, Supervision. AA: Methodology, Data curation, Investigation, Resources, Conceptualization, Project administration, Visualization, Validation, Writing – original draft, Funding acquisition, Supervision, Formal analysis, Writing – review & editing, Software. BA: Visualization, Formal analysis, Investigation, Validation, Methodology, Software, Writing – review & editing. YA: Visualization, Formal analysis, Writing – review & editing, Software, Methodology, Investigation, Data curation. KA: Formal analysis, Methodology, Conceptualization, Validation, Funding acquisition, Writing – review & editing, Resources, Visualization, Investigation.
Conflict of interest
The author(s) declared that this work was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.
Generative AI statement
The author(s) declared that generative AI was not used in the creation of this manuscript.
Any alternative text (alt text) provided alongside figures in this article has been generated by Frontiers with the support of artificial intelligence and reasonable efforts have been made to ensure accuracy, including review by the authors wherever possible. If you identify any issues, please contact us.
Publisher's note
All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.
Supplementary material
The Supplementary Material for this article can be found online at: https://www.frontiersin.org/articles/10.3389/froh.2025.1686090/full#supplementary-material
References
- 1.Sissodia R, Dwivedi V. Multidisciplinary approaches to AI integration in education and healthcare systems. In: Sissodia R, Dwivedi V, editors. AI in Mental Health: Innovations, Challenges, and Collaborative Pathways. Hershey, PA: IGI Global Scientific Publishing; (2025). p. 435–64. [Google Scholar]
- 2.Kothinti RR. Deep learning in healthcare: transforming disease diagnosis, personalized treatment, and clinical decision-making through AI-driven innovations. World J Adv Res Rev. (2024) 24(2):2841–56. 10.30574/wjarr.2024.24.2.3435 [DOI] [Google Scholar]
- 3.Srivastava R, Tangade P, Priyadarshi S. Transforming public health dentistry: exploring the digital foothold for improved oral healthcare. Int Dent J Stud Res. (2023) 11(2):61–7. 10.18231/j.idjsr.2023.013 [DOI] [Google Scholar]
- 4.Puleio F, Lo Giudice G, Bellocchio AM, Boschetti CE, Lo Giudice R. Clinical, research, and educational applications of ChatGPT in dentistry: a narrative review. Appl Sci. (2024) 14(23):10802. 10.3390/app142310802 [DOI] [Google Scholar]
- 5.Batra AM, Reche A. A new era of dental care: harnessing artificial intelligence for better diagnosis and treatment. Cureus. (2023) 15(11):e49319. 10.7759/cureus.49319 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 6.Illangarathne P, Jayasinghe N, de Lima AD. A comprehensive review of transformer-based models: chatGPT and bard in focus. 2024 7th International Conference on Artificial Intelligence and Big Data (ICAIBD) (2024). p. 543–54 [Google Scholar]
- 7.Whitbeck K, Brown L, Abernathy S. Evaluating the Utility-Truthfulness Trade-off in Large Language Model Agents: A Comparative Study of ChatGPT, Gemini, and Claude.
- 8.Adinath DR, Smiju IS. Manus AI, Gemini, Grok AI, DeepSeek, and ChatGPT: A Comparative Analysis of Advancements in NLP (March 19, 2025).
- 9.Huang H, Zheng O, Wang D, Yin J, Wang Z, Ding S, et al. ChatGPT for shaping the future of dentistry: the potential of multi-modal large language model. Int J Oral Sci. (2023) 15(1):29. 10.1038/s41368-023-00239-y [DOI] [PMC free article] [PubMed] [Google Scholar]
- 10.Wu Y, Zhang Y, Xu M, Jinzhi C, Xue Y, Zheng Y. Effectiveness of various general large language models in clinical consensus and case analysis in dental implantology: a comparative study. BMC Med Inform Decis Mak. (2025) 25(1):147. 10.1186/s12911-025-02972-2 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 11.Heverly MA, Fitt DX, Newman FL. Constructing case vignettes for evaluating clinical judgment: an empirical model. Eval Program Plann. (1984) 7(1):45–55. 10.1016/0149-7189(84)90024-7 [DOI] [Google Scholar]
- 12.Abdulrab S, Abada H, Mashyakhy M, Mostafa N, Alhadainy H, Halboub E. Performance of 4 artificial intelligence chatbots in answering endodontic questions. J Endod. (2025) 51(1):13. 10.1016/j.joen.2025.01.002 [DOI] [PubMed] [Google Scholar]
- 13.Freire Y, Laorden AS, Pérez JO, Sánchez MG, García VD, Suárez A. ChatGPT performance in prosthodontics: assessment of accuracy and repeatability in answer generation. J Prosthet Dent. (2024) 131(4):659.e1. 10.1016/j.prosdent.2024.01.018 [DOI] [PubMed] [Google Scholar]
- 14.Yamaguchi S, Morishita M, Fukuda H, Muraoka K, Nakamura T, Yoshioka I, et al. Evaluating the efficacy of leading large language models in the Japanese national dental hygienist examination: a comparative analysis of ChatGPT, Bard, and Bing Chat. J Dent Sci. (2024) 19(4):2262–7. 10.1016/j.jds.2024.02.019 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 15.Ghaffari M, Zhu Y, Shrestha A. A review of advancements of artificial intelligence in dentistry. Dent Rev. (2024) 4(2):100081. 10.1016/j.dentre.2024.100081 [DOI] [Google Scholar]
- 16.Brozović J, Mikulić B, Tomas M, Juzbašić M, Blašković M. Assessing the performance of Bing Chat artificial intelligence: dental exams, clinical guidelines, and patients’ frequent questions. J Dent. (2024) 144:104927. 10.1016/j.jdent.2024.104927 [DOI] [PubMed] [Google Scholar]
- 17.Danesh A, Pazouki H, Danesh K, Danesh F, Danesh A. The performance of artificial intelligence language models in board-style dental knowledge assessment: a preliminary study on ChatGPT. J Am Dent Assoc. (2023) 154(11):970–4. 10.1016/j.adaj.2023.07.016 [DOI] [PubMed] [Google Scholar]
- 18.Danesh A, Pazouki H, Danesh F, Danesh A, Vardar-Sengul S. Artificial intelligence in dental education: ChatGPT's performance on the periodontic in-service examination. J Periodontol. (2024) 95(7):682–7. 10.1002/JPER.23-0514 [DOI] [PubMed] [Google Scholar]
- 19.Sabri H, Saleh MH, Hazrati P, Merchant K, Misch J, Kumar PS, et al. Performance of three artificial intelligence (AI)-based large language models in standardized testing; implications for AI-assisted dental education. J Periodontal Res. (2024) 59(1):42–52. 10.1111/jre.13203 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 20.Sismanoglu S, Capan BS. Performance of artificial intelligence on Turkish dental specialization exam: can ChatGPT-4.0 and Gemini advanced achieve comparable results to humans? BMC Med Educ. (2025) 25(1):214. 10.1186/s12909-024-06389-9 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 21.Kung TH, Cheatham M, Medenilla A, Sillos C, De Leon L, Elepaño C, et al. Performance of ChatGPT on USMLE: potential for AI-assisted medical education using large language models. PLoS Digital Health. (2023) 2(2):e0000198. 10.1371/journal.pdig.0000198 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 22.Aljindan FK, Al Qurashi AA, Albalawi IA, Alanazi AM, Aljuhani HA, Almutairi FF, et al. ChatGPT conquers the Saudi medical licensing exam: exploring the accuracy of artificial intelligence in medical knowledge assessment and implications for modern medical education. Cureus. (2023) 15(9):e45043. 10.7759/cureus.45043 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 23.Anthropic. Claude AI: Model Capabilities and Safety Optimizations. San Francisco, CA: Anthropic; (2024). Available online at: https://www.anthropic.com (Accessed May 15, 2024). [Google Scholar]
- 24.Manus AI. Technical Documentation and Clinical Benchmarking Report. Beijing: Manus AI, Inc; (2024). [Google Scholar]
- 25.Rao A, Kim J, Kamineni M, Pang M, Lie W, Dreyer KJ, et al. Evaluating GPT as an adjunct for radiologic decision making: GPT-4 versus GPT-3.5 in a breast imaging pilot. J Am Coll Radiol. (2023) 20(10):990–7. 10.1016/j.jacr.2023.05.003 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 26.Gilson A, Safranek C, Huang T, Socrates V, Chi L, Taylor RA, et al. How does ChatGPT perform on the medical licensing exams? The implications of large language models for medical education and knowledge assessment. MedRxiv. (2022):2022–12. [Google Scholar]
- 27.Huang X, Ruan W, Huang W, Jin G, Dong Y, Wu C, et al. A survey of safety and trustworthiness of large language models through the lens of verification and validation. Artif Intell Rev. (2024) 57(7):175. 10.1007/s10462-024-10824-0 [DOI] [Google Scholar]
- 28.Alhazmi N, Alshehri A, BaHammam F, Philip M, Nadeem M, Khanagar S. Can large language models serve as reliable tools for information in dentistry? A systematic review. Int Dent J. (2025) 75(4):100835. 10.1016/j.identj.2025.04.015 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 29.Wang X, Ye H, Zhang S, Yang M, Wang X. Evaluation of the performance of three large language models in clinical decision support: a comparative study based on actual cases. J Med Syst. (2025) 49(1):23. 10.1007/s10916-025-02152-9 [DOI] [PubMed] [Google Scholar]
- 30.Nori H, King N, McKinney SM, Carignan D, Horvitz E. Capabilities of gpt-4 on Medical Challenge Problems. arXiv preprint arXiv:2303.13375 (March 20, 2023).
- 31.Alshammari AF, Madfa AA, Anazi BA, Alenezi YE, Alkurdi KA. Comparison of accuracy and consistency of AI language models when answering standardised dental MCQs. BMC Med Educ. (2025) 25:1507. 10.1186/s12909-025-07624-7 [DOI] [PMC free article] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Data Availability Statement
The datasets presented in this study can be found in online repositories. The names of the repository/repositories and accession number(s) can be found in the article/Supplementary Material.


