Skip to main content
BMC Oral Health logoLink to BMC Oral Health
. 2026 Mar 3;26:943. doi: 10.1186/s12903-026-07996-2

Comparative evaluation of artificial intelligence language models on knowledge of dental antibiotic use

Neslihan Yılmaz Çırakoğlu 1,✉, Mustafa Doğan 2, Barış Can Duymaz 2
PMCID: PMC13231621  PMID: 41776530

Abstract

Background

Artificial intelligence–based large language models (LLMs) are increasingly being used across multiple disciplines. In healthcare, as well as in education, research, and information access, the quality, reliability, and accuracy of the responses generated by these models have become a growing concern. This study aimed to comparatively evaluate the ability of two prominent LLMs, ChatGPT-3.5 and Gemini-1.0, using the most advanced publicly accessible versions at the time of AI access (February 9, 2025), to inform both the public and dental professionals about antibiotic use in dentistry, and to assess the content quality, accuracy, and comprehensiveness of their responses.

Methods

A total of 36 questions—comprising 12 multiple-choice, 12 true/false, and 12 open-ended items—were developed and posed to both ChatGPT-3.5 and Gemini-1.0, which were the most advanced publicly accessible versions at the time of AI access (February 9, 2025). The responses were independently scored by four expert endodontist dentists on a scale of 1 to 5. Data analyses were performed using SPSS. In addition to descriptive statistics, the Wilcoxon signed-rank test was used to compare model scores for identical questions, the Kruskal–Wallis test was applied to examine score differences across question types, and the intraclass correlation coefficient (ICC) was calculated to assess inter-rater reliability.

Results

Overall, responses generated by ChatGPT-3.5 and Gemini-1.0 were evaluated, and Gemini’s responses received higher scores than those of ChatGPT. The mean score for ChatGPT was 4.08 ± 0.83 (median = 4.0, interquartile range [IQR] = 1.25), whereas the mean score for Gemini was 4.65 ± 0.58 (median = 5.0, IQR = 1.00), indicating an approximate 0.57-point difference in favor of Gemini. The difference between the two models was statistically significant based on the Wilcoxon signed-rank test (p < 0.05). When analyzed by question type, Gemini scored higher than ChatGPT across all categories, including multiple-choice, true/false, and open-ended questions.

Conclusions

While large language models (LLMs) and AI chatbots, specifically ChatGPT-3.5 and Gemini-1.0, demonstrate reasonable performance in providing basic knowledge about antibiotic use in dentistry, their reliability remains limited in areas directly influencing clinical decision-making—particularly those involving patient-specific contraindications and individualized treatment considerations. These findings should be interpreted in the context of the specific model versions evaluated in this study. As LLMs continue to evolve rapidly, future studies using updated models and real-world clinical validation are warranted to further clarify their role as decision-support tools rather than standalone authorities.

Supplementary Information

The online version contains supplementary material available at 10.1186/s12903-026-07996-2.

Keywords: Artificial intelligence, Antibiotics, Large language models

Introduction

Artificial intelligence (AI)–based conversational agents capable of interacting with users through natural language are software applications that can perform a wide range of tasks, including generating presentations, educational content, computer code, academic text, and healthcare-related services [1–3]. Large language models (LLMs) are neural networks trained on extensive text corpora sourced from Wikipedia, digitized books, academic publications, and online materials [4]. These systems interpret and generate human language using AI techniques such as natural language processing (NLP), machine learning (ML), and deep learning (DL) [5].

Chatbots can analyse user input, identify contextual cues, learn from previous interactions, and generate dynamic and contextually appropriate responses. They can answer questions, provide detailed explanations, identify inaccuracies, reject inappropriate content, and offer apologies when errors occur [6].

Historically, early chatbots developed between the 1960s and 1980s, including ELIZA and PARRY, were rule-based systems. In contrast, modern AI assistants—such as ChatGPT, Gemini, and Copilot—operate using advanced large language models with contextual reasoning and multimodal capabilities [7]. Over time, numerous new features have been integrated, and diverse AI applications have emerged.

The widely used ChatGPT system was originally built on the GPT-3 model (2020). Its first large-scale conversationally adapted version, GPT-3.5 (November 2022), reached millions of users and introduced improved conversational memory and contextual coherence [8–10]. The launch of GPT-4 marked the beginning of the ChatGPT Plus subscription service, whereas the GPT-4o (Omni) model, available shortly after the study’s data collection, offers more human-like conversational flow and generates safer, more accurate content [11, 12]. Therefore, the GPT-3.5 version was selected for this study as the most advanced publicly accessible model at the time of data collection (Late 2024 – Early 2025).

Another prominent chatbot is Gemini, developed by Google/DeepMind. Similar to ChatGPT, Gemini is a large language model and multimodal AI system capable of processing text, images, and audio inputs [13, 14]. It is freely accessible, supports unlimited queries, and enables interactive dialogue. While Gemini generally provides comprehensive and informative responses, ChatGPT models often produce more creative and coherent outputs. Both systems are trained using principles of interactive learning [15, 16].

AI-based chatbots are increasingly being integrated into healthcare services [17, 18]. In dentistry, these technologies support multiple domains, including patient communication (appointment scheduling, frequently asked questions, pre- and post-treatment guidance), education (clinical simulation, self-directed learning), clinical process management (clinical decision support tools, patient data management, emergency triage), and research [19]. These tools provide patients, practitioners, and students with rapid responses and continuous access to information. However, they also have limitations, including an inability to fully evaluate complex medical conditions, the risk of misinformation, and unresolved ethical and legal concerns [20].

Traditionally, patients have relied on search engines such as Google or Bing to obtain health-related information. However, the emergence of publicly accessible LLMs, including ChatGPT and Gemini, has introduced a new era in information-seeking behaviours, with many users—including patients—now directing their questions to these models [21]. Previous studies evaluating the performance of LLMs in dental licensing examinations and discipline-specific question sets have shown that chatbot accuracy and reliability vary considerably depending on question format, clinical complexity, and evaluation methodology [22–26].

Although similar studies have been conducted, no research to date has specifically examined the competence of LLMs regarding antibiotic use in dentistry [27]. Therefore, this study aims to comparatively evaluate the responses of ChatGPT-3.5 and Gemini-1.0 to both technical and patient-oriented questions on antibiotic use in dentistry, thereby assessing their current evidence-based capabilities. Using open-ended, multiple-choice, and true/false formats, the study further explores how accurately each model responds to different question types. These findings are intended to provide a baseline for future studies evaluating newer model versions and refined prompting strategies.

Methods

Question dataset

To ensure clinical relevance and comprehensiveness, a set of 36 questions was developed to mirror real-world inquiries commonly posed by patients, dentists, and dental students regarding antibiotic use in dentistry. To evaluate the models’ versatility across different testing formats, the dataset was equally stratified into three categories: open-ended, multiple-choice, and true/false (n = 12 for each). The questions are provided in the supplementary file for transparency and reproducibility.

The study utilized a cross-sectional design to comparatively evaluate two prominent large language models: ChatGPT (powered by the GPT-3.5 architecture) and Gemini (Google DeepMind). Data collection was conducted between late 2024 and early 2025, and the GPT-3.5 version was selected as the most advanced publicly accessible model during that period.

The question set was designed to encompass both technical and patient-oriented scenarios, allowing assessment of the models’ capacity to provide evidence-based and practically relevant responses. These findings also provide a baseline for future studies evaluating newer model versions and refined prompting strategies.

Response generation

Access to both AI systems was obtained on 9 February 2025, and a new account was created specifically for the study. To ensure procedural consistency, all questions were entered by a single researcher. The versions assessed were ChatGPT (GPT-3.5) and Gemini, reflecting the most advanced publicly accessible models at the time of data collection (Late 2024 – Early 2025). Before interacting with the chatbots, all browsing history and cookies were cleared. Each question was entered in a new chat window to minimise the influence of previous interactions, and all outputs were recorded for subsequent analysis.

To ensure independence and guideline-appropriate responses, both systems received the following instruction:

"Please answer the following questions about antibiotic use in dentistry. Be specific and respond according to current standard guidelines."

Each question was asked once, in the same order, without any follow-up prompts, rephrasing, or clarification, and no questions were repeated by another researcher. This single-pass approach was chosen to reduce bias and reflect an initial user inquiry scenario, while acknowledging that real-world users may refine their prompts in practice.

All responses were categorised according to question type (open-ended, multiple-choice, true/false) and recorded for subsequent scoring and statistical analysis.

Performance evaluation

Responses generated by ChatGPT and Gemini were independently assessed by four expert endodontists using a 1–5 scoring scale. The evaluators were blinded to the identity of the chatbot generating each response to minimize bias and ensure impartiality. Assessment criteria, based on standards from the American Dental Association (ADA) and the Centers for Disease Control and Prevention (CDC), included accuracy, comprehensiveness, explanatory quality, and clarity.

Questions on antibiotic use in dentistry were categorised under two main categories—patient-oriented and technical questions—and distributed across open-ended, multiple-choice, and true/false formats, with 12 questions per format. The scoring protocol was developed to provide consistent evaluation across raters, and a detailed rubric is included in the supplementary material.

All questions were asked once in a single-pass format to reduce bias, reflecting the initial inquiry scenario, and the responses were directly linked to the corresponding questions in the analysis. The questions and corresponding responses from both models are presented in Appendix 1.

Global Quality Scale (GQS) [23–25]

Responses were evaluated using the 5-point Global Quality Scale (GQS), which assesses the overall quality, information flow, and practical usefulness of the responses. Scoring was performed independently by four expert raters. The scale is defined as follows:

  • Score 1: Low quality; poor information flow; insufficient content.

  • Score 2: Generally poor quality; some information presented, but many key aspects missing.

  • Score 3: Moderate quality; important points adequately addressed; somewhat useful for patient education and learning purposes.

  • Score 4: Good quality; most relevant information included; useful for patients, practitioners, and educational contexts.

  • Score 5: Excellent quality and flow; highly informative; maximally useful for patient education, practitioner guidance, and learning.

Raters were trained using example responses and a detailed scoring rubric (provided in the supplementary material) to promote consistency and minimize inter-rater variability. The GQS scores were subsequently used in combination with question-type stratification (open-ended, multiple-choice, true/false) for statistical analysis.

Statistical analysis

Statistical analyses were performed using IBM SPSS Statistics software (version 29.0; IBM Corp., Armonk, NY, USA). Descriptive statistics, including means, standard deviations (SDs), medians, and interquartile ranges (IQRs), were calculated for all variables to provide a comprehensive overview of the data distribution.

The non-parametric Wilcoxon signed-rank test was employed to compare the scores of the two models for identical questions, while the Kruskal–Wallis test was used to examine differences across the three question formats (open-ended, multiple-choice, and true/false). Post-hoc pairwise comparisons with Bonferroni correction were performed when applicable to account for multiple testing.

Inter-rater reliability among the four evaluators was assessed using the Intraclass Correlation Coefficient (ICC). Given the moderate ICC values observed, caution was applied in interpreting small differences between models; low ICC may reflect inherent subjectivity in scoring and variability among raters.

Although formal power analysis was not conducted prior to the study, the sample size (36 questions × 2 models × 4 raters) provides sufficient data points for preliminary comparative assessment. All statistical tests were two-sided, and significance was set at P < 0.05.

Data were analyzed separately for patient-oriented versus technical questions to explore potential differences in model performance across content type.

Results

Expert agreement analysis

The Intraclass Correlation Coefficient (ICC) for ChatGPT was calculated as 0.314 (95% CI: 0.151–0.503; p < 0.001). Although this falls within the range of “poor agreement” according to standard interpretive guidelines, the statistical significance suggests a measurable degree of concordance among the evaluators regarding ChatGPT’s responses.

In contrast, Gemini yielded an ICC of 0.001 (95% CI: − 0.110–0.163; p > 0.05). The negative lower bound of the confidence interval and lack of statistical significance indicate no meaningful agreement among the expert raters. This finding suggests that, although Gemini generally achieved higher mean performance scores (see Section 3.2), the evaluators did not demonstrate a systematic consensus on specific questions.

The moderate-to-poor ICC values likely reflect inherent subjectivity in scoring complex AI-generated responses, variability in expert interpretation, and the absence of repeated scoring or calibration sessions. Despite these limitations, the ICC analysis provides insight into inter-rater alignment and highlights the need for standardized scoring protocols in future studies.

Overall, these findings reveal that ChatGPT displayed a weak yet statistically significant alignment with expert evaluations, whereas Gemini showed no consistent inter-rater reliability. These results should be interpreted with caution and in the context of the study’s exploratory design and single-pass evaluation approach.

Table 1 presents the inter-rater reliability analysis, illustrating the correlation among the scores assigned by the four evaluators. This assessment was conducted to determine which AI system demonstrated a closer alignment with expert consensus.

Table 1.

Inter-rater reliability

ICC(%95 CI) p
ChatGPT 0.314 (0.151–0.503) < 0.001
Gemini 0.001 (-0.110-0.163) > 0,05

Interpretation: ICC values were interpreted as follows: poor (< 0.5), moderate (0.5–0.75), good (0.75–0.90), and excellent (> 0.90)

Table 2 presents a summary of the scores assigned to ChatGPT and Gemini based on the predefined evaluation criteria. Gemini demonstrated mean and median values predominantly ranging from 4.5 to 5.0, accompanied by low standard deviations (0–0.5), indicating a high degree of consistency in its responses. In contrast, ChatGPT’s scores ranged between 3.0 and 4.5, with higher standard deviations (0.5–1.0), reflecting greater variability in evaluator ratings. The prevalence of median scores of 5 for Gemini suggests a higher level of expert satisfaction regarding the quality of its outputs.

Table 2.

Descriptive statistics of model scores

ChatGPT Gemini
Mean ± SD Median Mean ± SD Median
1 4 ± 0.81 4 5 ± 0 5
2 4.25 ± 0.5 4 5 ± 0 5
3 3.75 ± 0.95 3.5 4.5 ± 0.57 4.5
4 3.25 ± 0.95 3.5 4.75 ± 0.5 5
5 3.75 ± 0.95 3.5 4.5 ± 0.57 4.5
6 4 ± 1.15 4 4.5 ± 1 5
7 3.75 ± 0.5 4 4.5 ± 0.57 4.5
8 4.5 ± 0.57 4.5 4.75 ± 0.5 5
9 4 ± 0.81 4 5 ± 0 5
10 4 ± 0.81 4 4.75 ± 0.5 5
11 4 ± 0.81 4 4.75 ± 0.5 5
12 4.25 ± 0.5 4 5 ± 0 5
13 4.25 ± 0.5 4 4.5 ± 1 5
14 4 ± 0.81 4 4.5 ± 0.57 4.5
15 4.25 ± 0.95 4.5 4.75 ± 0.5 5
16 4.75 ± 0.5 5 4.75 ± 0.5 5
17 4.5 ± 0.57 4.5 4.25 ± 0.5 4
18 4.75 ± 0.5 5 4.75 ± 0.5 5
19 4.5 ± 1 5 4.75 ± 0.5 5
20 4.5 ± 0.57 4.5 4.25 ± 0.95 4.5
21 4 ± 1.15 4 4.5 ± 1 5
22 4.25 ± 0.95 4.5 4.25 ± 0.5 4
23 4 ± 0 4 4.5 ± 0.57 4.5
24 4.5 ± 0.57 4.5 4.75 ± 0.5 5
25 4.75 ± 0.5 5 5 ± 0 5
26 4.25 ± 0.95 4.5 5 ± 0 5
27 4.25 ± 0.95 4.5 4.5 ± 1 5
28 4.25 ± 0.5 4 5 ± 0 5
29 4.5 ± 0.57 4.5 4.25 ± 0.5 4
30 4.25 ± 0.5 4 4.5 ± 0.57 4.5
31 3.75 ± 1.5 4 5 ± 0 5
32 4.5 ± 0.57 4.5 4.75 ± 0.5 5
33 4 ± 1.15 4 4.25 ± 1.5 5
34 2.75 ± 0.5 3 4.75 ± 0.5 5
35 3.25 ± 0.5 3 4.25 ± 0.5 4
36 2.75 ± 0.5 3 4.75 ± 0.5 5

It should be noted that the moderate-to-low ICC values reported in Section "Expert agreement analysis" indicate some variability among raters, which may influence the interpretation of these descriptive statistics. Nonetheless, these findings provide a preliminary comparative evaluation of the models’ response quality.

Table 3 presents statistical comparisons of ChatGPT and Gemini responses across multiple-choice (MCQ), true/false (T/F), and open-ended questions.

Table 3.

Model performance by question type

ChatGPT Gemini
Mean ± SD Median(IQR) Mean ± SD Median(IQR) Test statistics p (Wilcoxon test)
Section
 MCQ 3.89 ± 0.97 4(2) 4,62 ± 0,67 5(1) -3.892 < 0.001
 T/F 4.04 ± 0.79 4(1.250) 4,7 ± 0,54 5(1) -4.056 < 0.001
 Open-ended 4.31 ± 0.65 4(1) 4,62 ± 0,53 5(1) -2.435 0.007
Test statistics 7.92 1.04
p (Kruskal–Wallis Test) 0.019
  • Multiple-choice questions: Gemini performed significantly better than ChatGPT (Wilcoxon signed-rank test: Z = − 3.892, p < 0.001).

  • True/false questions: Gemini significantly outperformed ChatGPT (Wilcoxon signed-rank test: Z = − 4.056, p < 0.001).

  • Open-ended questions: Gemini again demonstrated superior performance (Wilcoxon signed-rank test: Z = − 2.435, p = 0.007).

Across all formats, Gemini demonstrated statistically significant superiority over ChatGPT. These results should be interpreted in the context of single-pass questioning, model versions (ChatGPT-3.5 and Gemini-1.0), and exploratory study design.

Table 4 reports the overall performance of both LLMs, independent of question type. ChatGPT had a mean score of 4.08 ± 0.83, with a median of 4.0 (IQR: 1.25). In comparison, Gemini achieved a higher mean score of 4.65 ± 0.58 and a median of 5.0 (IQR: 1.00), demonstrating lower variability.

Table 4.

Overall performance comparison

ChatGPT Gemini
Mean ± SD Median(IQR) Mean ± SD Median(IQR) Test statistics p (Wilcoxon Test)
4.08 ± 0.83 4(2–5) 4,65 ± 0,58 5(2–5) -6.031 < 0.001
p (Mann–Whitney U Test) < 0.001
U Value (Mann Whitney U Test) 6.154

The difference between the models was statistically significant according to both the Mann–Whitney U test (U = 6,154, p < 0.001) and the Wilcoxon signed-rank test (Z = − 6.031, p < 0.001). These findings suggest that, under the study conditions, Gemini consistently outperformed ChatGPT in terms of overall response quality and consistency.

However, it is important to consider that these results reflect the models available at the time of data collection and the single-pass question design. Future studies with newer model versions and iterative querying may provide additional insights.

Discussion

Dentistry is a discipline characterised by the rapid adoption of emerging technologies, with artificial intelligence (AI) increasingly incorporated into diagnostic and clinical workflows. AI systems have been applied in areas such as the detection of dental caries, periapical pathosis, and root fractures; orthognathic surgery planning; assessment of alveolar and peri-implant bone loss; and anatomical interpretation, among other applications [28–32]. Beyond enhancing diagnostic accuracy and operational efficiency, AI contributes to more personalised and patient-centred care by enabling tailored responses to individual clinical presentations [33]. In addition, AI-driven innovations hold significant potential to improve overall health outcomes and quality of life. Nonetheless, these benefits must be balanced against ethical concerns, particularly those relating to privacy, data security, algorithmic fairness, and transparency, which remain central to responsible AI integration in healthcare [34]. As the technological landscape evolves, ensuring the ethical and trustworthy deployment of AI will be essential for its sustained acceptance in dentistry and broader clinical practice [35].

Despite the advantages associated with AI, the accuracy and reliability of large language models (LLMs) require rigorous evaluation. Publicly accessible chatbots such as ChatGPT and Gemini are increasingly used for health-related inquiries, making the validation of their outputs particularly important [36]. Previous research has highlighted the risk of internal inconsistencies within model-generated content. For instance, GPT-4 and Claude 3 Opus correctly reported that clindamycin is no longer recommended for infective endocarditis prophylaxis under the 2021 AHA guidelines; however, contradictory statements were subsequently produced. Such inconsistencies have the potential to mislead both clinicians and patients, underscoring the need for caution when utilizing LLMs in clinical decision-making [37].

Earlier studies have shown mixed performance among LLMs in dentistry. Shrivastava et al. [38] assessed ChatGPT’s accuracy across 243 questions from multiple dental specialties and concluded that open-ended questions generated the highest accuracy, followed by binary and multiple-choice formats, suggesting potential applications in dental education and clinical reasoning. Similarly, Ngo et al. [39] reported that only 32% of pathology-related multiple-choice questions answered by ChatGPT-3.5 were accurate, indicating substantial limitations in structured question formats. Similarly, Rokhshad et al. demonstrated that chatbot performance in dentistry is highly dependent on clinical context and question complexity, and may differ substantially from clinician responses, underscoring the need for cautious interpretation of LLM-generated outputs [26].

The methodology used in the present study aligns with previous research protocols [40, 41], whereby each question was posed only once without follow-up clarification, avoiding unintended bias introduced by iterative prompting [42]. This single-pass questioning approach reflects real-world user interactions but may limit performance optimisation achievable through iterative clarification. Our aim was to evaluate the accuracy and quality of ChatGPT and Gemini’s responses to questions concerning antibiotic use in dentistry, a topic of ongoing clinical and public health relevance.

Overall, Gemini demonstrated superior performance, particularly in multiple-choice and true/false formats. The performance gap between the models narrowed for open-ended questions, indicating that ChatGPT may exhibit greater consistency in tasks requiring reasoning or elaboration. Nevertheless, both models achieved relatively high mean scores, suggesting that LLMs are capable of generating academically valuable content when responding to general inquiries.

Recent studies further contextualise these findings. Bayraktar and Göllü reported that ChatGPT provided lower accuracy levels than current guidelines when addressing antibiotic use in paediatric dental cases, whereas Microsoft Copilot demonstrated a level of accuracy comparable to that of paediatric dentists and dental students [43]. Additionally, Rewthamrongsris et al. evaluated seven contemporary LLMs—including GPT-3.5 Turbo, GPT-4o, Claude 3 models, and Gemini 1.5 variants—regarding antibiotic prophylaxis recommendations for dental procedures. Their results indicated that advanced models such as GPT-4o may serve as useful “assistive tools” in specific decision-support contexts [37]. Complementing this, Antonie’s review of AI chatbots in antibiotic stewardship suggested that although LLMs may support appropriate antibiotic use and enhance patient outcomes, limitations persist, including variability in handling clinical nuance, susceptibility to unsafe recommendations, algorithmic biases, privacy risks, and insufficient clinical validation [44].

A notable finding in the present study was the low inter-rater reliability observed among evaluators, particularly regarding Gemini. Gemini frequently received high scores (e.g., median values often reported as 5), resulting in restricted variance (ceiling effect) which contributed to a markedly lower Intraclass Correlation Coefficient (ICC). This suggests that, while Gemini generally achieved high performance scores, expert raters did not consistently agree on individual responses. Accordingly, future research should employ more explicit scoring rubrics, incorporate calibration exercises among evaluators, and explore iterative questioning approaches to improve standardisation, reproducibility, and alignment with real-world clinical use.

Importantly, the current study evaluated the models available at the time of data collection (ChatGPT-3.5 and Gemini-1.0). Subsequent model updates—such as GPT-4o, GPT-5, and Gemini 3.0—may address some of the observed limitations. Our study thus provides a baseline assessment, which can be complemented by future work using newer versions and expanded datasets.

Conclusion

In conclusion, while large language models (LLMs) and AI chatbots demonstrate reasonable performance in providing basic knowledge about antibiotic use, their reliability remains limited in areas directly influencing treatment decisions—particularly those involving patient-specific contraindications. Previous studies have emphasized that although these models can offer generally accurate information, they tend to perform poorly in addressing nuanced aspects such as dosage, duration, and special conditions (e.g., pregnancy, allergy, pediatrics) [43]. Review studies further highlight that for LLMs to be safely integrated into comprehensive clinical decision-support systems—beyond antibiotic stewardship—they must undergo rigorous clinical validation, localisation, and continuous monitoring of their real-world impact [44]. Similarly, Rokhshad et al. reported that, despite promising performance, chatbot-generated responses may lack the consistency and clinical reliability required to replace expert judgement, reinforcing the need for cautious interpretation in patient-facing contexts [26].

The results of this study demonstrate that Gemini generally achieved higher scores than ChatGPT, particularly excelling in multiple-choice and true/false question formats. However, the observed low inter-rater reliability limits the robustness of these findings, and these results reflect the performance of the models available at the time of data collection (ChatGPT-3.5 and Gemini-1.0). This underscores the importance of standardising evaluation procedures and employing explicit scoring rubrics in future research.

Importantly, future studies using newer model versions (e.g., GPT-4o, GPT-5, Gemini 3.0), iterative prompting, and expanded datasets are warranted to validate and extend these findings. Such work will further clarify the potential of LLMs to support clinical decision-making and educational applications while addressing current limitations.

As the use of AI-based models expands in education and research, it becomes essential to rigorously assess not only model performance but also the methodologies used to evaluate their outputs, thereby ensuring transparency, reproducibility, and safe integration into clinical practice.

Supplementary Information

Supplementary Material 1. (18.9KB, docx)

Acknowledgements

We would like to express our gratitude to TUBITAK for supporting this research.

Abbreviations

ADA

American Dental Association

AHA

American Heart Association

AI

Artificial Intelligence

CDC

Centers for Disease Control and Prevention

CI

Confidence Interval

DL

Deep Learning

GQS

Global Quality Scale

GPT

Generative Pre-trained Transformer

ICC

Intraclass Correlation Coefficient

IQR

Interquartile Range

LLM

Large Language Model

ML

Machine Learning

NLP

Natural Language Processing

SD

Standard Deviation

SPSS

Statistical Package for the Social Sciences

Authors’ contributions

NYÇ contributed to the study concept and design, collected data, and was involved in drafting and critically revising the manuscript. BD contributed to data analysis and interpretation, and participated in drafting the manuscript. MD performed all asking the questions to language models and participated in drafting the manuscript. All authors have approved the final manuscript and agreed to accountable for all aspects of the work.

Funding

This research is supported by the Scientific and Technological Research Council of Turkey (TUBITAK) institution.

Data availability

The datasets used and/or analysed during the current study are available from the corresponding author on reasonable request.

Declarations

Ethics approval and consent to participate

Not applicable.

Consent for publication

Not applicable.

Competing interests

The authors declare no competing interests.

Footnotes

Publisher’s note

Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.

References

  • 1.Aggarwal A, Tam CC, Wu D, Li X, Qiao S. Artificial Intelligence- Based Chatbots for Promoting Health Behavioral Changes: Systematic Review. J Med Internet Res. 2023;25:e40789. 10.2196/40789. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 2.Wailthare S, Gaikwad T, Khadse K, Dubey P. Artificial In telligence Based Chat- Bot. Artif Intell. 2018;5:2305–6. 10.22214//ijraset.2018.4393.
  • 3.Krishnan C, Gupta A, Gupta A, Singh G. Impact of Artificial Intelligence- Based Chatbots on Customer Engagement and Business Growth. in Deep Learning for Social Media Data Analytics. Springer International Publishing; 2022. pp. 195–210. 10.1007/978-3-031-10869-3_11.
  • 4.Wikipedia contributors. (n.d.). Large language model. In Wikipedia, The Free Encyclopedia. Retrieved 12 Oct 2025 from https://en.wikipedia.org/wiki/Large_language_model.
  • 5.Otter DW, Medina JR, Kalita JK. A survey of the usages of deep learning in natural language processing. arXiv. 2018. https://arxiv.org/abs/1807.10854. [DOI] [PubMed]
  • 6.Harland H, Senaratne H, Dazeley R, Vamplew P, Cruz F. A critical review of apology in AI systems. arXiv. 2025. https://arxiv.org/abs/2412.15787.
  • 7.Weizenbaum J. ELIZA—a computer program for the study of natural language communication between man and machine. Commun ACM. 1966;9(1):36–45. 10.1145/365153.365168.
  • 8.Brown TB, Mann B, Ryder N, Subbiah M, Kaplan J, Dhariwal P, et al. Language models are few-shot learners. arXiv [Preprint]. 2020. 10.48550/arXiv.2005.14165.
  • 9.OpenAI. (2022, Kasım 30). Introducing ChatGPT. OpenAI Blog. https://openai.com/blog/chatgpt.
  • 10.OpenAI. GPT-4 technical report. arXiv [Preprint]. 2023. 10.48550/arXiv.2303.08774.
  • 11.Brown TB, Mann B, Ryder N, Subbiah M, Kaplan J, Dhariwal P, Amodei D. Language models are few-shot learners. arXiv preprint arXiv: 2020;2005.14165. https://arxiv.org/abs/2005.14165.
  • 12.OpenAI. . Introducing ChatGPT. OpenAI Blog. 2022. https://openai.com/blog/chatgpt.
  • 13.Gemini. A Family of Highly Capable Multimodal Models. arXiv preprint. 2023. https://arxiv.org/abs/2312.11805.
  • 14.Google. (2024, Aralık). Introducing Gemini 2.0: our new AI model for the agentic era [Blog yazısı]. Google Blog. https://blog.google/technology/google-deepmind/google-gemini-ai-update-december-2024/.
  • 15.Google. (2023, Mayıs 10). Introducing Gemini: our most capable and general AI model. Google Blog. https://blog.google/technology/ai/google-gemini-ai/.
  • 16.Google Cloud. (n.d.). Multimodal AI | Google Cloud. Google Cloud. https://cloud.google.com/use-cases/multimodal-ai.
  • 17.Yılmaz E. Sağlık Hizmetlerinde Yapay Zekâ: Bir Araştırma (Yüksek Lisans Tezi, Sakarya Üniversitesi). 2022. https://acikerisim.sakarya.edu.tr/bitstream/handle/20.500.12619/102468/T11135.pdf.
  • 18.Frontiers in Digital Health. AI in healthcare: Navigating opportunities and challenges in digital communication. 2023. https://www.frontiersin.org/articles/10.3389/fdgth.2023.1291132/full. [DOI] [PMC free article] [PubMed]
  • 19.Alhaidry, A., … ve diğerleri. ChatGPT in Dentistry: A Comprehensive Review. 2023. https://www.ncbi.nlm.nih.gov/pmc/articles/PMC10230850/. [DOI] [PMC free article] [PubMed]
  • 20.Tudor Car L, Dhinagaran DA, Kyaw BM, Lim K. Ethical considerations of using ChatGPT in health care. JMIR Med Educ. 2023;9(1):e48098. https://pmc.ncbi.nlm.nih.gov/articles/PMC10457697/. [Google Scholar]
  • 21.Xie Z, Li X, Zhang Q. & ark. Laypeople’s use of and attitudes toward large language models for health information seeking. 2025. https://www.ncbi.nlm.nih.gov/pmc/articles/PMC11888097/.
  • 22.Chau RCW, Thu KM, Yu OY, Hsung RTC, Lo ECM, Lam WYH. Performance of generative artificial intelligence in dental licensing examinations. Int Dent J. 2024;74(3):616–21. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 23.Chau RCW, Thu KM, Yu OY, Hsung RT-C, Wang DCP, Man MWH, Wang JJ, Lam WYH. Evaluation of Chatbot Responses to Text-Based Multiple-Choice Questions in Prosthodontic and Restorative Dentistry. Dentistry J. 2025;13(7):279. 10.3390/dj13070279. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 24.Rokhshad R, Mohammad FD, Nomani M, Mohammad-Rahimi H, Schwendicke F. Chatbots for conducting systematic reviews in pediatric dentistry. J Dent. 2025;158:105733. [DOI] [PubMed] [Google Scholar]
  • 25.Rokhshad, R., Khoury, Z. H., Mohammad-Rahimi, H., Motie, P., Price, J. B., Tavares,T., … Sultan, A. S. Efficacy and empathy of AI chatbots in answering frequently asked questions on oral oncology. Oral surgery, oral medicine, oral pathology and oral radiology, 2025;139(6):719–728. 10.1016/j.oooo.2024.12.028. [DOI] [PubMed]
  • 26.Rokhshad R, Zhang P, Mohammad-Rahimi H, Pitchika V, Entezari N, Schwendicke F. Accuracy and consistency of chatbots versus clinicians for answering pediatric dentistry questions: A pilot study. J Dent. 2024;144:104938. [DOI] [PubMed] [Google Scholar]
  • 27.Liu M, Okuhara T, Huang W, Ogihara A, Nagao HS, Okada H, Kiuchi T. Innovation and application of Large Language Models (LLMs) in dentistry: a scoping review. 2024. https://pubmed.ncbi.nlm.nih.gov/39617779/. [DOI] [PMC free article] [PubMed]
  • 28.Ossowska A, Kusiak A, Świetlik D. Artificial intelligence in dentistry—narrative review. Int J Environ Res Public Health. 2022;19(6):3449. 10.3390/ijerph19063449. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 29.Pauwels R, Brasil DM, Yamasaki MC, Jacobs R, Bosmans H, Freitas DQ, Haiter-Neto F. Artificial intelligence for detection of periapical lesions on intraoral radiographs: Comparison between convolutional neural networks and human observers. Oral Surg Oral Med Oral Pathol Oral Radiol. 2021;131(5):610–6. 10.1016/j.oooo.2020.12.010. [DOI] [PubMed] [Google Scholar]
  • 30.Revilla-León M, Gómez-Polo M, Barmak AB, Inam W, Kan JYK, Kois JC, Akal O. Artificial intelligence models for diagnosing gingivitis and periodontal disease: A systematic review. J Prosthet Dent. 2023;130(6):816–24. 10.1016/j.prosdent.2022.05.006. [DOI] [PubMed] [Google Scholar]
  • 31.Murav AA, Guseynov NA, Tsay PA, Kibardin IA, Burnchev DV, Ivanov SS, Oborotistov NY, Matuta MA, Grachev NS, Larin SS. Artificial neural networks in dental and maxillofacial radiology: A review. Klinicheskaia Stomatologiia. 2020;3:72–80. [Google Scholar]
  • 32.Chau RCW, Li G-H, Tew IM, Thu KM, McGrath C, Lo W-L, Ling W-K, Hsung RT-C, Lam WYH. Accuracy of artificial intelligence-based photographic detection of gingivitis. Int Dent J. 2023;73(5):724–30. 10.1016/j.identj.2023.03.006. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 33.Al-Khafaji H, Al-Khafaji A, Al-Khafaji A. Generative AI in improving personalized patient care plans: Opportunities and barriers towards its wider adoption. Appl Sci. 2024;14(23):10899. 10.3390/app142310899. [Google Scholar]
  • 34.Pham T. Ethical and legal considerations in healthcare AI: Innovation and policy for safe and fair use. R Soc Open Sci. 2025;12(5):241873. 10.1098/rsos.241873. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 35.Sunar A, Taş Ç, Sır E, Polater H, Bağlıoğlu N. Diş hekimliğinde yapay zeka uygulamaları. J Kocaeli Health Technol Univ. 2024;2(3):41–57. https://dergipark.org.tr/tr/download/article-file/4106717. [Google Scholar]
  • 36.Balel Y. Can ChatGPT Be Used in Oral and Maxillofacial Sur gery? J Stomatology Oral Maxillofacial Surg. 2023;124:101471. 10.1016/j.jormas.2023.101471. [DOI] [PubMed] [Google Scholar]
  • 37.Rewthamrongsris P, Burapacheep J, Trachoo V, Porntaveetus T. Accuracy of large language models for infective endocarditis prophylaxis in dental procedures. Int Dent J. 2025;75(2):206–12. 10.1016/j.identj.2024.09.033. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 38.Shrivastava PK, Rai A, Injety RJ, Singh S, Jain A, Mahuli AV et al. Performance of ChatGPT in dentistry: a cross-sectional, multi-specialty and multi-centric study. Braz. J. Oral Sci. 2025;24(00):e254954. https://periodicos.sbu.unicamp.br/ojs/index.php/bjos/article/view/8674954. Cited 9 Feb 2026.
  • 39.Ngo A, Gupta S, Perrine O, Reddy R, Ershadi S, Remick D. ChatGPT 3.5 fails to write appropriate multiple choice practice exam questions. Acad Pathol. 2024;11(1):100099. 10.1016/j.acpath.2023.100099. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 40.Sybil D, Shrivastava P, Rai A et al. Performance of ChatGPT in Dentistry: Multi- Specialty and Multi- Centric Study. 2023. 10.21203/rs.3.rs-3247663/v1.
  • 41.Özbay Y. Evaluation of ChatGPT as a Multiple- Choice Question Generator in Dental Traumatology. Med Record. 2024;6:235–8. 10.37990/medr.1446396. [Google Scholar]
  • 42.Tokgöz Kaplan T, Cankar M. Evidence-Based Potential of Generative Artificial Intelligence Large Language Models on Dental Avulsion: ChatGPT Versus Gemini. Dent Traumatol. 2025;41(2):178–86. 10.1111/edt.12999. [DOI] [PubMed] [Google Scholar]
  • 43.Bayraktar C, Göllü S. Evaluation of Chatbots in Antibiotic Use in Children. Int Dent J. 2024;74:S153. 10.1016/j.identj.2024.07.1042. [Google Scholar]
  • 44.Antonie N, Iacobus, Gheorghe G, Ionescu VA, Tiucă L-C, Diaconu CC. The Role of ChatGPT and AI Chatbots in Optimizing Antibiotic Therapy: A Comprehensive Narrative Review. Antibiotics. 2025;14(1):60. 10.3390/antibiotics14010060. [DOI] [PMC free article] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Supplementary Material 1. (18.9KB, docx)

Data Availability Statement

The datasets used and/or analysed during the current study are available from the corresponding author on reasonable request.


Articles from BMC Oral Health are provided here courtesy of BMC

RESOURCES