Abstract
Background
This study aims to conduct a comparative evaluation of the accuracy, reliability, quality, and readability of chatbot-generated responses from five widely used artificial intelligence (AI) chatbots when addressing frequently asked questions (FAQs) on zirconia and pediatric zirconia crowns.
Methods
Twenty FAQs on zirconia crowns were derived from two Google searches (“frequently asked questions about zirconia” and “frequently asked questions about pediatric zirconia”). Five chatbots (ChatGPT-5, ChatGPT-4o, Gemini-2.5 Flash, DeepSeek-V3, and Microsoft Copilot) were queried independently, and responses were anonymized and evaluated. Accuracy was rated on a 5-point Likert scale, reliability using a modified DISCERN tool, quality with the Global Quality Scale (GQS), and readability using the Flesch Reading Ease Score (FRES). Statistical analyses included Mann−Whitney U and Kruskal−Wallis tests, with intraclass correlation coefficients (ICC) used for inter-rater reliability.
Results
Inter-rater agreement was strong (ICC: 0.78 − 0.98). Gemini achieved the highest scores in accuracy, quality, and reliability (p < 0.001), while ChatGPT-4o, ChatGPT-5, and DeepSeek demonstrated superior readability. Microsoft Copilot scored lowest across domains, particularly in reliability and readability. No significant differences emerged between prosthodontic and pediatric evaluations, except for higher GQS ratings for DeepSeek in pediatric dentistry (p = 0.035).
Conclusion
Gemini showed the highest accuracy, reliability, and quality, indicating its strong potential for clinician use in generating evidence-aligned information. ChatGPT-4o, ChatGPT-5, and DeepSeek offered more readable outputs suitable for explanations. Given the substantial between-platform variability, clinicians should critically appraise and, when necessary, adapt chatbot responses to ensure alignment with current evidence before recommending them to patients.
Keywords: Dental restorations, Esthetics, Patient information, Prosthodontics, Pediatric dentistry
Introduction
The rapid development of artificial intelligence (AI) in healthcare and dentistry has changed access to health-related information for clinicians, students, and patients [1]. AI-based chatbots, such as ChatGPT, Gemini, DeepSeek, and Microsoft Copilot, have become common methods of obtaining almost instantaneous conversational outcomes for complex questions [2]. These chatbots and digital assistants use large language models (LLMs) trained on large datasets to imitate human-style interaction to respond to patient prompts, assist with clinical decisions, and facilitate education [3].
Digitized health information tools can improve access to dental knowledge; however, previous studies have shown considerable variability in AI-generated outputs, particularly in terms of accuracy, reliability, quality, and readability [4–7]. Such variability increases the risk of misinformation. In dentistry, this is a significant concern because inaccurate information may negatively influence patient care and clinical decision-making. These concerns are relevant for zirconia crowns, as patients and parents increasingly seek information when making restorative decisions. Zirconia crowns are a rapidly growing treatment option in clinical practice. AI systems may oversimplify clinical indications, present zirconia crowns as universally suitable, or inadequately address material limitations. In pediatric dentistry, insufficient differentiation between adult and primary dentition may further lead to misleading explanations. These issues underscore the need to evaluate the accuracy, reliability, quality, and readability of AI-generated information on zirconia crowns [4–6].
Zirconia-based materials represent some of the most notable advances in restorative dentistry over the past decade, improvements in the esthetics of zirconia restorations have positioned them among the leading options compared to other restorative materials [8, 9]. Zirconia-based crowns are a popular restoration in adult prosthodontics as well as pediatric dentistry since adult patients and parents request them as an option outside of metal-based or composite restorations [10]. There are many advantages of zirconia restorations, including more esthetical properties, comparable biocompatibility, excellent mechanical properties, superb long-term durability, minimal plaque accumulation, and minimal risk of allergic reaction [9, 11, 12]. Recently, both adult patients and parents have increasingly sought information from their clinicians regarding zirconia and pediatric zirconia crowns, leading to a growing demand for guidance on their indications, a reliable restorative option [10, 13]. The demand is apparent as more prosthodontic patients and caregivers are obtaining their information from the internet and searching for answers from AI platforms to get real-time responses to their questions about dental restorations [14, 15]. The number of studies investigating the trustworthiness and quality of these responses generated through AI compared to human clinicians has been steadily increasing in recent years [15–19]. Multiple studies across medical and dental disciplines have shown that, although AI chatbots can provide health-related information, their outputs should be interpreted with caution. Chatbot-generated responses may be variable, occasionally inaccurate, lack appropriate source attribution, exhibit bias, or provide insufficient contextual depth [15, 17–21].
Currently, no studies have specifically evaluated how AI chatbots perform when responding to questions about zirconia crowns in both adult prosthodontics and pediatric dentistry contexts. This represents a critical gap, as AI-generated responses may lead to misconceptions such as inaccurate guidance on material selection, clinical indications, or procedural considerations. Zirconia crowns have distinct clinical implications in adults and children: in adult prosthodontics, considerations include long-term durability, esthetics, and functional load, whereas in pediatric dentistry, factors such as physiological tooth abrasion, growth, and patient/parent preference are particularly relevant. Given these differences, the reliability, accuracy, readability, and quality of AI-generated content may vary depending on the domain. Accordingly, this study adopts a comparative two-domain design, evaluating AI chatbot responses to frequently asked questions related to (1) prosthodontic zirconia crowns and (2) pediatric zirconia crowns. The present study aims to evaluate and compare the performance of five commonly used AI chatbots (ChatGPT-5, ChatGPT-4o, Gemini-2.5 Flash, DeepSeek-V3, and Microsoft Copilot) by analyzing their responses to frequently asked questions regarding zirconia crowns in adults and pediatric patients. The null hypothesis is that there is no statistically significant difference among the five AI chatbots (ChatGPT-5, ChatGPT-4o, Gemini-2.5 Flash, DeepSeek-V3, and Microsoft Copilot) in terms of accuracy, reliability, quality, and readability when responding to frequently asked questions about zirconia and pediatric zirconia crowns.
Materials and methods
Two independent Google searches were conducted using the queries “frequently asked questions about zirconia” and “frequently asked questions about pediatric zirconia”. Following protocols used in prior research [19, 22], the first 30 websites from each search were screened, yielding 205 zirconia-related and 196 pediatric zirconia-related questions. After eliminating duplicates, irrelevant entries, sponsored content, and non-English material, the 20 FAQs were selected (Fig. 1). The final set of 20 questions represented the most frequently recurring and consistently appearing patient inquiries across both searches. This number is aligned with previous methodological approaches in the literature evaluating chatbot performance in dentistry, and it ensures inclusion of the highest-frequency, clinically relevant questions while avoiding redundancy [23–25]. The selection process by two prosthodontists and two pediatric dentists also served as expert validation, confirming that the retained items accurately reflected the most common patient concerns regarding zirconia and pediatric zirconia crowns. As the study used only publicly available, non-identifiable data, ethics committee approval was not required.
Fig. 1.

Distribution of the 20 most frequently asked questions regarding zirconia crowns, categorized into prosthodontic and pediatric domains for evaluation
The performance of five AI-based chatbots (ChatGPT-5, ChatGPT-4o (OpenAI), Gemini-2.5 Flash (Google), DeepSeek-V3, and Microsoft Copilot) were assessed. As chatbot outputs may vary based on timing or repeated prompts, each question was posed once. To avoid potential bias or inter-platform influence, each chatbot was accessed via a separate user account and browser session.
Prosthetic content was evaluated by two prosthodontists (N.K.A., S.K.), while pediatric content was assessed by two pediatric dentists (F.S., P.C.). To ensure objectivity, all chatbot responses were exported, anonymized, and assigned random letter codes (A − E); all platform names and identifying information were removed. The anonymized question−answer sets were then compiled into a single PDF file and distributed to the evaluators for scoring, ensuring that all assessors were fully blinded to chatbot identity throughout the accuracy, reliability, quality, and readability assessments. Before initiating the full evaluation, all assessors participated in a calibration session using representative responses to standardize the scoring criteria. A pilot evaluation of five randomly chosen responses, selected by numerical randomization, was conducted by each assessor independently, with any discrepancies resolved by consensus. Prosthodontic and pediatric evaluators scored only the content specific to their respective specialties, and their assessments were not pooled. Instead, scores from each specialty were analyzed separately to preserve domain-specific evaluation, and cross-specialty comparisons were subsequently performed using the Mann−Whitney U test to identify any significant differences between the two groups.
Accuracy of responses was rated using a 5-point Likert scale (1: entirely incorrect or irrelevant; 5: entirely accurate and informative) [19, 26]. Reliability was measured using a modified version of the DISCERN instrument, covering objectives, relevance, source documentation, neutrality, date of publication, provision of additional resources, and acknowledgment of uncertainty [27, 28]. Responses were scored as follows: 1 (no), 2 − 4 (partial), and 5 (yes). Quality was assessed with the 5-point Global Quality Scale (GQS) (1: poor and incomplete; 5: comprehensive and excellent) [29]. Readability was evaluated using an online readability calculator, Character Calculator (charactercalculator.com), which performs automated syllable, word, and sentence counts to generate the Flesch Reading Ease Score (FRES). The FRES was calculated as follows:
206.835 − 1.015 × (total words / total sentences) − 84.6 × (total syllables / total words), where values range from 0 (very difficult) to 100 (very easy) [30]. Accuracy, reliability, quality, and readability scoring systems and score interpretations are shown in Fig. 2.
Fig. 2.

Accuracy (5-Point Likert Scale), reliability (Modified DISCERN Scale), quality (GQS), and readability (FRES) scoring systems and score interpretations
All statistical analyses were performed using IBM SPSS Statistics (version 26.0; IBM Corp., Armonk, NY, USA). Continuous variables were summarized as median and interquartile range (IQR). Comparisons between the prosthodontics and pediatric groups, given the non-normal distribution of data, were performed using the Mann−Whitney U test. Differences in chatbot performance across all four systems were analyzed with the Kruskal−Wallis test, followed by Bonferroni-adjusted pairwise comparisons when appropriate. Inter-rater reliability was quantified using the Intraclass Correlation Coefficient (ICC) for scores assigned by the two independent evaluators within each specialty. ICC results were reported with 95% confidence intervals (CI), with values > 0.75 considered indicative of strong agreement and > 0.90 reflecting excellent agreement [31]. A significance threshold of p < 0.05 was applied for all analyses.
Results
The inter-rater reliability analysis conducted for the evaluation of contents produced by five different artificial intelligence engines in terms of accuracy, reliability, and quality revealed ICC values ranging between 0.78 and 0.98 (Table 1).
Table 1.
Intraclass correlation coefficients (ICC) of evaluators
| Pediatric dentistry | Prosthodontics | |||
|---|---|---|---|---|
| ICC | (95% CI) | ICC | (95% CI) | |
| Accuracy | 0,94 | 0.88 − 0.97 | 0,95 | 0.91 − 0.98 |
| Reliability | 0,92 | 0.85 − 0.96 | 0,85 | 0.78 − 0.91 |
| Quality | 0,88 | 0.81 − 0.93 | 0,91 | 0.86 − 0.95 |
The responses provided by five different artificial intelligence engines (ChatGPT-4o, ChatGPT-5, Gemini-2.5Flash, DeepSeek, Copilot) to various questions regarding zirconia crowns were evaluated by Prosthodontics and Pediatric Dentistry specialists in terms of accuracy (Likert scale), quality (GQS), reliability (DISCERN), and readability (FRES), and the results are presented in Table 2.
Table 2.
Statistical comparison of chatbot responses based on Likert, modified DISCERN, GQS, and FRES scores
| ChatGPT-4o | ChatGPT-5 | DeepSeek | Gemini | Copilot | p | Effect Size | ||
|---|---|---|---|---|---|---|---|---|
| Prosthodontics | Accuracy (Likert) | 4 (4–4)ab | 4 (4–4)a | 4 (3–4)b | 4 (4–5)a | 3.5 (3–4)b | < 0.001 | 0.260 |
| Quality (GQS) | 4 (3.8-4)c | 5 (4–5)ab | 4 (4-4.3)c | 5 (4.8-5)a | 4 (4–5)bc | < 0.001 | 0.165 | |
| Reliability (DISCERN) | 19 (17-20.6)bc | 20 (19–21)b | 19 (18.8–20)bc | 27 (25–28)a | 18 (17–19)c | < 0.001 | 0.596 | |
| Readability (FRES) | 52 (45.5–61.7)a | 49.7 (44.1–59.7)a | 55.23 (47.18-61)a | 45.5 (34.2–52.8)ab | 36 (27.7–44.7)b | < 0.001 | 0.254 | |
| Pediatric Dentistry | Accuracy (Likert) | 4 (4–4)ab | 4 (4–4)a | 4 (3–4)b | 4 (4–5)ab | 3.5 (3–4)b | < 0.001 | 0.133 |
| Quality (GQS) | 4 (3.8-4)b | 5 (4–5)ab | 4 (4-4.3)a | 5 (4.8-5)a | 4 (4–5)ab | 0.008 | 0.175 | |
| Reliability (DISCERN) | 19 (17-22.5)b | 20 (19–21)b | 19 (18.8–20)b | 27 (25–28)a | 18 (17–19)b | < 0.001 | 0.628 | |
| Readability (FRES) | 52.02 (45.5–61.7)a | 49.67 (44.1–59.7)ab | 55.23 (47.2–61)a | 45.5 (34.2–52.8)b | 36 (27.7–44.7)c | < 0.001 | 0.489 |
Different superscript letters indicate statistically significant differences between groups according to Dunn post hoc test (p < 0.05) after significant Kruskal−Wallis test
IQR Interquartile range
According to the Likert scale results in the Prosthodontics group, Gemini and ChatGPT-5 demonstrated significantly higher accuracy scores than DeepSeek and Copilot (p < 0.001). In terms of GQS scores, Gemini and ChatGPT-5 received the highest quality scores, showing a significant difference from ChatGPT-4o and DeepSeek, which had the lowest scores (p < 0.001). In the DISCERN analysis, Gemini achieved a significantly higher reliability score compared to all other engines (p < 0.001), while the lowest score was observed for Copilot. Regarding Flesch readability scores, ChatGPT-4o, ChatGPT-5, and DeepSeek provided significantly higher readability than Copilot (p < 0.001).
In the Likert scale evaluations of the Pediatric Dentistry group, ChatGPT-5 exhibited a significantly higher accuracy score compared to DeepSeek and Copilot (p < 0.001). Regarding GQS scores, Gemini and DeepSeek achieved significantly higher scores than ChatGPT-4o (p = 0.008). In terms of DISCERN, Gemini provided a significantly higher reliability score compared to all other engines (p < 0.001). For Flesch readability assessments, ChatGPT-4o and DeepSeek demonstrated statistically higher readability scores than Gemini and Copilot (p < 0.001). Copilot showed a statistically significantly lower readability score than all other groups (p < 0.001).
When the responses of the artificial intelligence engines in the fields of Prosthodontics and Pediatric Dentistry were compared (Table 3), no statistically significant differences were observed in the overall scores assigned to the engines (p > 0.05). The only significant difference was identified for the DeepSeek engine in terms of overall quality evaluation (GQS) (p = 0.035). This indicates that the content generated by DeepSeek was rated as higher in quality by participants in the field of Pediatric Dentistry compared to those in Prosthodontics. For all other engines, no significant differences were found between the two fields in perceived usefulness (Likert scale), content reliability (DISCERN), or readability (FRES) scores (p > 0.05).
Table 3.
p-Values for the comparison of artificial intelligence engines’ responses in prosthodontics and pediatric dentistry (Mann−Whitney U test Results)
| Accuracy (Likert) | ChatGPT-4o | ChatGPT-5 | Gemini | DeepSeek | Copilot |
|---|---|---|---|---|---|
| 0.211 | 0.183 | 0.327 | 0.429 | 0.429 | |
| Quality (GQS) | 0.512 | 0.429 | 0.659 | 0.035* | 0.149 |
| Reliability (DISCERN) | 0.157 | 0.398 | 0.640 | 0.086 | 0.114 |
| Readability (FRES) | 0.114 | 0.201 | 0.369 | 0.355 | 0.841 |
Discussion
Although AI chatbots are increasingly used to obtain information about clinical procedures in dentistry, the accuracy, reliability, and readability of such public-facing information remain uncertain. The healthcare information provided by these chatbots may pose challenges, including inaccuracies, biases, overconfidence, limited scope, and ethical concerns [25]. Systematic investigation of the reliability and accuracy of these chatbots in generating clinical content is critical for their informed and safe use in dentistry. In this context, the aim of this study is to evaluate, across prosthodontics and pediatric dentistry, the accuracy, quality, reliability, and readability of information on zirconia crowns provided by five distinct chatbot models (ChatGPT-5, ChatGPT-4o, Gemini 2.5 Flash, DeepSeek-V3, and Microsoft Copilot). The findings revealed statistically significant differences among the models in terms of accuracy, reliability, quality, and readability; therefore, the study’s null hypothesis was rejected. The findings of this study highlighted both the strengths and weaknesses of these models, particularly in terms of accuracy, consistency, and patient education.
In this study, when comparing AI-based chatbots’ responses regarding zirconia materials, topics frequently queried by patients, across prosthodontics and pediatric dentistry, we found that overall the models performed at similar levels between the two domains, with only the DeepSeek-V3 model showing a statistically significant difference in quality (GQS) scores. The higher quality scores achieved by DeepSeek-V3 in pediatric dentistry, compared with prosthodontics, suggest that the model’s training data may include a greater representation of more up-to-date pediatric content. The higher GQS scores observed for DeepSeek in pediatric dentistry may have several explanations. One possibility is that the model includes a greater amount of pediatric or child-focused health content, resulting in more structured responses. Another explanation may relate to assumed user type, as pediatric queries are often directed toward parents or caregivers, which may encourage more explanatory and educational outputs. In contrast, for all other models, no significant differences between the two domains were observed in accuracy (Likert), reliability (DISCERN), or readability (FRES) scores, indicating that chatbot performance remained stable and consistent across prosthodontics and pediatric dentistry. These findings suggest that while current chatbots are largely consistent in cross-domain content generation, certain models (e.g., DeepSeek-V3) may demonstrate more proficient performance in specific specialties. Existing literature includes studies evaluating the accuracy and reliability of responses provided by ChatGPT chatbots to questions across various healthcare domains [7, 32]. In the present study, it was observed that the latest version of ChatGPT-5 attained high scores for accuracy, whereas ChatGPT-4o demonstrated better performance in readability. Similarly, in the systematic review by Tordjman et al. [33], it was reported that while certain versions achieved high accuracy on medical examinations, they lagged behind in readability.
According to the findings of the present study, Gemini stood out relative to the other chatbots, particularly in accuracy, content quality (GQS), and reliability (DISCERN). Similarly, studies by Dursun et al. [34] in orthodontics and by Arpacı et al. [7] addressing oral-health questions likewise emphasized Gemini’s capacity to provide reliable information in dentistry. This may be attributable to Gemini’s being supported by more up-to-date databases and advanced information-filtering algorithms, advantages that may collectively explain its higher accuracy, reliability, and quality scores. Furthermore, Gemini incorporates advanced retrieval-augmented mechanisms and safety-filtered knowledge integration, allowing the model to minimize hallucinations and generate more coherent, evidence-aligned responses [35, 36]. Therefore, Gemini may be considered a more functional AI system in clinical settings, particularly in terms of patient education and the clear, comprehensible delivery of information. Notably, Copilot received the lowest scores across all evaluated domains, particularly in accuracy, reliability, and readability. This suggests that the datasets used to train the model may be limited in dentistry-related content, or that its domain specificity may be inadequate. The present findings indicate that, in its current version, Copilot may have limited utility as a reliable information source in clinical contexts such as dental education and patient education.
In the literature, readability is recognized as facilitating the comprehension of documents and is considered a fundamental component of health literacy [37]. Several studies have reported that ChatGPT-generated responses exhibit suboptimal readability [32, 34]. Poor readability of chatbot-generated content may limit its usability among the general public. Although AI-based patient information materials are often written at a level suitable for clinicians, they require simplification for patient use. In the present study, the DeepSeek- and ChatGPT-based models achieved higher FRES scores than Gemini and Copilot, indicating a greater potential to generate comprehensible and patient-friendly content. As patients and parents increasingly rely on online sources for information about dental treatment options, even modest differences in readability may influence comprehension of instructions, evaluation of risks and benefits, and decision-making. Readability therefore has direct clinical implications, particularly in pediatric dentistry. Poorly readable information may limit caregiver understanding and contribute to misinterpretation of treatment options, whereas clearer explanations may facilitate shared decision-making and improve communication between clinicians and families.
Overall, Gemini-2.5 performed the most consistently across domains, achieving the highest reliability scores and strong accuracy and quality outcomes, except for pediatric readability. Both ChatGPT-5 and ChatGPT-4o showed lower reliability compared with Gemini, although ChatGPT-5 exhibited a slight improvement in prosthodontic quality. DeepSeek-V3 achieved the best readability but ranked behind Gemini and ChatGPT-5 in accuracy and quality. Copilot showed the lowest overall performance. These findings indicate that AI chatbot outputs vary substantially by model, emphasizing the importance of considering model-specific strengths and limitations in clinical dental information.
This study has several limitations. Conducting the assessments solely in English restricts the generalizability of the findings to other languages. In the absence of an expert-validated gold standard, the study evaluated only relative performance rather than absolute accuracy. Querying each model only once may also have limited the capture of inherent response variability. Future studies should incorporate repeated testing and expert-approved reference answers. Additionally, ethical considerations (such as the risk of biased or incomplete information, potential misinformation due to outdated content, and concerns regarding data privacy) must be taken into account. Because the models are trained on different datasets and undergo continuous updates, cross-platform comparisons should be interpreted with caution. Finally, the rapid evolution of scientific knowledge in pediatric dentistry and prosthodontics may result in chatbot responses that do not align with future guidelines. Therefore, the findings should be interpreted cautiously, and future research should encompass a wider range of dental specialties, include patient-based readability assessments, and explore hybrid model approaches.
Conclusion
The study demonstrated that AI chatbots have the potential to augment patient education about both adult and pediatric zirconia. According to the results of the study, the Gemini engine stood out in both clinical fields in terms of information accuracy (Likert), content quality (GQS), and information reliability (DISCERN), while ChatGPT-4o, ChatGPT-5, and DeepSeek demonstrated higher scores particularly in readability (FRES). The persistent underperformance of Copilot emphasizes the need for cautious evaluation before adopting chatbots in clinical contexts. These findings collectively highlight that careful engine selection is crucial when using AI-based content in healthcare, as different engines may yield outputs that vary significantly in both perceived usefulness and content quality. While the results highlight the comparative strengths of different AI engines, these findings reflect potential usefulness rather than clinical validation. Future studies should incorporate patient-based assessments of comprehension and trust, evaluate multilingual performance, and explore the integration of AI tools into clinical education workflows to better determine their real-world applicability.
Acknowledgements
Not applicable.
Abbreviations
- AI
Artificial Intelligence
- LLM
Large Language Model
- FAQ
Frequently Asked Question
- GQS
Global Quality Scale
- FRES
Flesch Reading Ease Score
- DISCERN
Quality Criteria for Consumer Health Information
- ICC
Intraclass Correlation Coefficient
- CI
Confidence Interval
- IQR
Interquartile Range
Authors’ contributions
Conceptualization: [N.K.A.,F.S],;Methodology: [P.C, E.B, F.S],;Formal analysis and investigation: [F.S, P.C],;Writing - original draft preparation: [E.B., P.C] , Writing - review and editing: [N.K.A, E.B, P.C],;Resources: [F.S, P.C],;Supervision: [N.K, P.C, F.S]
Funding
The authors reported there is no funding associated with the work featured in this article.
Data availability
The datasets are available from the corresponding author on reasonable request.
Declarations
Ethics approval and consent to participate
Not applicable. Open-source public data was used in this study.
Consent for publication
Not applicable since there was no direct human contact.
Competing interests
The authors declare no competing interests.
Footnotes
Publisher’s Note
Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
References
- 1.Surlari Z, Budală DG, Lupu CI, et al. Current progress and challenges of using artificial intelligence in clinical dentistry—A narrative review. J Clin Med. 2023;12:73–8. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 2.Ozcivelek T, Ozcan B. Comparative evaluation of responses from DeepSeek-R1, ChatGPT-o1, ChatGPT-4, and dental GPT chatbots to patient inquiries about dental and maxillofacial prostheses. BMC Oral Health. 2025;25(1):871. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 3.Su Z, Tang G, Huang R, Qiao Y, Zhang Z, Dai X. Based on medicine, the now and future of large Language models. Cel Mol Bioeng. 2024;17(4):263–77. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 4.Fadahunsi KP, O’Connor S, Akinlua JT, Wark PA, Gallagher J, Carroll C, Car J, Majeed A, O’Donoghue J. Information quality frameworks for digital health technologies: systematic review. J Med Internet Res. 2021;23(5):e23479. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 5.Yau JY-S, Saadat S, Hsu E, Murphy LS-L, Roh JS, Suchard J, Tapia A, Wiechmann W, Langdorf MI. Accuracy of prospective assessments of 4 large Language model chatbot responses to patient questions about emergency care: experimental comparative study. J Med Internet Res. 2024;26:e60291. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 6.Pan A, Musheyev D, Bockelman D, Loeb S, Kabarriti AE. Assessment of artificial intelligence chatbot responses to top searched queries about cancer. JAMA Oncol. 2023;9(10):1437–40. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 7.Arpaci A, Ozturk AU, Okur I, Sadry S. Evaluation of the accuracy of ChatGPT-4 and gemini’s responses to the world dental federation’s frequently asked questions on oral health. BMC Oral Health. 2025;25(1):1293. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 8.Luo F, Hong G, Wan Q. Artificial intelligence in biomedical applications of zirconia. Front Dent Med. 2021;2:689288. [Google Scholar]
- 9.Alqutaibi AY, Ghulam O, Krsoum M, Binmahmoud S, Taher H, Elmalky W, Zafar MS. Revolution of current dental zirconia: A comprehensive review. Molecules. 2022;27(5):1699. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 10.Hamrah MH, Mokhtari S, Hosseini Z, Khosrozadeh M, Hosseini S, Ghafary ES. Evaluation of the clinical, child, and parental satisfaction with zirconia crowns in maxillary primary incisors: a systematic review. Int J Dent. 2021;2021(1):7877728. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 11.Bapat RA, Yang HJ, Chaubal TV, Dharmadhikari S, Abdulla AM, Arora S, Rawal S, Kesharwani P. Review on synthesis, properties and multifarious therapeutic applications of nanostructured zirconia in dentistry. RSC Adv. 2022;12(20):12773–93. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 12.Jabber HN, Ali R, Al-Delfi MN. Monolithic zirconia in dentistry: evolving aesthetics, durability, and cementation techniques–an in-depth review. Future Dent Res. 2023;1(1):26–36. [Google Scholar]
- 13.Pani SC, Saffan AA, AlHobail S, Bin Salem F, AlFuraih A, AlTamimi M. Esthetic concerns and acceptability of treatment modalities in primary teeth: a comparison between children and their parents. Int J Dent. 2016;2016(1):3163904. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 14.Thorat V, Rao P, Joshi N, Talreja P, Shetty AR. Role of artificial intelligence (AI) in patient education and communication in dentistry. Cureus. 2024;16:59799. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 15.Esmailpour H, Rasaie V, Babaee Hemmati Y, Falahchai M. Performance of artificial intelligence chatbots in responding to the frequently asked questions of patients regarding dental prostheses. BMC Oral Health. 2025;25(1):574. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 16.Shiferaw MW, Zheng T, Winter A,Mike AL, Chan NL. Assessing the accuracy and quality of artificial intelligence (AI) chatbot-generated responses in making patient-specific drug-therapy and healthcare-related decisions. BMC Med Inf Decis Mak. 2024;24(1):404. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 17.Eraslan R, Ayata M, Yagci F, Albayrak H. Exploring the potential of artificial intelligence chatbots in prosthodontics education. BMC Med Educ. 2025;25(1):321. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 18.Rokhshad R, Zhang P, Mohammad-Rahimi H, Pitchika V, Entezari N, Schwendicke F. Accuracy and consistency of chatbots versus clinicians for answering pediatric dentistry questions: A pilot study. J Dent. 2024;144:104938. [DOI] [PubMed] [Google Scholar]
- 19.Topdagı B, Kavaz T. Assessment of information quality in contemporary artificial intelligence systems for digital smile design: A comparative analysis. J Prosthet Dent. 2025;134(4):1279. e1-1279.e8. [DOI] [PubMed] [Google Scholar]
- 20.Taymour N, Fouda SM, Abdelrahaman HH, Hassan MG. Performance of the ChatGPT-3.5, ChatGPT-4, and Google gemini large Language models in responding to dental implantology inquiries. J Prosthet Dent. 2025;134(6):2427–34. [DOI] [PubMed]
- 21.Ayers JW, Poliak A, Dredze M, Leas EC, Zhu Z, Kelley JB, Faix DJ, Goodman AM, Longhurst CA, Hogarth M. Comparing physician and artificial intelligence chatbot responses to patient questions posted to a public social media forum. JAMA Intern Med. 2023;183(6):589–96. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 22.Freire Y, Laorden AS, Pérez JO, Sánchez MG, García VD-F, Suárez A. ChatGPT performance in prosthodontics: assessment of accuracy and repeatability in answer generation. J Prosthet Dent. 2024;131(4):651–9. [DOI] [PubMed] [Google Scholar]
- 23.Mohammad-Rahimi H, Ourang SA, Pourhoseingholi MA, Dianat O, Dummer PMH, Nosrat A. Validity and reliability of artificial intelligence chatbots as public sources of information on endodontics. Int Endod J. 2024;57(3):305–14. [DOI] [PubMed] [Google Scholar]
- 24.Akpınar H. Comparison of responses from different artificial intelligence-powered chatbots regarding the All-on-four dental implant concept. BMC Oral Health. 2025;25(1):922. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 25.Gheisarifar M, Shembesh M, Koseoglu M, Fang Q, Afshari FS, Yuan JCC, Sukotjo C. Evaluating the validity and consistency of artificial intelligence chatbots in responding to patients’ frequently asked questions in prosthodontics. J Prosthet Dent. 2025;134(1):199–206. [DOI] [PubMed] [Google Scholar]
- 26.Joshi A, Kale S, Chandel S, Pal DK. Likert scale: explored and explained. BJAST. 2015;7(4):396. [Google Scholar]
- 27.Aurlene N, Shaik SS, Dickson-Swift V, Tadakamadla SK. Assessment of usefulness and reliability of YouTube™ videos on denture care. Int J Dent Hyg. 2024;22(1):106–15. [DOI] [PubMed] [Google Scholar]
- 28.Rees CE, Ford JE, Sheard CE. Evaluating the reliability of DISCERN: a tool for assessing the quality of written patient information on treatment choices. PEC. 2002;47(3):273–5. [DOI] [PubMed] [Google Scholar]
- 29.Yagci F. Evaluation of YouTube as an information source for denture care. J Prosthet Dent. 2023;129(4):623–9. [DOI] [PubMed] [Google Scholar]
- 30.Flesch R. Flesch-Kincaid readability test. Retrieved. 2007;26(3):2007. [Google Scholar]
- 31.Koo TK, Li MY. A guideline of selecting and reporting intraclass correlation coefficients for reliability research. J Chiropr Med. 2016;15(2):155–63. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 32.Onder CE, Koc G, Gokbulut P, Taskaldiran I, Kuskonmaz SM. Evaluation of the reliability and readability of ChatGPT-4 responses regarding hypothyroidism during pregnancy. Sci Rep. 2024;14(1):243. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 33.Tordjman M, Liu ZL, Yuce M, Fauveau V, Mei YH, Hadjadj J, Bolger I, Almansour H, Horst C, Parihar AS, et al. Comparative benchmarking of the deepseek large Language model on medical tasks and clinical reasoning. Nat Med. 2025;31(8):2550–5. [DOI] [PubMed] [Google Scholar]
- 34.Dursun D, Bilici Gecer R. Can artificial intelligence models serve as patient information consultants in orthodontics? BMC Med Inf Decis Mak. 2024;24(1):211. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 35.Tokgoz Kaplan T, Cankar M. Evidence-Based potential of generative artificial intelligence large Language models on dental avulsion: ChatGPT versus gemini. Dent Traumatol. 2025;41(2):178–86. [DOI] [PubMed] [Google Scholar]
- 36.Raj M, Ravindran V, Arthanari A. Assessing the utility of large Language models in guiding dental practitioners on pediatric patient care: A comparative AI study. J Clin Exp Dent. 2025;17(9):e1099. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 37.Meade MJ, Dreyer CW. Orthodontic treatment consent forms: A readability analysis. J Orthod. 2022;49(1):32–8. [DOI] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Data Availability Statement
The datasets are available from the corresponding author on reasonable request.
