Skip to main content
BMC Oral Health logoLink to BMC Oral Health
. 2026 Apr 30;26:1160. doi: 10.1186/s12903-026-08448-7

Evaluating large language models for orthodontic consultation in patients with periodontitis: a study of reliability, quality, and readability

Jianing Li 1,2,#, Jianhui Ni 1,2,#, Tong Zheng 1,2, Zhigang Zuo 1,2,✉, Yue Wang 1,2,✉
PMCID: PMC13326460  PMID: 42062980

Abstract

Background

This study aimed to evaluate and compare the performance of five publicly accessible large language models (LLMs)-based chatbots, ChatGPT-4o, DeepSeek-V3, Claude-Sonnet-4, Gemini-2.0 Flash, and Grok-3, in addressing inquiries from patients with periodontitis seeking orthodontic treatment. The primary objective was to assess the reliability, quality, and readability of the LLM-generated responses.

Methods

Thirty frequently asked questions regarding orthodontic treatment for patients with periodontitis were sourced from social media platforms and health-related websites and compiled for this study. Each LLM response was evaluated for reliability using the modified DISCERN (mDISCERN) tool, quality using the Global Quality Score (GQS), and readability using the Flesch Reading Ease (FRE) and Flesch–Kincaid Grade Level (FKGL) scores. Differences among models were analysed using linear mixed-effects models, with model treated as a fixed effect and question as a random effect. Post-hoc pairwise comparisons of estimated marginal means were performed with Bonferroni’s adjustment. Significance was set at P < 0.05.

Results

Among the evaluated LLMs, significant performance differences were observed across all metrics (P < 0.001). Grok-3 provided the highest reliability and quality (mDISCERN: 4.20 ± 0.48; GQS: 4.38 ± 0.61), whereas Claude-Sonnet-4 scored the lowest (mDISCERN: 3.54 ± 0.50; GQS: 3.63 ± 0.59). DeepSeek-V3 was rated as most readable (FRE: 33.61 ± 6.11; FKGL: 10.10 ± 1.14), whereas Claude-Sonnet-4 was the least readable (FRE: 4.73 ± 4.14; FKGL: 13.72 ± 1.22). All models produced responses with university-level readability.

Conclusions

Grok-3 demonstrates higher reliability and quality, whereas DeepSeek-V3 generates more readable responses. All models exceed recommended readability thresholds for patient education. However, given the risks of misinformation and readability limitations, these should be considered supplementary educational resources, rather than primary sources of medical information.

Supplementary Information

The online version contains supplementary material available at 10.1186/s12903-026-08448-7.

Keywords: Large language models, Orthodontic treatment, Periodontitis, Artificial intelligence

Background

Large language models (LLMs) represent a major recent breakthrough in artificial intelligence (AI) [1]. Trained on vast text and code datasets, LLMs excel at recognising, predicting, and generating human languages with unprecedented fluency [2]. Building on these capabilities, LLMs have rapidly become the core engines of next-generation conversational agents and chatbots, which are now extensively deployed in different settings through user-friendly dialogue interfaces [3]. Since 2022, rapid progress in LLMs has accelerated their adoption in medicine [4], where on-demand access to information can meaningfully influence communication between patients and healthcare professionals, particularly in complex specialties [5].

In dentistry, orthodontics represents a particularly relevant application area because clinical decision-making and outcomes rely on complex diagnostics and sustained patient-clinician communication [6–8]. Accordingly, LLM-based tools are being explored to support clinical reasoning and patient education [9, 10]. Among these applications, patient-facing consultation is especially practical, offering immediate, 24/7 responses [11, 12]. This capability may be particularly valuable for patients with special medical backgrounds who seek reassurance about treatment safety.

Recently, a growing number of adult patients with varying degrees of periodontitis have sought orthodontic treatment to meet their higher aesthetic and functional demands [13]. Orthodontic treatment for patients with periodontitis exemplifies a scenario of increased complexity and risk [14, 15]. As the integrity of the periodontium is compromised in these patients, the biomechanical environment for tooth movement is inherently fragile [16]. Applying orthodontic forces in the presence of active inflammation can exacerbate alveolar bone loss and potentially lead to irreversible damage, including tooth mobility and loss [17].

These orthodontic treatment-related risks motivate patients with periodontitis to pose questions to LLMs that not only encompass the routine inquiries of general orthodontic patients but are also more profound, specific, and personalised. Their queries are more oriented on the ‘if’ and ‘how safely’ it can be done, given their condition. Consequently, the informational needs of patients with periodontitis undergoing orthodontic consultation are distinct from those of the general patient population. However, existing evaluations of LLMs as orthodontic patient support tools have been based predominantly on the perspectives of general patients, focusing on their ability to answer generic questions [18–20]. These studies rarely segmented specific patient cohorts and thus failed to adequately consider groups with special medical backgrounds and unique concerns, such as those with periodontitis, thereby overlooking the significant disparities in informational needs among different patient populations. Without addressing the unique characteristics of periodontitis-specific enquiries, such evaluations cannot accurately reflect the true capacity of LLMs to serve this vulnerable subgroup of patients seeking orthodontic treatment.

Therefore, a study that focuses on this specific patient cohort is necessary. Here, we conducted a targeted, condition-specific evaluation. We systematically assessed the reliability, quality, and readability of LLM-generated responses to highly specialised and risk-oriented questions posed by patients with periodontitis seeking orthodontic treatment.

Materials and methods

Ethics consideration

Ethical approval for this study was not required from the institutional review board because the study protocol involved neither direct engagement with patients nor the processing of personally identifiable information.

Overview of the workflow

To facilitate understanding of the overall study design and analytical procedures, we provide an overview of the workflow. The comprehensive workflow of this study is shown in Fig. 1.

Fig. 1.

Fig. 1

Overview of the study design and workflow

Formulation of evaluation question sets

To ensure transparency and clinical relevance, the question set was developed through a structured review of platforms including Instagram, TikTok, YouTube, Rednote, and WebMD from 26 October 2025 to 3 November 2025. Two doctoral candidates in orthodontics performed the search using English and Chinese keywords, such as ‘orthodontics with periodontitis’, ‘periodontitis and braces’, and ‘Invisalign for patients with periodontitis’. Inquiries were included if they pertained to treatment eligibility, clinical risks, or long-term outcomes; promotional content, duplicate queries, and non-clinical topics (e.g., pricing) were excluded. Data from the sources were synthesised and consolidated; the 47 most frequently asked questions were identified and compiled. The final list constituted the evaluation set from the perspective of patients with periodontitis. To accurately capture the typical anxieties and enquiries of individuals with periodontitis regarding orthodontics, each question was deliberately framed using simple language to maximise relevance.

To evaluate the validity of the questionnaire, it was distributed to four specialists, three orthodontists, and one periodontist, each with at least 10 years of academic and clinical experience in their respective fields. These specialties were selected to provide an interdisciplinary assessment of the items. Validation followed the Item-Level Content Validity Index (I-CVI) guidelines of Yusoff et al. [21] Using a 4-point Likert scale, responses rated as 3 (‘quite relevant’) or 4 (‘highly relevant’) were classified as acceptable (score = 1). A question was retained only if all four experts assigned a score of 3 or 4 with an I-CVI of 1.00 [22]. Based on the expert’s suggestions, certain items were simplified to enhance clarity and comprehension. Following validation, four items were removed because of their limited scientific relevance, resulting in a final set of 30 validated questions. Two doctoral orthodontic candidates categorised the questions into five domains: evaluation and preparation (n = 6), treatment design (n = 6), management and challenges (n = 8), periodontal surgery (n = 4), and retention and stability (n = 6) (Appendix 1).

LLM evaluation

The selection protocol was guided by a multifaceted set of criteria, including superior performance on recent authoritative benchmarks, public availability for research, established reputation of their developing institutions, and demonstrated proficiency in specialised domains, such as medical question answering. This diverse selection, which includes both performance-optimised and speed-optimised architectures, ensured a comprehensive analysis of various AI models that are characterised by heterogeneous design orientations. Accordingly, five leading LLMs were chosen for evaluation: ChatGPT-4o (OpenAI Global, San Francisco, CA, USA), DeepSeek-V3 (DeepSeek, Beijing, China), Claude-Sonnet-4 (Sonnet; Anthropic, San Francisco, CA, USA), Gemini-2.0 Flash (Gemini; Google, Mountain View, CA, USA), and Grok-3 (xAI, Bay Area, CA, USA).

For reproducibility, each prompt was processed via the default web interfaces of the respective models, using standard platform configurations and without any manual adjustments to settings such as temperature, top-p, or token limits. This approach reflects typical user interactions and reduces potential bias introduced by parameter tuning. Technical specifications for each model are provided in Supplementary Material 1.

A standardised initial prompt (‘I am a patient with periodontitis and want to ask some questions’) was used to simulate clinical scenarios. Five language models were presented with the same questions from 6 November 2025 to 20 November 2025. All responses were generated solely based on their native pre-trained capabilities. To mitigate confounding variables, such as memory bias and inter-query influence, a strict protocol was enforced: each question was submitted in a new, isolated session, effectively clearing any conversational context from the model. Responses from the LLMs were categorised into five distinct formats (A, B, C, D, and E). Outputs were exported and transcribed verbatim into plain-text files with only user interface artifacts and explicit model removed; no rewriting or structural edits were made, and all titles, numbering, bullet points, and information order were preserved. Citations and original response length were retained. Selected examples of the models’ responses are available in Supplementary Material 2.

Evaluation of responses

Three expert orthodontists independently evaluated the models’ responses. Their evaluations were based on comparisons with clinical practice guidelines for the interdisciplinary management of periodontitis [23, 24]. and expert consensus statements regarding orthodontic treatment in patients with periodontitis [25], together with their own clinical experience and professional judgement.

For reliability, we used a modified version of the DISCERN instrument (mDISCERN), which is widely recognised as a standard tool for gauging the quality of digital health information [26]. As reported previously, we used this modified version because the original criteria were not uniformly centred on treatment content [27]. Each item was scored based on the presence of a ‘yes’ response, with a maximum score of 5 (Table 1).

Table 1.

Description of the modified DISCERN 5-point scale

Characteristics
1. Are the aims clear and achieved?
2. Are reliable sources of information used? (i.e., publication cited, the responses are from valid studies/sources)
3. Is the information presented by the LLM balanced and unbiased?
4. Are additional sources of information listed for patient references?
5. Do the responses of the LLM address areas of uncertainty?
∗One point is given for every yes, with a maximum number of five points achievable

Abbreviation: LLM large language model

For quality, the Global Quality Scale (GQS) was selected based on its established use in previous similar research [11], providing a validated framework to assess the clinical applicability of LLM-generated advice for periodontally compromised orthodontic patients. Based on this scale, they were classified as either low-to-moderate (score < 3) or good-to-excellent (score > 3) (Table 2) [28].

Table 2.

Description of the global quality score 5-point scale

GQS Score Characteristics Classification
Score 1 Poor flow of the site. Most information missing, not at all useful for patients. Poor quality
Score 2 Poor flow. Some information listed but many important topics missing, of very limited use to patients. Generally poor quality
Score 3 Suboptimal flow. Some important information is adequately discussed but others poorly discussed, somewhat useful for patients. Moderate quality
Score 4 Generally good flow. Most of the relevant information is listed, but some topics not covered, useful for patients. Good quality
Score 5 Excellent flow. Very useful for patients. Excellent quality

Abbreviation: GQS Global Quality Score

Readability of the responses was evaluated using the Flesch-Kincaid readability test to generate a reading ease score, and the Flesch Reading Ease (FRE) and Flesch–Kincaid Grade Level (FKGL) scores were recorded. The FRE score, calculated as 206.835 − 1.015 × (total words/total sentences) − 84.6 × (total syllables/total words), produces a value between 0 and 100, with higher scores indicating better readability [29]. Texts scoring, 60–70 were generally considered appropriate by the general public [30]. The FKGL applies the formula 0.39 × (total words/total sentences) + 11.8 × (total syllables/total words) − 15.59 to determine the U.S. educational grade level required to comprehend the material [31].

Statistical analysis

Descriptive statistics were presented as means/standard deviations, medians, first and third quartiles (Q1 and Q3), and ranges (minimum–maximum). Data distribution was assessed using the Shapiro–Wilk test. Inter-rater reliability among the three orthodontists was evaluated using the intraclass correlation coefficient (ICC), based on a two-way random-effects model with absolute agreement for average measurements, and the corresponding 95% confidence interval (CI) were reported.

Although the mDISCERN and GQS scores are ordinal variables, both instruments comprise five ordered categories and have been commonly analysed as approximately continuous outcomes in studies assessing the quality of health information [32]. In addition, score distributions were examined before analysis, and based on the results, linear mixed-effects models (LMMs) were applied.

To account for the clustered structure of the data, differences among LLMs were analysed using LMMs, with LLM included as a fixed effect and question as a random effect. Post-hoc pairwise comparisons of estimated marginal means were performed using the Bonferroni’s adjustment for multiple testing. Effect sizes for each mixed-effects model were also calculated, including marginal R², conditional R², and the corresponding Cohen’s f², to quantify the magnitude of model effects.

Statistical significance was set at P < 0.05. All statistical analyses were performed using IBM SPSS Statistics for Windows, version 25 (IBM Corp., Armonk, NY, USA). Graphical visualisation was conducted using GraphPad Prism for macOS, version 9 (GraphPad Software, San Diego, CA, USA).

Results

ICC

Inter-rater reliability among the three orthodontists was high. The ICCs were 0.981 (95%CI: 0.967–0.989) for ChatGPT-4o, 0.835 (95%CI: 0.716–0.884) for DeepSeek-V3, 0.847 (95%CI: 0.724–0.920) for Claude-Sonnet-4, 0.872 (95%CI: 0.804–0.920) for Gemini-2.0 Flash, and 0.851 (95%CI: 0.817–0.893) for Grok-3, indicating good-to-excellent inter-rater reliability.

mDISCERN and GQS

Based on the combined responses generated by the five LLMs across all 30 questions, the analysis revealed a significant effect of model on mDISCERN scores [F(4,116) = 11.01, P < 0.001], with a moderate effect size (Table 3). The estimated marginal means ranged from 3.54 for Claude-Sonnet-4 to 4.20 for Grok-3 (Table 4). Bonferroni-adjusted pairwise comparisons indicated that Grok-3 achieved significantly higher scores than several other models (P < 0.0001) (Fig. 2A).

Table 3.

Effect sizes derived from the linear mixed-effects models

F(df1, df2) Marginal R2 Conditional R2 Cohen’s f2 Effect size
mDISCERN 11.01 (4,116) 0.21 0.28 0.26 Moderate
GQS 13.35 (4,116) 0.25 0.31 0.33 Moderate–large
FKGL 44.21 (4,116) 0.51 0.57 1.05 Large
FRE 143.96 (4,116) 0.79 0.80 3.74 Very large

Marginal R2 represents the proportion of variance explained by the fixed effect (model), whereas conditional R2 represents the variance explained by the full model including both fixed and random effects. Cohen’s f2 quantifies the magnitude of the model effect

Abbreviations: mDISCERN Modified DISCERN, GQS Global Quality Score, FRE Flesch Reading Ease, FKGL Flesch–Kincaid Grade Level

Table 4.

mDISCERN, GQS, FRE and FKGL scores of five large language models

ChatGPT-4o DeepSeek-V3 Claude-Sonnet-4 Gemini-2.0 Flash Grok-3 P
mDISCERN Mean ± SD 3.61 ± 0.56 3.64 ± 0.69 3.54 ± 0.50 3.62 ± 0.55 4.20 ± 0.48 <0.001
Median (Q1–Q3) 4 (3–4) 4 (3–4) 4 (3–4) 4 (3–4) 4 (4–4)
Min-Max 2–4 2–5 3–4 2–4 3–5
GQS Mean ± SD 4.37 ± 0.68 3.90 ± 0.67 3.63 ± 0.59 4.31 ± 0.57 4.38 ± 0.61 <0.001
Median (Q1–Q3) 4 (4–5) 4 (3–4) 4 (3–4) 4 (4–5) 4 (4–5)
Min-Max 2–5 2–5 3–5 2–5 3–5
FRE Mean ± SD 31.73 ± 6.31 33.61 ± 6.11 4.73 ± 4.14 27.29 ± 5.57 27.64 ± 4.44 <0.001
FKGL Mean ± SD 12.06 ± 1.10 10.10 ± 1.14 13.72 ± 1.22 13.25 ± 1.55 13.70 ± 1.60 <0.001

Abbreviation: SD Standard deviation, Q1 1st quartile, Q3 3rd quartile, mDISCERN Modified DISCERN, GQS Global Quality Score, FRE Flesch Reading Ease, FKGL Flesch–Kincaid Grade Level

Fig. 2.

Fig. 2

Pairwise comparisons of model scores on (A) mDISCERN and (B) GQS. P values from Bonferroni-adjusted post-hoc pairwise comparisons of estimated marginal means from the linear mixed-effects model. ***: P < 0.001, ****: P < 0.0001. Abbreviations: mDISCERN, modified DISCERN; GQS, Global Quality Score

Similarly, a significant effect of model was observed for GQS scores [F(4,116) = 13.35, P < 0.001], with a moderate-to-large effect size. Post-hoc comparisons with Bonferroni’s adjustment demonstrated that Grok-3 performed significantly better than DeepSeek-V3 (P < 0.0001) and Claude-Sonnet-4 (P < 0.0001). In addition, significant differences were identified among the remaining four models, ChatGPT-4o, DeepSeek-V3, Claude-Sonnet-4, and Gemini-2.0 Flash, in pairwise comparisons (Fig. 2B).

A comparative analysis of the five LLMs’ performance across five domains revealed markedly different patterns between mDISCERN and the GQS. A clear pattern of uniform dominance was observed in the mDISCERN scores. Grok-3 achieved the highest score in every domain (P = 0.003 for periodontal surgery; P < 0.001 for all other domains). This indicates its comprehensive superiority under the mDISCERN evaluation criteria. In contrast, the GQS scores suggested the specialised strengths of the models. No single LLM dominated the board. Detailed statistical comparisons for each domain are shown in Figs. 3 and 4.

Fig. 3.

Fig. 3

Pairwise comparisons of model mDISCERN scores within each of five domains: (A) evaluation and preparation, (B) treatment design, (C) management and challenges, (D) periodontal surgery, and (E) retention and stability. P values from Bonferroni-adjusted post-hoc pairwise comparisons of estimated marginal means from the linear mixed-effects model. *: P < 0.05, ***: P < 0.001. Abbreviation: mDISCERN Modified DISCERN

Fig. 4.

Fig. 4

Pairwise comparisons of model GQS scores within each of the five domains: (A) evaluation and preparation, (B) treatment design, (C) management and challenges, (D) periodontal surgery, and (E) retention and stability. P values from Bonferroni-adjusted post-hoc pairwise comparisons of estimated marginal means from the linear mixed-effects model. *: P < 0.05,**: P < 0.01, ***: P < 0.001, ****: P < 0.0001. Abbreviation: GQS, Global Quality Score

FRE and FKGL

Readability analysis revealed significant differences among the responses generated by the five LLMs (Table 4). For FRE scores, where higher values indicate better readability, significant differences among models were observed [F(4,116) = 143.96, P < 0.001], indicating a very large effect size. DeepSeek-V3 achieved the highest score (33.61 ± 6.11), suggesting that its responses were the most readable. In contrast, Claude-Sonnet-4 recorded the lowest score (4.73 ± 4.14), indicating that its outputs were the most difficult to read.

For FKGL scores, where lower values correspond to better readability, the model effect was also statistically significant [F(4,116) = 44.21, P < 0.001], with a large effect size. DeepSeek-V3 showed the lowest grade level requirement (10.10 ± 1.14). By comparison, Claude-Sonnet-4 exhibited the highest FKGL score (13.72 ± 1.22), indicating that its responses required a reading level beyond high school and were therefore the least accessible.

For a more detailed analysis, Claude-Sonnet-4 exhibited the lowest FRE scores among the five categories (P < 0.001 for all domains) (Table 5). This indicates that among the models tested, Claude-Sonnet-4 generated content with the highest textual complexity and the poorest readability. In the FKGL assessment, DeepSeek-V3 consistently achieved the lowest scores across all five domains (P = 0.006 for periodontal surgery; P < 0.001 for all other domains). Detailed inter-model comparisons for both metrics are illustrated in Figs. 5 and 6.

Table 5.

FRE and FKGL scores of five large language models

ChatGPT-4o DeepSeek-V3 Claude-Sonnet-4 Gemini-2.0 Flash Grok-3 P
FRE
 Evaluation and preparation 28.68 ± 5.73a, B 38.68 ± 3.08a, C 4.97 ± 5.27a, A 27.07 ± 8.31a, B 25.28 ± 4.52a, B <0.001
 Treatment design 30.65 ± 6.37a, B 31.05 ± 10.67a, B 4.53 ± 1.56a, A 30.88 ± 8.28a, B 29.65 ± 4.98a, B <0.001
 Management and challenges 35.14 ± 8.61a, D 32.66 ± 4.83a, CD 4.98 ± 4.95a, C 26.06 ± 1.58a, B 27.89 ± 2.51a, BC <0.001
 Periodontal surgery 31.68 ± 1.50a, C 31.45 ± 2.01a, C 2.88 ± 2.13a, AB 25.40 ± 2.86a, B 24.68 ± 3.11a, B 0.006
 Retention and stability 31.37 ± 4.56a, B 33.78 ± 3.77a, B 5.62 ± 5.33a, BC 26.83 ± 3.54a, B 29.65 ± 5.57a, B <0.001
 Total 31.73 ± 6.31C 33.61 ± 6.11C 4.73 ± 4.14C 27.29 ± 5.57B 27.64 ± 4.44B <0.001
 P 0.436 0.208 0.904 0.519 0.201
FKGL
 Evaluation and preparation 12.50 ± 0.76a, B 9.33 ± 0.40a, A 13.80 ± 1.18a, B 12.75 ± 1.10ab, B 13.45 ± 0.92a, B <0.001
 Treatment design 12.07 ± 0.78a, B 10.18 ± 1.39a, A 13.95 ± 1.47a, B 11.98 ± 1.15a, B 12.93 ± 1.29a, B <0.001
 Management and challenges 11.56 ± 1.49a, B 9.95 ± 0.87a, A 13.71 ± 1.65a, C 12.86 ± 0.90ab, BC 14.48 ± 1.81a, C <0.001
 Periodontal surgery 12.40 ± 0.70a, AB 11.15 ± 1.69a, A 13.50 ± 0.75a, AB 15.23 ± 1.16c, B 15.00 ± 2.36a, B <0.001
 Retention and stability 12.05 ± 1.32a, B 10.27 ± 0.98a, A 13.55 ± 0.86a, BC 14.23 ± 1.67bc, C 12.82 ± 0.69a, BC <0.001
 Total 12.06 ± 1.10B 10.10 ± 1.14A 13.72 ± 1.22C 13.25 ± 1.55C 13.70 ± 1.60C <0.001
 P 0.590 0.164 0.979 0.002 0.086

Values are presented as mean±standard deviation. Similar superscript letters indicate no statistically significant differences (P > 0.05). Lowercase letters indicate differences between rows, and uppercase letters indicate differences between columns

Abbreviations: FKGL Flesch–Kincaid Grade Level, FRE Flesch Reading Ease

Fig. 5.

Fig. 5

Bar chart of FRE scores for five large language models across the five domains. Abbreviation: FRE, Flesch Reading Ease

Fig. 6.

Fig. 6

Bar chart of FKGL scores for five large language models across the five domains. Dashed red line indicates the recommended FKGL threshold (Grade 7) for healthcare materials. Abbreviation: FKGL, Flesch–Kincaid Grade Level

Additionally, we counted the number of words in the LLMs’ answers. Overall, the word counts for ChatGPT-4o, DeepSeek-V3, Claude-Sonnet-4, Gemini-2.0 Flash, and Grok-3 were 414.55 ± 134.14, 465.82 ± 172.11, 473.91 ± 153.22, 810.18 ± 159.44, and 1045.00 ± 486.62, respectively.

Discussion

This study provides a targeted assessment of the reliability, quality, and readability of LLMs specifically regarding orthodontic enquiries for patients with periodontitis, extending current research into this specialised clinical context. The key distinction of this study is its shift in focus from general orthodontic questions to those originating from a clinically distinct and vulnerable patient subgroup. This focused investigation revealed the capabilities and limitations of LLMs in a specialised medical scenario, a dimension largely overlooked in previous studies that did not differentiate between patient populations.

As LLMs increasingly demonstrate the ability to emulate human-like dialogue, their application in healthcare requires careful consideration. The potential of these systems to propagate flawed or partial information, which could subsequently misdirect patients or compromise their well-being, underscores the importance of validating their performance in medical applications [33]. In this study, Grok-3 achieved the highest mean scores for both reliability (mDISCERN = 4.20 ± 0.48) and quality (GQS = 4.38 ± 0.61), whereas Claude-Sonnet-4 had significantly lower scores on both measures (mDISCERN = 3.54 ± 0.50, GQS = 3.63 ± 0.59). Among the five LLMs evaluated, only individual responses from DeepSeek-V3 and Grok-3 reached a perfect mDISCERN score of 5. Conversely, all five models produced responses that attained the maximum GQS score.

In this study, Grok-3 responses more frequently incorporated periodontal health considerations into answers to orthodontic enquiries. These responses were also more often supported by relevant citations, whereas this feature was less consistent in some of the other models. This pattern may partly explain the relatively higher reliability and quality scores observed for Grok-3 in this specific clinical context. Similar findings have been reported in other medical fields. For example, a study evaluating LLMs in the context of laparoscopic cholecystectomy reported comparatively higher reliability and quality scores for Grok-3 [29]. Another study on radiofrequency ablation for varicose veins found that Grok-3 produced the highest-quality answers in 51.6% of cases, with a higher proportion of high-quality responses than the other evaluated models [34]. However, given the limited academic evaluation of the Grok series, further research is needed to determine whether these findings are reproducible, particularly in the field of oral health.

Among all models, ChatGPT-4o, DeepSeek-V3, and Gemini-2.0 Flash each produced responses that received a low score of 2 on both the mDISCERN and GQS metrics, indicating that, in some scenarios, their outputs fell substantially below acceptable standards. A closer review suggested that these shortcomings could be broadly grouped into four categories: clinical errors, factual errors, omissions, and overgeneralisations. ChatGPT-4o’s claim that orthodontic treatment could stimulate gingival regeneration and eliminate black triangles represented a factual error. DeepSeek-V3’s recommendation to lengthen follow-up intervals for patients with both periodontitis and diabetes constituted a clinical error judged to be inconsistent with standard practice. Gemini-2.0 Flash’s suggestion to switch from fixed appliances to clear aligners during treatment for gingival recession reflected overgeneralisation, with insufficient consideration of biomechanics, cost, and treatment continuity. In addition, some responses failed to adequately emphasise periodontal stability before orthodontic treatment or under addressed supportive periodontal care and long-term retention, indicating that omission was also a recurring problem. These results indicate that even responses that sound clinically reasonable can miss key prerequisites for safe combined periodontal and orthodontic care, or describe them inaccurately. Common misinformation themes are summarised in Supplementary Material 3.

Notably, this study found that the LLM performance rankings were strongly dependent on the chosen evaluation metrics. Under the structured mDISCERN criteria, Grok-3 showed consistently higher scores across domains, suggesting relative strengths in generating technically complete and well-referenced texts. However, this advantage diminished under the GQS framework, which prioritised clinical utility, where models such as ChatGPT-4o and Gemini-2.0 Flash excel in specialised domains. This discrepancy highlights the trade-off between technical completeness (mDISCERN) and practical value (GQS), underscoring the context dependency of model evaluation. Response length may have also acted as a confounder in this study. Longer responses may have achieved higher mDISCERN and GQS scores by covering more relevant points, but they may also have reduced readability. Future studies should consider controlling for or standardising response length when comparing LLM performance. Moreover, absolute score differences (especially in mDISCERN) are modest; therefore, their clinical significance should be interpreted cautiously in the appropriate context. Therefore, a crucial direction for future research is to deconstruct existing metrics and develop novel hybrid frameworks that combine their strengths. This ensures that the selected model is the ‘best fit’ for a specific scenario, ultimately achieving an optimal match between the model capabilities and application needs.

Readability describes the ease with which readers comprehend a written piece. It is influenced by many elements, including the reader’s cognitive ability, educational background, reading environment, interests, purpose, and content of the text, as well as the vocabulary, style, and structure chosen by the writer. Ideally, LLMs intended for patient information and guidance should possess high FRE and low FKGL scores to benefit the users. In this study, Claude-Sonnet-4 demonstrated the poorest readability compared to the other LLMs, with a mean FRE score of 4.73 and a mean FKGL of 13.72. Corroborating our results, Taşyürek et al. similarly found that Claude-Sonnet-4 generated the most complex text in their comparative evaluation of chatbot responses to endodontic questions, noting that its outputs contained ‘very complex sentence structures, long phrases, and advanced-level vocabulary’ [35] This convergence of evidence suggests that Claude-Sonnet-4 may have an inherent bias towards generating academically dense text, potentially reflecting the nature of its training data, which poses a significant challenge to its application in patient-facing communication.

The FKGL scores for all evaluated LLMs ranged from 10.10 to 13.72, indicating a universally high level of reading difficulty. Although DeepSeek-V3 demonstrated comparatively better readability of the data, its responses still corresponded to a U.S. ninth-grade reading level, which exceeded the recommended threshold for health information materials [36]. Both the National Institutes of Health and the American Medical Association advocate setting the readability of patient education materials at or below the sixth- to seventh-grade level to ensure accessibility and understanding [37]. The readability issues identified in this study are mirrored across dental specialties. Research on chatbot responses for dental prostheses by Özcivelek et al. revealed FKGL scores (9.42–10.70) indicative of content that is fairly difficult to read [38]. This challenge is further substantiated by the findings of Guven et al. in traumatic dental injuries and another similar study in orthodontics, both of which concluded that the readability of chatbot-generated content is unfavourable [39, 40]. This deficit hinders informed decision-making. While readability formulas do not fully equate to conceptual comprehension, they remain essential proxies for information accessibility. Consequently, future LLMs should prioritise readability to serve as effective tools for patient education.

Orthodontic treatment of patients with periodontitis is uniquely complex because compromised periodontal tissues create a biomechanically fragile environment for tooth movement [41]. Consequently, any intervention must be predicated on achieving periodontal stability, while the patient’s ability to make informed decisions depends on the precision and completeness of the information provided. While LLMs such as ChatGPT-4 demonstrate substantially evolved capabilities, their reliance on training data can replicate biases and errors, necessitating patient awareness of their abilities and limitations when used for diagnostic or therapeutic guidance [42]. Prior research on LLMs has been largely confined to assessing general enquiries and has failed to adequately account for clinical heterogeneity among patient populations. This has led to an oversight of the specific informational needs of high-risk subgroups, such as those with periodontal disease. To address this gap, the present study adopted a more focused evaluative framework. Rather than applying a one-size-fits-all assessment approach, it provided a contextualised analysis tailored to patients with periodontal disease and their specific clinical concerns.

The findings of this study offer several important considerations for the future application of LLMs in orthodontic consultations for patients with periodontitis. Suboptimal scores for completeness and clarity among certain LLMs reflect the persistent challenge of producing thorough and lucid content in complex medical fields. This performance inconsistency necessitates the refinement of LLMs to balance factual accuracy with comprehensive and coherent delivery of information [43]. Ultimately, the goal is not only to create a technologically advanced tool but also to responsibly integrate it into clinical workflows, ensuring that it serves as a reliable partner in providing patients with dependable treatment information and facilitating shared decision-making.

Limitations

This study has several limitations. First, the question set validation and expert evaluation were conducted by a relatively small panel. Although providing high-level expertise, the limited number of evaluators and inherent subjectivity may restrict the generalisability of the findings. Future studies should employ larger, more diverse expert panels across clinical specialties and geographical regions, potentially utilising a Delphi-based approach for systematic evaluation, to establish a more comprehensive representative assessment framework.

Second, our analysis was confined to single-turn questions, which did not replicate the progressive multi-turn dialogue characteristics of actual clinical consultations. Subsequent studies should develop a more sophisticated test set that features multi-turn conversational scenarios. Such a dataset should integrate multimodal information, including clinical findings, systemic conditions such as diabetes, radiographic images, and digital photographs to rigorously evaluate the ability of LLMs to synthesise complex data and effectively address patient enquiries.

Third, the use of publicly accessible web-based LLM interfaces may limit reproducibility because key model parameters are not disclosed and outputs can vary over time as models are updated. In addition, we did not examine hallucinations, response consistency across repeated queries, stochastic reproducibility, or stratify misinformation by clinical severity; future studies should incorporate these dimensions to provide a more comprehensive safety and performance profile.

Finally, although the mDISCERN instrument used in this study is frequently used in similar research, it may not fully capture all dimensions of information quality, particularly those related to treatment options, risk–benefit balance, and prognostic uncertainty. Future studies should consider incorporating complementary validated instruments or developing domain-specific metrics to address these gaps.

Given the exceptionally rapid iterative cycles of AI, the model versions evaluated in this study may not represent the most recent iterations available at the time of publication. Nevertheless, our findings provide a critical performance baseline and rigorous evaluative framework for the ongoing assessment of evolving LLMs in specialised dental contexts.

Conclusions

In this study, in terms of response reliability and quality, Grok-3 demonstrates the strongest overall performance, whereas Claude-Sonnet-4 has a relatively weaker performance. Regarding the readability of the generated responses, DeepSeek-V3 shows relatively better performance than the other models. However, the outputs produced by all the evaluated models, including DeepSeek-V3, consistently have a universally high level of reading difficulty. Furthermore, LLMs exhibit errors and biases when addressing orthodontic questions from patients with periodontitis. Therefore, their outputs should be treated as adjunctive information only and should be verified by qualified clinicians with individualised assessments before making decisions.

Supplementary Information

Supplementary Material 2. (59.2KB, xlsx)
Supplementary Material 3. (16.5KB, docx)
Supplementary Material 4. (14.9KB, docx)

Acknowledgements

We want to thank the experts who identified and curated the relevant questions and who evaluated and scored the large language models’ performance.

Abbreviations

AI

Artificial intelligence

CI

Confidence intervals

FKGL

Flesch–Kincaid Grade Level

FRE

Flesch Reading Ease

GQS

Global Quality Score

ICC

Intraclass correlation coefficient

I-CVI

Item-Level Content Validity Index

LLMs

Large language models

mDISCERN

Modified DISCERN

Authors’ contributions

JL: original draft preparation, methodology, and investigation; JN: software and investigation; TZ: data curation and validation; ZZ: manuscript review and editing; YW: supervision, methodology, and manuscript review and editing.

Funding

This work was supported by the Tianjin Science and Technology Plan Project (No. 25YFXTHZ00490).

Data availability

The datasets are available from the corresponding author on reasonable request.

Declarations

Ethics approval and consent to participate

Only publicly available information was evaluated in this study; thus, ethical approval was not needed.

Consent for publication

Not applicable.

Competing interests

The authors declare no competing interests.

Footnotes

Publisher’s Note

Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.

Jianing Li and Jianhui Ni contributed equally to this work.

Contributor Information

Zhigang Zuo, Email: zzuo@tmu.edu.cn.

Yue Wang, Email: wangyue1@tmu.edu.cn.

References

  • 1.Fahrner LJ, Chen E, Topol E, Rajpurkar P. The generative era of medical AI. Cell. 2025;188:3648–60. 10.1016/j.cell.2025.05.018. [DOI] [PubMed] [Google Scholar]
  • 2.Shah NH, Entwistle D, Pfeffer MA. Creation and adoption of large language models in medicine. JAMA. 2023;330:866–9. 10.1001/jama.2023.14217. [DOI] [PubMed] [Google Scholar]
  • 3.Gong R, Ding Y, Wang Z, Lv C, Zheng X, Du J, et al. A survey of low-bit large language models: basics, systems, and algorithms. Neural Netw. 2025;192:107856. 10.1016/j.neunet.2025.107856. [DOI] [PubMed] [Google Scholar]
  • 4.Lee S, Jung S, Park JH, Cho H, Moon S, Ahn S. Performance of ChatGPT, Gemini and DeepSeek for non-critical triage support using real-world conversations in emergency department. BMC Emerg Med. 2025;25:176. 10.1186/s12873-025-01337-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 5.Dave T, Athaluri SA, Singh S. ChatGPT in medicine: an overview of its applications, advantages, limitations, future prospects, and ethical considerations. Front Artif Intell. 2023;6:1169595. 10.3389/frai.2023.1169595. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6.Eggmann F, Weiger R, Zitzmann NU, Blatz MB. Implications of large language models such as ChatGPT for dental medicine. J Esthet Restor Dent. 2023;35:1098–102. 10.1111/jerd.13046. [DOI] [PubMed] [Google Scholar]
  • 7.Morishita M, Fukuda H, Muraoka K, Nakamura T, Hayashi M, Yoshioka I, et al. Evaluating GPT-4V’s performance in the Japanese national dental examination: a challenge explored. J Dent Sci. 2024;19:1595–600. 10.1016/j.jds.2023.12.007. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 8.Zhu G, Zhang X, Chen C. Assessing and enhancing the reliability of Chinese large language models in dental implantology. BMC Oral Health. 2025;25:1242. 10.1186/s12903-025-06648-1. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 9.Ryu J, Kim YH, Kim TW, Jung SK. Evaluation of artificial intelligence model for crowding categorization and extraction diagnosis using intraoral photographs. Sci Rep. 2023;13:5177. 10.1038/s41598-023-32514-7. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 10.King MR. The future of AI in medicine: a perspective from a Chatbot. Ann Biomed Eng. 2023;51:291–5. 10.1007/s10439-022-03121-w. [DOI] [PubMed] [Google Scholar]
  • 11.Kurt Demirsoy K, Buyuk SK, Bicer T. How reliable is the artificial intelligence product large language model ChatGPT in orthodontics? Angle Orthod. 2024;94:602–7. 10.2319/031224-207.1. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12.Pupong K, Hunsrisakhun J, Pithpornchaiyakul S, Naorungroj S. Development of chatbot-based oral health care for young children and evaluation of its effectiveness, usability, and acceptability: mixed methods study. JMIR Pediatr Parent. 2025;8:e62738. 10.2196/62738. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 13.Hirschfeld J, Reichardt E, Sharma P, Hilber A, Meyer-Marcotty P, Stellzig-Eisenhauer A, et al. Interest in orthodontic tooth alignment in adult patients affected by periodontitis: a questionnaire-based cross-sectional pilot study. J Periodontol. 2019;90:957–65. 10.1002/JPER.18-0578. [DOI] [PubMed] [Google Scholar]
  • 14.Papageorgiou SN, Antonoglou GN, Michelogiannakis D, Kakali L, Eliades T, Madianos P. Effect of periodontal-orthodontic treatment of teeth with pathological tooth flaring, drifting, and elongation in patients with severe periodontitis: a systematic review with meta-analysis. J Clin Periodontol. 2022;49(Suppl 24):102–20. 10.1111/jcpe.13529. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15.Rath-DeschnerB, Nogueira AVB, Beisel-Memmert S, Nokhbehsaim M, Eick S, Cirelli JA et al. Interaction of periodontitis and orthodontic tooth movement-an in vitro and in vivo study. Clin Oral Investig. 2022;26:171–81. 10.1007/s00784-021-03988-4. [DOI] [PMC free article] [PubMed]
  • 16.Zhang J, Zhang AM, Zhang ZM, Jia JL, Sui XX, Yu LR, Liu HT. Efficacy of combined orthodontic-periodontic treatment for patients with periodontitis and its effect on inflammatory cytokines: A comparative study. Am J Orthod Dentofac Orthop. 2017;152:494–500. 10.1016/j.ajodo.2017.01.028. [DOI] [PubMed] [Google Scholar]
  • 17.Chaushu S, Klein Y, Mandelboim O, Barenholz Y, Fleissig O. Immune changes induced by orthodontic forces: a critical review. J Dent Res. 2022;101:11–20. 10.1177/00220345211016285. [DOI] [PubMed] [Google Scholar]
  • 18.Daraqel B, Wafaie K, Mohammed H, Cao L, Mheissen S, Liu Y, et al. The performance of artificial intelligence models in generating responses to general orthodontic questions: ChatGPT vs Google Bard. Am J Orthod Dentofac Orthop. 2024;165:652–62. 10.1016/j.ajodo.2024.01.012. [DOI] [PubMed] [Google Scholar]
  • 19.Fan Z, Lei J, Shi W, Lin Y, Wang Q, Bao L. Can AI chatbots accurately provide information on orthodontic risks? Angle Orthod. 2025;95:483–9. 10.2319/121424-1021.1. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 20.Martina S, Cannatà D, Paduano T, Schettino V, Giordano F, Galdi M. Reliability of large language model-based chatbots versus clinicians as sources of information on orthodontics: a comparative analysis. Dent J (Basel). 2025;13:343. 10.3390/dj13080343. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 21.Yusoff MSB. ABC of content validation and content validity index calculation. EIMJ. 2019;11:49–54. 10.21315/eimj2019.11.2.6. [Google Scholar]
  • 22.Büker M, Sümbüllü M, Arslan H. Comparative performance of chatbots in endodontic clinical decision support: a 4-day accuracy and consistency study. Int Dent J. 2025;75:100920. 10.1016/j.identj.2025.100920. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 23.Sanz M, Herrera D, Kebschull M, Chapple I, Jepsen S, Beglundh T, Sculean A, Tonetti MS, EFP Workshop Participants and Methodological Consultants. Treatment of stage I–III periodontitis-The EFP S3 level clinical practice guideline. J Clin Periodontol; EFP Workshop Participants and Methodological Consultants. 2020;47 Suppl 22:4–60. 10.1111/jcpe.13290. [DOI] [PMC free article] [PubMed]
  • 24.Herrera D, Sanz M, Kebschull M, Jepsen S, Sculean A, Berglundh T, et al. Treatment of stage IV periodontitis: the EFP S3 level clinical practice guideline. J Clin Periodontol. 2022;49(Suppl 24):4–71. 10.1111/jcpe.13639. [DOI] [PubMed] [Google Scholar]
  • 25.Zhong W, Zhou C, Yin Y, Feng G, Zhao Z, Pan Y, et al. Expert consensus on orthodontic treatment of patients with periodontal disease. Int J Oral Sci. 2025;17:27. 10.1038/s41368-025-00356-w. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 26.Arslan HÇ, Arslan K, Gün M, Karkın PO, Arslanoğlu T. Evaluation of ChatGPT-5 responses in obstetric and gynecological emergencies: concordance, readability, and clinical reliability. BMC Emerg Med. 2025;25:220. 10.1186/s12873-025-01396-5. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 27.Sari F, Çelik Z. Evaluation of the quality, reliability, and readability of ChatGPT-4 responses on exercise and rehabilitation strategies for adolescent myositis. J Adolesc Health. 2025;78:183–9. 10.1016/j.jadohealth.2025.09.015. [DOI] [PubMed] [Google Scholar]
  • 28.Onder CE, Koc G, Gokbulut P, Taskaldiran I, Kuskonmaz SM. Evaluation of the reliability and readability of ChatGPT-4 responses regarding hypothyroidism during pregnancy. Sci Rep. 2024;14:243. 10.1038/s41598-023-50884-w. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 29.Korkmaz YY, Aydın O, Güngör F, Kudaş İ, Sarigoz T, Bostanci O. ChatGPT and other large language models in laparoscopic cholecystectomy: a multidimensional audit of reliability, quality, and readability. Surg Endosc. 2026;40:480–90. 10.1007/s00464-025-12315-x. [DOI] [PubMed] [Google Scholar]
  • 30.Morony S, Flynn M, McCaffery KJ, Jansen J, Webster AC. Readability of written materials for CKD Patients: a systematic review. Am J Kidney Dis. 2015;65:842–50. 10.1053/j.ajkd.2014.11.025. [DOI] [PubMed] [Google Scholar]
  • 31.Trapp C, Schmidt-Hegemann N, Keilholz M, Brose SF, Marschner SN, Schönecker S, et al. Patient- and clinician-based evaluation of large language models for patient education in prostate cancer radiotherapy. Strahlenther Onkol. 2025;201:333–42. 10.1007/s00066-024-02342-3. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 32.Norman G. Likert scales, levels of measurement and the laws of statistics. Adv Health Sci Educ Theory Pract. 2010;15:625–32. 10.1007/s10459-010-9222-y. [DOI] [PubMed] [Google Scholar]
  • 33.Suri G, Slater LR, Ziaee A, Nguyen M. Do large language models show decision heuristics similar to humans? A case study using GPT-3.5. J Exp Psychol Gen. 2024;153:1066–75. 10.1037/xge0001547. [DOI] [PubMed] [Google Scholar]
  • 34.Zyada A, Fakhry A, Nagib S, Seken RA, Farrag M, Abouelseoud A, et al. How well do different AI language models inform patients about radiofrequency ablation for varicose veins? Cureus. 2025;17:e86537. 10.7759/cureus.86537. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 35.Taşyürek M, Adıgüzel Ö, Ortaç H. Comparative evaluation of responses from ChatGPT-5, Gemini 2.5 Flash, Grok 4, and Claude Sonnet-4 chatbots to questions about endodontic iatrogenic events. Healthc (Basel). 2025;13:2615. 10.3390/healthcare13202615. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 36.Helvacioglu-Yigit D, Demirturk H, Ali K, Tamimi D, Koenig L, Almashraqi A. Evaluating artificial intelligence chatbots for patient education in oral and maxillofacial radiology. Oral Surg Oral Med Oral Pathol Oral Radiol. 2025;139:750–9. 10.1016/j.oooo.2025.01.001. [DOI] [PubMed] [Google Scholar]
  • 37.Badarudeen S, Sabharwal S. Readability of patient education materials from the American Academy of Orthopaedic Surgeons and Pediatric Orthopaedic Society of North America web sites. J Bone Joint Surg Am. 2008;90:199–204. 10.2106/JBJS.G.00347. [DOI] [PubMed] [Google Scholar]
  • 38.Özcivelek T, Özcan B. Comparative evaluation of responses from DeepSeek-R1, ChatGPT-o1, ChatGPT-4, and dental GPT chatbots to patient inquiries about dental and maxillofacial prostheses. BMC Oral Health. 2025;25:871. 10.1186/s12903-025-06267-w. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 39.Guven Y, Ozdemir OT, Kavan MY. Performance of artificial intelligence chatbots in responding to patient queries related to traumatic dental injuries: a comparative study. Dent Traumatol. 2025;41:338–47. 10.1111/edt.13020. [DOI] [PubMed] [Google Scholar]
  • 40.Kılınç DD, Mansız D. Examination of the reliability and readability of Chatbot Generative Pretrained Transformer’s (ChatGPT) responses to questions about orthodontics and the evolution of these responses in an updated version. Am J Orthod Dentofac Orthop. 2024;165:546–55. 10.1016/j.ajodo.2023.11.012. [DOI] [PubMed] [Google Scholar]
  • 41.Jepsen K, Sculean A, Jepsen S. Complications and treatment errors involving periodontal tissues related to orthodontic therapy. Periodontol 2000. 2023;92:135–58. 10.1111/prd.12484. [DOI] [PubMed] [Google Scholar]
  • 42.Zhang Q, Wu Z, Song J, Luo S, Chai Z. Comprehensiveness of large language models in patient queries on gingival and endodontic health. Int Dent J. 2025;75:151–7. 10.1016/j.identj.2024.06.022. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 43.Tuzlalı M, Baki N, Aral K, Aral CA, Bahçe E. Evaluating the performance of AI chatbots in responding to dental implant FAQs: a comparative study. BMC Oral Health. 2025;25:1548. 10.1186/s12903-025-06863-w. [DOI] [PMC free article] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Supplementary Material 2. (59.2KB, xlsx)
Supplementary Material 3. (16.5KB, docx)
Supplementary Material 4. (14.9KB, docx)

Data Availability Statement

The datasets are available from the corresponding author on reasonable request.


Articles from BMC Oral Health are provided here courtesy of BMC

RESOURCES