Skip to main content
Medicine logoLink to Medicine
. 2026 Sep 25;105(39):e50839. doi: 10.1097/MD.0000000000050839

Flatfoot and artificial intelligence

Can language models be trusted for patient education?

Hasan Emirhan Usta a, Ruhat Ünlü b,*
PMCID: PMC13619182  PMID: 42798064

Abstract

Flatfoot (pes planus) is common in children and adults. Although often asymptomatic, it may alter lower-limb biomechanics and contribute to pain or injury. Patients and caregivers increasingly use artificial intelligence chatbots such as chat generative pretrained transformer (ChatGPT) and Google Gemini for medical information, yet their responses on flatfoot remain insufficiently studied. This study compared the factual accuracy, added value, omissions, and readability of responses generated by ChatGPT and Google Gemini to standardized patient-oriented questions. In this cross-sectional comparative study, 15 standardized patient-oriented questions were developed from clinical guidelines, peer-reviewed literature, and educational resources from established orthopedic societies, then refined by 2 orthopedic surgeons. Each question was submitted to ChatGPT (OpenAI) and Google Gemini (Google LLC) in July 2025 using identical wording in independent chat sessions. Responses were anonymized and independently rated by 2 board-certified orthopedic surgeons for factual accuracy, added value, and omissions using an investigator-developed 5-point rubric. Discrepant assessments were reviewed by a 3rd senior orthopedic surgeon. Primary analyses used mean scores of the 2 reviewers. Readability was assessed using the Flesch–Kincaid grade level and Flesch reading ease score. Paired differences were analyzed using the Wilcoxon signed-rank test, with P < .05 considered significant. Both models produced generally accurate responses, with mean factual accuracy scores above 4/5. Gemini scored higher than ChatGPT for factual accuracy (4.7 ± 0.2 vs 4.2 ± 0.3, P = .01), added value (4.6 ± 0.3 vs 4.0 ± 0.4, P = .02), and omissions (4.8 ± 0.2 vs 3.9 ± 0.4, P < .001), with higher scores indicating fewer clinically relevant omissions. Gemini also generated longer responses (165 ± 20 vs 120 ± 15 words, P < .001) and showed better readability, with lower Flesch–Kincaid grade level (9.2 ± 1.1 vs 11.8 ± 1.4, P < .001) and higher Flesch reading ease score (59.4 ± 6.2 vs 42.5 ± 5.3, P < .001). Both models generated generally accurate answers to standardized flatfoot questions. Gemini performed better across evaluation domains and readability measures. As these findings are time-specific, repeated assessments and direct patient and caregiver evaluations are needed before broader patient-facing use can be recommended.

Keywords: accuracy, artificial intelligence, ChatGPT, flatfoot, Google Gemini, orthopedics, patient education, pes planus, readability

1. Introduction

Flatfoot (pes planus) is among the most common reasons for pediatric orthopedic visits and remains a frequent finding in adults.[1] It occurs when the medial longitudinal arch of the foot is reduced or completely absent, causing the inner border of the foot to approach or directly touch the ground.[2] The medial longitudinal arch, formed by ligaments, tendons, and fascia that link the forefoot and hindfoot, provides both durability and flexibility. This structure is critical for lower-limb biomechanics, as it maintains balance, absorbs forces during weight-bearing, and stores elastic energy throughout the gait cycle to enhance movement efficiency.[3]

Diagnosing flatfoot in early childhood, particularly in the 1st 3 years of life, can be difficult because soft tissues and physiological fat pads can obscure the arch. Most cases during this period are flexible, symptom-free, and do not require treatment. However, some patients present with more serious underlying conditions. Findings such as achilles tendon tightness, restricted motion at the ankle or subtalar joint, or associated pain warrant closer evaluation and, when necessary, further investigation. While often benign, flatfoot may alter the biomechanics of the lower limbs and spine, predisposing individuals to pain, functional limitations, and even injury. Although conservative treatment remains the mainstay of management, surgical procedures such as subtalar arthroereisis have been proposed for carefully selected symptomatic pediatric patients, with implant selection and long-term outcomes remaining subjects of ongoing debate. For this reason, distinguishing high-risk cases, using appropriate diagnostic tools, and selecting the right management strategies are clinically significant.[4,5]

Despite its high prevalence, patients and families often struggle to access precise and reliable information about flatfoot. Many turn to online resources, including artificial intelligence (AI)-based chatbots, which have become increasingly popular sources of health information. Since the public release of chat generative pretrained transformer (ChatGPT) in 2022 and the launch of Google Gemini soon after, these platforms have quickly become widely used for health information. Because these models are continuously updated, the quality and consistency of their responses may change over time, making periodic evaluation of their medical content increasingly important. This is particularly important for common orthopedic conditions such as flatfoot, for which patients and caregivers frequently seek information before or after clinical consultation.[6]

Yet concerns remain about the reliability of such chatbots when it comes to medical accuracy. Prior orthopedic studies have examined AI performance in areas such as spine disorders and joint replacement, but little attention has been given to common orthopedic conditions such as pes planus, despite their high prevalence in both pediatric and adult populations. Moreover, very few investigations have incorporated expert orthopedic review to evaluate factual accuracy, added value, clinically relevant omissions, and readability of chatbot-generated responses. Furthermore, because the outputs of large language models may evolve following model updates, condition-specific evaluations should be interpreted as time-specific assessments rather than permanent measures of performance.[6,7]

To address this gap, this study aimed to compare the factual accuracy, added value, omissions, and readability of responses generated by ChatGPT and Google Gemini to standardized patient-oriented questions on flatfoot. Readability was included to assess the linguistic characteristics of the generated text, while recognizing that readability metrics alone do not directly measure patient comprehension. The findings are intended to provide an evidence-based assessment of AI-generated information on flatfoot and to inform the appropriate use of large language models in orthopedic patient communication.[7–9]

2. Methods

2.1. Study design

As this study evaluated publicly available AI-generated responses and did not involve human participants, patient data, or identifiable information, institutional ethics committee approval and informed consent were not required. This cross-sectional comparative study evaluated responses generated by GPT-4o through the ChatGPT web interface (OpenAI) and Gemini 2.5 Flash through the Google Gemini web interface (Google LLC) to standardized patient-oriented questions on flatfoot during a predefined data collection period in July 2025. Both platforms were accessed using the default models available at the time. All prompts were submitted using identical wording in independent chat sessions to ensure consistency of data collection and minimize contextual bias between successive responses.

2.2. Question development

A total of 15 standardized questions were developed to capture the concerns most commonly raised by patients and families about flatfoot. These questions were developed from international clinical guidelines, peer-reviewed literature, and educational materials published by established orthopedic societies and institutions. Two fellowship-trained orthopedic surgeons specializing in pediatric foot and ankle disorders reviewed and refined the final list.

The questions were designed to reflect the type of information most commonly sought by patients and caregivers during routine clinical practice. These questions covered the full spectrum of patient concerns, including definition, epidemiology, causes, risk factors, clinical features, diagnostic methods, conservative and surgical treatments, recovery expectations, and long-term outcomes.[10] Examples included: “How is flatfoot diagnosed?” and “What treatment options are available for children with flatfoot?” The complete list of questions is provided in Table 1.

Table 1.

Standardized patient-oriented questions on flatfoot used for chatbot evaluation.

No. Question
1 What is flatfoot and how does it affect the anatomy of the foot?
2 What are the different types of flatfoot (flexible, rigid, and acquired)?
3 How common is flatfoot in children and adults?
4 Does flatfoot cause pain or other health problems?
5 How is flatfoot diagnosed clinically and with imaging methods?
6 At what age is flatfoot considered normal in children, and when is it pathological?
7 What are the main causes of flatfoot (congenital vs acquired)?
8 What are the major risk factors for adult-acquired flatfoot?
9 Can flatfoot be prevented or managed through lifestyle modifications?
10 What nonsurgical treatment options are available (orthotics, physiotherapy, and supportive footwear)?
11 When is surgery indicated in flatfoot, and what surgical procedures are commonly performed?
12 What is subtalar arthroereisis, and in which cases is it used?
13 What is the recovery and rehabilitation process after flatfoot surgery?
14 Can flatfoot lead to problems in other parts of the body, such as the knees, hips, or spine?
15 When should a patient with flatfoot seek medical consultation?

2.3. Response collection

Each of the 15 questions was submitted to GPT-4o through the ChatGPT web interface and Gemini 2.5 Flash through the Google Gemini web interface in July 2025 using identical wording. Each question was submitted in a separate, newly initiated chat session. The platforms were accessed in incognito browser mode to reduce the potential influence of locally stored cookies, browsing history, and previous interactions. No follow-up prompts or requests for clarification were submitted. Responses were collected in their original form without editing, anonymized, and randomly coded as Chatbot A or Chatbot B before being shared with the reviewers.

2.4. Evaluation framework

Two board-certified orthopedic surgeons independently reviewed the anonymized responses while blinded to chatbot identity. Each response was evaluated across 3 domains: factual accuracy, added value, and omissions. Each domain was scored using an investigator-developed 5-point ordinal rubric, with higher scores indicating better performance. For the omissions domain, higher scores indicated fewer clinically relevant omissions.

The evaluation framework was developed by the study investigators with reference to international clinical guidelines, peer-reviewed literature, and orthopedic patient-education resources. Because, to our knowledge, no validated scoring instrument specifically designed to evaluate AI-generated responses on flatfoot was available, the investigator-developed framework was applied consistently to all responses. The scoring criteria and anchor definitions are presented in Table 2.

Table 2.

Investigator-developed 5-point scoring rubric for the evaluation of chatbot responses on flatfoot.

Domain Score Anchor definition
Factual accuracy 1 (Very poor) Substantially inaccurate; contains major medical errors or contradictions.
2 (Poor) Partially inaccurate; contains clinically relevant misconceptions or several errors.
3 (Moderate) Mostly accurate but contains minor errors or incomplete factual information.
4 (Good) Accurate and generally consistent with current clinical guidance, with no major contradictions.
5 (Excellent) Highly accurate and fully consistent with current clinical guidance, with no clinically relevant factual errors.
Added value 1 (Very poor) Provides no meaningful supplementary information or includes largely irrelevant content.
2 (Poor) Provides minimal supplementary information, with predominantly generic or repetitive content.
3 (Moderate) Provides some relevant supplementary information beyond the basic answer.
4 (Good) Provides relevant and clinically useful supplementary information.
5 (Excellent) Provides comprehensive, relevant, and clinically informative supplementary content.
Omissions 1 (Very poor) Critical omissions; multiple essential diagnostic, management, or counseling points are absent.
2 (Poor) Substantial omissions; several clinically important points are absent.
3 (Moderate) Some relevant information is omitted, but the principal elements are addressed.
4 (Good) Only minor omissions; nearly all essential information is addressed.
5 (Excellent) No clinically relevant omissions; all essential information identified by the evaluation framework is addressed.

Higher scores indicate better performance across all 3 domains. For the omissions domain, higher scores indicate fewer clinically relevant omissions.

The reviewers completed their initial evaluations independently and did not discuss their scores during the assessment process. Responses were anonymized, randomly ordered, and coded before evaluation to minimize potential assessment bias. A 3rd senior pediatric orthopedic surgeon reviewed discrepant assessments for methodological oversight; however, the primary quantitative analyses were based on the mean scores assigned by the 2 independent reviewers.

2.5. Readability assessment

Readability was also evaluated to assess the linguistic accessibility of chatbot-generated responses, recognizing that readability metrics do not directly measure patient comprehension. Each response was analyzed using a publicly available online readability tool (www.readabilityformulas.com), and the Flesch–Kincaid grade level (FKGL), Flesch reading ease score (FRES), and word count were recorded.[11] Lower FKGL values and higher FRES scores were interpreted as indicating lower linguistic complexity, with an 8th-grade reading level generally regarded as an appropriate target for patient education materials.

2.6. Bias mitigation

To minimize evaluation bias, all chatbot-generated responses were anonymized, randomly ordered, and coded before assessment. Reviewers were blinded to chatbot identity and evaluated the responses independently without discussing their ratings with each other.

2.7. Statistical analysis

All analyses were performed using Statistical Package for the Social Sciences Statistics version 28 (IBM Corp.). Scores assigned by the 2 independent reviewers were averaged to generate a mean score for each evaluation domain for each chatbot response. The 5-point domain scores for factual accuracy, added value, and omissions, and the readability measures were summarized as mean ± standard deviation. Because responses generated by ChatGPT and Gemini corresponded to the same 15 standardized questions, between-model differences in the evaluation domain scores and readability measures were analyzed using the Wilcoxon signed-rank test for paired observations. Inter-rater and intra-rater agreement for the ordinal rubric scores was evaluated using quadratic-weighted Cohen kappa (κw) together with percentage exact agreement. Ninety-five percent confidence intervals (95% CIs) for the kappa coefficients were calculated using asymptotic standard errors. Intra-rater reliability was assessed by having both reviewers rescore a randomly selected subset of 6 responses (20% of the 30 responses) 1 week after the initial assessment. All tests were 2-tailed, and statistical significance was set at P < .05.

3. Results

A total of 15 standardized questions on flatfoot were evaluated across both AI language models. Overall, both ChatGPT and Gemini achieved high factual accuracy scores, with differences observed across the evaluated domains and readability measures.

Inter-rater reliability between the 2 independent reviewers was assessed using quadratic-weighted Cohen kappa (κw). Domain-specific inter-rater agreement was κw = 0.85 (95% CI, 0.76–0.94; 90.0% exact agreement [27/30]) for factual accuracy, κw = 0.80 (95% CI, 0.69–0.91; 86.7% exact agreement [26/30]) for added value, and κw = 0.81 (95% CI, 0.71–0.91; 86.7% exact agreement [26/30]) for omissions. Intra-rater reliability was assessed using a randomly selected subset of 6 responses (20% of the 30 responses) reevaluated after a 1-week interval. Quadratic-weighted kappa values were κw = 0.87 (95% CI, 0.78–0.96) for reviewer 1 and κw = 0.84 (95% CI, 0.73–0.95) for reviewer 2. These reliability findings are summarized in Table 3.

Table 3.

Inter-rater and intra-rater reliability of the investigator-developed scoring rubric.

Reliability parameter Evaluated scope/subset Quadratic-weighted Cohen κw 95% CI Exact agreement
Inter-rater reliability
 Factual accuracy All responses (n = 30) 0.85 0.76–0.94 90.0% (27/30)
 Added value All responses (n = 30) 0.80 0.69–0.91 86.7% (26/30)
 Omissions All responses (n = 30) 0.81 0.71–0.91 86.7% (26/30)
Intra-rater reliability
 Reviewer 1 Random subset (n = 6) 0.87 0.78–0.96 –
 Reviewer 2 Random subset (n = 6) 0.84 0.73–0.95 –

Exact agreement represents the proportion of responses assigned identical scores by both reviewers. Intra-rater reliability was evaluated in a randomly selected subset of 6 responses rescored after 1 week.

CI = confidence interval.

When assessed for factual accuracy, both models performed well, with mean scores above 4 out of 5. ChatGPT achieved a mean factual accuracy score of 4.2 ± 0.3, while Gemini achieved a higher mean score of 4.7 ± 0.2 (P = .01). For added value, Gemini achieved a higher mean score than ChatGPT (4.6 ± 0.3 vs 4.0 ± 0.4, P = .02).

Gemini achieved a higher mean score in the omissions domain than ChatGPT (4.8 ± 0.2 vs 3.9 ± 0.4, P < .001), indicating fewer clinically relevant omissions according to the scoring rubric. A summary of the evaluation domain scores is presented in Table 4.

Table 4.

Mean 5-point evaluation domain scores (±SD) for ChatGPT and Gemini across 15 standardized flatfoot questions.

Evaluation domain ChatGPT (mean ± SD) Gemini (mean ± SD) P value
Factual accuracy 4.2 ± 0.3 4.7 ± 0.2 .01
Added value 4.0 ± 0.4 4.6 ± 0.3 .02
Omissions 3.9 ± 0.4 4.8 ± 0.2 <.001

Scores range from 1 to 5, with higher scores indicating better performance. For the omissions domain, higher scores indicate fewer clinically relevant omissions.

ChatGPT = chat generative pretrained transformer, SD = standard deviation.

ChatGPT generated shorter responses than Gemini (120 ± 15 vs 165 ± 20 words, P < .001). ChatGPT had a higher FKGL (11.8 ± 1.4 vs 9.2 ± 1.1, P < .001) and a lower FRES (42.5 ± 5.3 vs 59.4 ± 6.2, P < .001). These findings indicate lower linguistic complexity in Gemini-generated responses according to the readability metrics used in this study. A summary of response length and readability measures is presented in Table 5.

Table 5.

Response length and readability measures for ChatGPT and Gemini across 15 standardized flatfoot questions.

Measure ChatGPT (mean ± SD) Gemini (mean ± SD) P value
Word count 120 ± 15 165 ± 20 <.001
FKGL (grade level) 11.8 ± 1.4 9.2 ± 1.1 <.001
FRES (ease score) 42.5 ± 5.3 59.4 ± 6.2 <.001

Lower FKGL values and higher FRES values indicate lower linguistic complexity according to the readability measures used.

ChatGPT = chat generative pretrained transformer, FKGL = Flesch–Kincaid grade level, FRES = Flesch reading ease score, SD = standard deviation.

4. Discussion

The aim of this study was to explore how 2 widely used AI language models, ChatGPT and Google Gemini, perform when asked standardized questions about flatfoot. Our analysis focused on 3 main aspects of their outputs – factual accuracy, added value, and omissions – together with readability. The results showed that both models generated generally accurate responses, although differences were observed in factual accuracy, added value, omissions, and readability.[6,7,12]

In terms of factual accuracy, both models achieved mean scores above 4 out of 5, indicating generally accurate responses within the investigator-developed evaluation framework. Gemini achieved a higher mean factual accuracy score than ChatGPT (4.7 ± 0.2 vs 4.2 ± 0.3, P = .01). Because large language models are continuously updated, these findings should be interpreted as a time-specific assessment of the models evaluated in July 2025 rather than evidence of stable performance over time.[7,13]

The greatest contrast appeared in the domain of added value. Gemini frequently went beyond the basic requirements of the question, including supplementary explanations about biomechanics, risk factors, and long-term outcomes. These findings indicate that Gemini more frequently generated supplementary information beyond the minimum content required to answer the standardized questions. ChatGPT, by comparison, was more restrained. Its answers were often correct but remained at the surface level, with less supporting detail. The potential impact of these differences on patient understanding was not directly evaluated in the present study.[6,14]

With regard to omissions, ChatGPT received lower scores than Gemini in the Omissions domain, indicating more frequent clinically relevant omissions within the study’s evaluation framework. Gemini’s higher scores reflected more complete coverage of the key information considered during expert assessment. These findings should be interpreted within the investigator-developed scoring framework, and the potential impact of such omissions on patient understanding or decision-making was not directly evaluated in the present study.[15]

Readability offered another layer of insight. ChatGPT’s responses were shorter but written at a higher grade level, closer to academic language. This pattern indicates greater linguistic complexity according to the readability metrics used in this study. Gemini’s content, although more verbose, showed more favorable readability metrics and was better aligned with recommended reading levels for patient education materials.[16] These findings should be interpreted cautiously, as readability indices reflect linguistic characteristics and do not directly measure patient comprehension or the ability to use medical information effectively.

This study has limitations. We focused only on 1 condition – flatfoot – so the findings may not generalize to other orthopedic diagnoses. Responses were analyzed based on single prompts, whereas real-world use often involves follow-up questions and iterative clarification. In addition, repeated responses to identical prompts were not evaluated, preventing assessment of within-model response variability and reproducibility. The responses were also evaluated at a single time point, and test–retest reliability across repeated queries or future model versions was not assessed. The scoring of factual accuracy, added value, and omissions was based on an investigator-developed framework and therefore remained dependent on expert judgment despite the standardized scoring criteria. AI models evolve rapidly, and the present findings therefore represent a time-specific assessment of the versions evaluated during the predefined data collection period rather than evidence of stable performance over time. Readability metrics also reflect linguistic characteristics rather than actual patient comprehension, which was not directly evaluated in this study.

Despite these caveats, our work highlights important trends. Both ChatGPT and Gemini generated broadly accurate responses, while Gemini achieved higher scores across the 3 evaluation domains and demonstrated more favorable readability metrics within the evaluation framework used in this study. These differences should not be interpreted as evidence of superior patient comprehension or clinical effectiveness, and AI-generated information should remain supplementary to individualized medical guidance.[17,18]

5. Conclusion

This study evaluated 2 widely used AI language models, ChatGPT and Google Gemini, by comparing their responses to standardized questions on flatfoot. Both systems generated generally accurate responses. ChatGPT generated more concise responses, whereas Gemini achieved higher scores across the 3 evaluation domains within the evaluation framework used in this study. Gemini also demonstrated more favorable readability metrics than ChatGPT. These metrics reflect linguistic characteristics rather than actual patient comprehension.

Taken together, these findings demonstrate that both models generated generally accurate responses to standardized questions on flatfoot, although differences were observed across the 3 evaluation domains and readability measures. Gemini achieved higher scores across the 3 evaluation domains and demonstrated more favorable readability metrics within the framework used in this study. These findings represent a time-specific assessment and should not be interpreted as evidence of superior performance across future model versions or broader clinical applications. Further studies incorporating repeated assessments, standardized external reference materials, and direct evaluation by patients and caregivers are needed to better define the role of AI-generated information in orthopedic patient education. Until such evidence is available, AI-generated information should be interpreted with appropriate clinical oversight.

Acknowledgments

The authors acknowledge the use of ChatGPT and Google Gemini as supportive tools for enhancing language clarity, fluency, and grammatical accuracy during manuscript preparation. All essential scientific responsibilities, including study conception, literature review, methodology, data analysis, and interpretation of results, were conducted solely by the authors. Any AI-generated suggestions were critically reviewed, verified, and incorporated only when considered scientifically appropriate. The authors take full responsibility for the scientific integrity and content of the final manuscript.

Author contributions

Conceptualization: Hasan Emirhan Usta, Ruhat Ünlü.

Data curation: Hasan Emirhan Usta, Ruhat Ünlü.

Formal analysis: Hasan Emirhan Usta, Ruhat Ünlü.

Funding acquisition: Hasan Emirhan Usta, Ruhat Ünlü.

Investigation: Hasan Emirhan Usta, Ruhat Ünlü.

Methodology: Hasan Emirhan Usta, Ruhat Ünlü.

Project administration: Hasan Emirhan Usta, Ruhat Ünlü.

Resources: Hasan Emirhan Usta, Ruhat Ünlü.

Software: Hasan Emirhan Usta, Ruhat Ünlü.

Supervision: Hasan Emirhan Usta, Ruhat Ünlü.

Validation: Hasan Emirhan Usta, Ruhat Ünlü.

Visualization: Hasan Emirhan Usta, Ruhat Ünlü.

Writing – original draft: Hasan Emirhan Usta, Ruhat Ünlü.

Writing – review & editing: Hasan Emirhan Usta, Ruhat Ünlü.

Abbreviations:

AI
artificial intelligence
ChatGPT
chat generative pretrained transformer
CI
confidence interval
FKGL
Flesch–Kincaid grade level
FRES
Flesch reading ease score

As this study did not involve human participants, patient data, or identifiable information, ethics committee approval and informed consent were not required in accordance with national regulations.

The authors have no funding and conflicts of interest to declare.

Data sharing not applicable to this article as no datasets were generated or analyzed during the current study.

How to cite this article: Usta HE, Ünlü R. Flatfoot and artificial intelligence: Can language models be trusted for patient education?. Medicine 2026;105:39(e50839).

References

  • [1].Carr JB, Yang S, Lather LA. Pediatric pes planus: a state-of-the-art review. Pediatrics. 2016;137:e20151230. [DOI] [PubMed] [Google Scholar]
  • [2].Kardm SM, Alanazi ZA, Aldugman TAS, Reddy RS, Gautam AP. Prevalence and functional impact of flexible flatfoot in school-aged children: a cross-sectional clinical and postural assessment. J Orthop Surg Res. 2025;20:783. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [3].Xu L, Gu H, Zhang Y, Sun T, Yu J. Risk factors of flatfoot in children: a systematic review and meta-analysis. Int J Environ Res Public Health. 2022;19:8247. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [4].Van Boerum DH, Sangeorzan BJ. Biomechanics and pathophysiology of flat foot. Foot Ankle Clin. 2003;8:419–30. [DOI] [PubMed] [Google Scholar]
  • [5].Vescio A, Testa G, Sapienza M, et al. Artificial intelligence in pediatric orthopedics: a comprehensive review. Medicina (Kaunas). 2025;61:954. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [6].Büker M, Mercan G. Readability, accuracy and appropriateness and quality of AI chatbot responses as a patient information source on root canal retreatment: a comparative assessment. Int J Med Inform. 2025;201:105948. [DOI] [PubMed] [Google Scholar]
  • [7].Ventresca HC, Davis HT, Gauthier CW, et al. ChatGPT-4 effectively responds to common patient questions on total ankle arthroplasty: a surgeon-based assessment of AI in patient education. Foot Ankle Orthop. 2025;10:24730114251322784. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [8].Sparks CA, Fasulo SM, Windsor JT, et al. ChatGPT is moderately accurate in providing a general overview of orthopaedic conditions. JB JS Open Access. 2024;9:e23.00129. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [9].Kim SE, Lee JH, Choi BS, Han H-S, Lee MC, Ro DH. Performance of ChatGPT on solving orthopedic board-style questions: a comparative analysis of ChatGPT 3.5 and ChatGPT 4. Clin Orthop Surg. 2024;16:669–73. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [10].Parekh AS, McCahon JAS, Nghe A, Pedowitz DI, Daniel JN, Parekh SG. Foot and ankle patient education materials and artificial intelligence chatbots: a comparative analysis. Foot Ankle Spec. 2026;19:19386400241235830. [DOI] [PubMed] [Google Scholar]
  • [11].Kincaid JP, Fishburne RP, Rogers RL, Chissom BS. Derivation of new readability formulas (automated readability index, fog count and flesch reading ease formula) for navy enlisted personnel. 1975. https://apps.dtic.mil/sti/pdfs/ADA006655.pdf. Accessed June 2025.
  • [12].Ibas M, Dursun S, Paksoy M, Ocal R, Karatas E. Accuracy and safety of ChatGPT-4o responses in rhinoplasty postoperative counseling: a panel-based study. Acta Otolaryngol. 2025;145:851–6. [DOI] [PubMed] [Google Scholar]
  • [13].Emile SH, Horesh N, Garoufalia Z, Gefen R, Boutros M, Wexner SD. Assessment of the utility of artificial intelligence-based chatbots in patient education: a systematic review and meta-analysis. Am Surg. 2026;92:258–69. [DOI] [PubMed] [Google Scholar]
  • [14].Jaques A, Abdelghafour K, Perkins O, Nuttall H, Haidar O, Johal K. A Study of orthopedic patient leaflets and readability of AI-generated text in foot and ankle surgery (SOLE-AI). Cureus. 2024;16:e75826. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [15].Marey A, Saad AM, Tanas Y, et al. Evaluating the accuracy and reliability of AI chatbots in patient education on cardiovascular imaging: a comparative study of ChatGPT, gemini, and copilot. Egypt J Radiol Nucl Med. 2025;56:37. [Google Scholar]
  • [16].Strzalkowski P, Strzalkowska A, Chhablani J, et al. Evaluation of the accuracy and readability of ChatGPT-4 and Google Gemini in providing information on retinal detachment: a multicenter expert comparative study. Int J Retina Vitreous. 2024;10:61. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [17].Collins CE, Giammanco PA, Guirgus M, et al. Evaluating the quality and readability of generative artificial intelligence (AI) chatbot responses in the management of achilles tendon rupture. Cureus. 2025;17:e78313. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [18].Stephenson-Moe CA, Behers BJ, Gibons RM, et al. Assessing the quality and readability of patient education materials on chemotherapy cardiotoxicity from artificial intelligence chatbots: an observational cross-sectional study. Medicine (Baltimore). 2025;104:e42135. [DOI] [PMC free article] [PubMed] [Google Scholar]

Articles from Medicine are provided here courtesy of Wolters Kluwer Health

RESOURCES