Abstract
Purpose
To compare the quality, transparency, educational value, and readability of cancer pain information generated by four AI chatbots and to assess whether the outputs met prespecified readability benchmarks for patient education.
Methods
This online cross-sectional comparative study was conducted on July 22, 2026. Nine unmodified Google-Trends-derived cancer-pain-related queries were submitted to ChatGPT 5.5, Microsoft Copilot, Google Gemini 3.5 Flash, and Perplexity, yielding 36 responses. Two oncology clinicians independently assessed the responses using DISCERN, Ensuring Quality Information for Patients (EQIP), Journal of the American Medical Association (JAMA) benchmark criteria, and the Global Quality Score (GQS). Six established indices assessed readability. Matched model comparisons used Friedman tests with query as the repeated unit and Kendall’s W as the omnibus effect size; significant quality outcomes were followed by Holm-adjusted paired Wilcoxon signed-rank tests.
Results
DISCERN did not differ significantly across models (χ2(3)=6.682, P=0.083, W=0.247). EQIP differed across models (χ2(3)=20.721, P<0.001, W=0.767), with Copilot scoring higher than ChatGPT, Gemini, and Perplexity (Holm-adjusted P=0.023 for each). JAMA also differed (χ2(3)=19.645, P<0.001, W=0.728); Copilot and Perplexity scored higher than Gemini (Holm-adjusted P=0.023 and.039, respectively). GQS differed overall (χ2(3)=10.500, P=0.015, W=0.389), although no pairwise comparison remained significant after Holm adjustment. All six readability indices differed overall across models after Holm correction; descriptively, Gemini showed the greatest estimated reading difficulty. Model medians exceeded the prespecified grade-level benchmarks, and FRES medians were below 80.
Conclusion
The four chatbots differed across information quality, source transparency, educational utility, and formula-based readability. Claim-level clinical accuracy and safety were not evaluated in this study and warrant separate guideline-based assessment.
Keywords: artificial intelligence, large language models, chatbots, cancer pain, patient education, readability, health information
Plain Language Summary
People with cancer often search online for explanations of pain, its possible causes, and treatment options. We asked four popular AI chatbots—ChatGPT, Copilot, Gemini, and Perplexity—nine common cancer-pain-related queries. Two oncology clinicians rated the answers with established tools for information quality, presentation, transparency, and usefulness. We also measured reading difficulty. The four chatbots differed across several measures, and no chatbot performed best in every domain. Source information was often limited, and the responses generally exceeded the sixth-grade reading level recommended for patient education. These findings describe differences in how cancer pain information was presented and how difficult it was to read. Future studies should also assess the factual accuracy and clinical safety of individual medical statements.
Introduction
Cancer remains a leading cause of morbidity and mortality worldwide, and more people now live during and after cancer treatment, increasing the need for supportive care across the disease trajectory.1 Pain is among the most prevalent and feared cancer-related symptoms. A large meta-analysis reported pain in approximately 39% of patients after curative treatment, 55% of those receiving anticancer treatment, and 66% of patients with advanced, metastatic, or terminal disease.2 A later update estimated an overall pain prevalence of 44.5%, with moderate-to-severe pain affecting nearly one third of patients.3 Prevalence varies by disease stage, tumor type, treatment exposure, and assessment method, yet pain remains a major clinical and public-health burden. Cancer pain may arise from tumor invasion, bone metastases, visceral obstruction, neuropathic mechanisms, procedures, surgery, radiotherapy, systemic treatment, or comorbid conditions. It is therefore a heterogeneous clinical problem that often requires multidimensional assessment and individualized management.4–6
Appropriate management can reduce suffering, but successful care depends on more than drug selection. Patients and caregivers need to understand why pain occurs, how to assess it, when to use regular and rescue medicines, which adverse effects need attention, and when new or rapidly worsening pain warrants urgent clinical review. Current guidelines recommend a mechanism-informed, stepwise, multimodal approach that may include nonopioid analgesics, opioids, adjuvant medicines, radiotherapy, interventional procedures, rehabilitation, psychological support, and treatment of the underlying cancer.4,7 Misunderstanding these elements can delay reporting, reduce adherence, prompt inappropriate dose changes, intensify fear of addiction, leave adverse effects unmanaged, or lead to avoidable emergency care. Patient education is part of safe symptom management and shared decision-making.
Patients now seek medical information through many channels. Although clinicians remain the preferred source for individualized advice, people increasingly consult websites, social media, search engines, mobile applications, and conversational AI before or after a clinical encounter.8–10 Online information can help users prepare questions, interpret unfamiliar terms, and take a more active role in care. It can also expose them to incomplete, outdated, commercially biased, or misleading claims. These risks matter in cancer pain because the same symptom may reflect expected treatment toxicity, undertreated chronic pain, an oncologic emergency, or a new noncancer diagnosis. Fluent information that lacks context or escalation advice may cause false reassurance or unnecessary alarm.
Large language models (LLMs) have shifted online health searching from documents toward conversation. ChatGPT, Copilot, Gemini, and Perplexity can answer natural-language questions with immediate, apparently tailored explanations. They are available at any time, respond quickly, support dialogue, and can reorganize complex information into lists or summaries. Studies show that LLMs can answer examination questions, draft patient-facing explanations, and sometimes produce responses rated as empathetic or useful.11–14 General-purpose models, however, predict plausible language rather than independently verify each medical claim. Their answers may omit qualifications, express uncertainty poorly, draw on outdated knowledge, fabricate references, or change when the same topic is phrased differently. Chatbot outputs therefore require systematic evaluation before they are used as patient education materials.
Studies across medical specialties report mixed performance. Comparative evaluations in urologic conditions, palliative care, low back pain, rhinoplasty, sexually transmitted diseases, and maintenance hemodialysis found coherent responses but variable reliability, citation practice, and clinical completeness.15–22 Search-oriented systems such as Perplexity or Copilot often scored better on instruments that reward references and organized presentation, whereas other models gave more conversational but less transparent answers. Poor readability recurs across this literature: content rated as acceptable often corresponds to secondary-school or college reading levels rather than the level recommended for patient information.16–22 Explicit prompts to simplify language can improve readability, but default responses often remain too difficult.23–25
Health information can support decisions only when users can read, understand, and apply it. Limited health literacy is associated with poorer knowledge, less effective self-management, lower use of preventive services, and worse health outcomes.26–28 Cancer care compounds these barriers because patients may be older, fatigued, distressed, cognitively burdened by treatment, or unfamiliar with pharmacological terms. Health-literacy guidance commonly recommends short sentences, familiar words, active voice, clear headings, and a reading level near the sixth grade for broadly accessible patient education.29–31 Readability formulas do not measure factual accuracy or comprehension, but they provide reproducible estimates based on sentence length, word length, syllabic complexity, and vocabulary burden. Using several formulas limits dependence on any one mathematical approach.
Quality also has several components. DISCERN assesses the reliability and balance of written treatment information; EQIP examines content, presentation, and usability; the JAMA benchmarks assess authorship, attribution, disclosure, and currency; and the GQS rates flow and usefulness.32–35 No instrument captures every feature of a chatbot answer. These instruments assess complementary aspects of information quality, while factual accuracy of individual medical statements requires direct verification against clinical evidence or guidelines. A response may be well organized yet cite no sources, or it may cite sources while remaining too technical for its audience.
Evidence on chatbot-generated cancer pain information remains limited. Earlier pain studies generally examined broad pain questions or low back pain rather than the specific information needs of people with cancer.19,20 Cancer pain queries concern symptom recognition, site-specific pain, treatment effects, and whether pain implies progression. They test whether a model can communicate uncertainty, avoid overgeneralization, and direct users to professional assessment. We compared four widely accessible chatbots using Google-Trends-derived cancer-pain-related queries. The primary objectives were to determine whether the models differed in quality, transparency, and educational value and whether their responses met the recommended sixth-grade reading standard. We expected quality differences, especially in citation-related measures, and expected default outputs from all models to exceed the recommended reading level.
Methods
Study Design and Reporting Approach
We conducted an online cross-sectional comparison of patient-facing information generated by four AI chatbots. Data were collected on July 22, 2026. The study language was English, the Google Trends region was worldwide, and all information came from publicly accessible web platforms. The design followed methods used in recent comparative studies of medical chatbot responses16–22 and, where applicable, the Chatbot Assessment Reporting Tool (CHART) recommendations.36 The unit of analysis was one chatbot response to one cancer pain query. The study recruited no human participants, accessed no identifiable private data, and involved no intervention in patient care; institutional ethics review and informed consent were therefore not required.
Identification of the Search Topic
The National Library of Medicine Medical Subject Headings database was used to identify a standardized topic term. We selected “cancer pain” because it directly represented the clinical construct and could be applied consistently across search and chatbot platforms. Google Trends was queried using “cancer pain” as a Search term over the preceding five years, with the region set to Worldwide, category set to all categories, and search type set to Web Search. The Top related queries were retrieved, and the 25 highest-ranked candidates were screened. Google Trends reports normalized relative public search interest rather than absolute search volume.37–39
Selection of Google-Trends-Derived Cancer-Pain-Related Queries
Two researchers independently screened the 25 candidate queries for relevance to cancer pain and then resolved disagreements through discussion. Queries clearly unrelated to cancer pain, unable to function reasonably as an information-seeking query or search prompt, or duplicating/highly overlapping another item were excluded. “Cancer symptoms” was retained because it appeared among the Top related queries for the Search term “cancer pain” and was considered relevant to symptom-recognition information seeking in which pain may constitute a cancer-related symptom. After screening, nine queries were retained: “cancer symptoms”, “cancer back pain”, “breast cancer pain”, “stomach cancer pain”, “lung cancer pain”, “colon cancer pain”, “bone cancer pain”, “does cancer cause pain”, and “pancreatic cancer pain”. The queries were entered exactly as retrieved, without modification.
Selection of AI Chatbots
We evaluated four widely used general-purpose AI chatbots: ChatGPT 5.5, Microsoft Copilot, Google Gemini 3.5 Flash, and Perplexity. All four systems were accessed from China through their official web interfaces under signed-out, free/public access conditions. Web search or retrieval was enabled for all four systems; Deep Research or other special reasoning modes were not used. Memory, personalization, and custom instructions were not used.
Standardized Prompting and Response Collection
Each of the nine queries was submitted once to each chatbot, yielding 36 responses. Queries were entered in their original form, without explanatory context, role prompting, requests for citations, or instructions to simplify the language. Each query began in a separate conversation, no follow-up questions were asked, and the first generated response was retained. Complete responses, including headings, lists, warnings, and citations generated by the chatbot, were saved without editing before quality assessment. Formatting was standardized only for the readability analysis.
Blinding and Rating Procedure
Two oncology clinicians, each with more than five years of clinical experience and relevant experience in cancer pain care, independently reviewed the responses. Model names and interface elements were removed from the rating copies to limit brand-related expectations. Before formal assessment, both evaluators received standardized training on the published instructions for each instrument and completed a calibration exercise using eight responses generated from two queries. They then scored all responses independently, and disagreements were resolved through discussion to consensus.
DISCERN
DISCERN is a 16-item instrument for assessing the reliability and quality of consumer health information about treatment choices.32 The first eight items assess reliability, the next seven address treatment options, and the final item gives an overall quality rating. Scores were interpreted as 63–80 excellent, 51–62 good, 39–50 fair, 27–38 poor, and 16–26 very poor. DISCERN emphasizes explicit aims, relevant sources, balanced discussion of benefits and risks, acknowledgement of uncertainty, and support for shared decision-making.
Ensuring Quality Information for Patients
EQIP assesses written health information from a patient-education perspective.33 The version used had 20 items scored yes (1) or no (0). The total was divided by 20 and multiplied by 100 to give a percentage. Scores of 76–100 were classified as excellent, 51–75 as good, 26–50 as poor or indicating serious quality problems, and 0–25 as indicating very serious quality problems. EQIP complements DISCERN by assessing whether the material provides practical, understandable information, follows a logical structure, and can be used by patients rather than merely listing biomedical facts.
JAMA Benchmark Criteria and Global Quality Score
The JAMA benchmark criteria assess four transparency features of online medical information: authorship, attribution or references, disclosure of ownership or conflicts, and currency or date of creation/update.34 Each criterion receives 0 or 1 point, for a maximum of 4. A low JAMA score does not necessarily indicate factual error, but it means the reader cannot readily identify who produced the information, which evidence supports it, whether conflicts exist, or how current it is. The GQS is a five-point rating of flow, quality, and usefulness: 1 indicates very poor, 2 poor, 3 fair, 4 good, and 5 excellent.35 It captures coherence and practical value in a single global judgment.
Text Preparation for Readability Analysis
Before readability measurement, each answer was converted to plain text. We removed Markdown symbols, bullet characters, list numbers, URLs, inline citation markers, standalone reference lists, and interface-generated labels. Headings were retained as ordinary text when they contained substantive information. We did not alter wording, sentence order, medical terms, warnings, or clinical content. This protocol reduced the influence of web formatting while preserving the language generated by each chatbot.
Readability Measures and Benchmark
Readability was calculated with Python 3.10.4 and textstat 0.7.13. We used six established indices: ARI, GFI, FKGL, CLI, SMOG, and FRES.40–45 ARI and CLI depend mainly on character counts and sentence length; FKGL and FRES use sentence length and syllables; GFI gives weight to long sentences and complex words; and SMOG counts polysyllabic words. For ARI, GFI, FKGL, CLI, and SMOG, higher values approximate a higher US school grade needed to understand the text. FRES runs in the opposite direction, with higher scores indicating easier reading. The prespecified acceptable threshold was below 6 for ARI, GFI, FKGL, CLI, and SMOG and at least 80 for FRES. This conservative threshold follows commonly cited guidance for broadly accessible patient education.29–31
Statistical Analysis
Quality and readability outcomes were summarized as medians and interquartile ranges (IQRs). Inter-rater reliability was calculated from the independent pre-consensus ratings. DISCERN and EQIP were assessed using ICC(2,1), while JAMA and GQS were assessed using quadratically weighted Cohen’s κ; raw percentage agreement was also reported for JAMA and GQS. Ninety-five percent confidence intervals were reported for the agreement statistics.
Because the same nine queries were submitted to all four chatbots, between-model comparisons were treated as matched repeated-measures analyses with query as the repeated unit. DISCERN, EQIP, JAMA, and GQS were compared using Friedman tests, with Kendall’s W reported as the omnibus effect size. When a quality outcome had a significant omnibus test, paired Wilcoxon signed-rank tests with Holm adjustment were used for post hoc comparisons.
Readability scores were compared with the prespecified thresholds using two-sided one-sample Wilcoxon signed-rank tests. Between-model differences in the six readability indices were assessed using Friedman tests with query as the repeated unit and Kendall’s W as the effect size; Holm adjustment was applied across the six omnibus P values. All tests were two sided, and P<0.05 was considered statistically significant.
Results
Query Selection and Dataset
Google Trends returned 25 candidate related queries. Screening yielded nine Google-Trends-derived cancer-pain-related queries (Table 1). Each query was submitted to all four chatbots, producing 36 complete responses. No response was missing or excluded after generation. Supplementary Data S1 provides query-level numerical values for all 36 model-query responses, including the final quality scores and six readability indices.
Table 1.
Google Trends Queries Related to “Cancer Pain”
| Rank | Query | Relative Relevance |
|---|---|---|
| 1 | Cancer symptoms | 100 |
| 2 | Cancer back pain | 61 |
| 3 | Breast cancer pain | 47 |
| 4 | Stomach cancer pain | 29 |
| 5 | Lung cancer pain | 24 |
| 6 | Colon cancer pain | 24 |
| 7 | Bone cancer pain | 21 |
| 8 | Does cancer cause pain | 19 |
| 9 | Pancreatic cancer pain | 18 |
Notes: Nine items were retained from 25 candidate queries after removal of unrelated, duplicate or highly similar, and otherwise ineligible items. Google Trends relative relevance is normalized and does not represent absolute search volume.
Inter-Rater Reliability
DISCERN ICC(2,1) was 0.805 (95% CI, 0.633–0.888), and EQIP ICC(2,1) was 0.864 (95% CI, 0.751–0.924). Quadratically weighted Cohen’s κ was 0.942 (95% CI, 0.810–1.000) for JAMA, with 97.2% raw agreement, and 0.730 (95% CI, 0.416–0.892) for GQS, with 77.8% raw agreement.
Overall Quality Comparisons
The matched analyses showed no significant overall difference in DISCERN (χ2(3)=6.682, P=0.083). Significant overall differences were observed for EQIP (χ2(3)=20.721, P<0.001), JAMA (χ2(3)=19.645, P<0.001), and GQS (χ2(3)=10.500, P=0.015) (Table 2).
Table 2.
Quality Scores and Matched Comparisons of AI Chatbots
| Instrument | ChatGPT | Copilot | Gemini | Perplexity | Friedman χ2(3) | Kendall’s W | P value |
|---|---|---|---|---|---|---|---|
| DISCERN | 34.00 (34.00, 43.00) | 42.00 (40.00, 42.00) | 35.00 (32.00, 38.00) | 41.00 (39.00, 43.00) | 6.682 | 0.247 | 0.083 |
| EQIP | 60.00 (55.00, 70.00) | 85.00 (80.00, 85.00) | 65.00 (60.00, 65.00) | 75.00 (70.00, 75.00) | 20.721 | 0.767 | <0.001 |
| JAMA | 0.00 (0.00, 1.00) | 1.00 (1.00, 1.00) | 0.00 (0.00, 0.00) | 1.00 (1.00, 1.00) | 19.645 | 0.728 | <0.001 |
| GQS | 4.00 (4.00, 4.00) | 4.00 (4.00, 4.00) | 4.00 (4.00, 4.00) | 3.00 (3.00, 4.00) | 10.500 | 0.389 | 0.015 |
Notes: Values are median (Q1, Q3). P values are from Friedman tests with query as the repeated unit. Kendall’s W is the omnibus repeated-measures effect size.
Abbreviations: EQIP, Ensuring Quality Information for Patients; GQS, Global Quality Score; JAMA, Journal of the American Medical Association.
DISCERN Results
Median DISCERN scores were 34.00 (IQR 34.00–43.00) for ChatGPT, 42.00 (40.00–42.00) for Copilot, 35.00 (32.00–38.00) for Gemini, and 41.00 (39.00–43.00) for Perplexity. ChatGPT and Gemini were in the poor category, and Copilot and Perplexity were in the fair category. The overall Friedman test was not significant (P=0.083).
EQIP Results
Median EQIP scores were 60.00 (IQR 55.00–70.00) for ChatGPT, 85.00 (80.00–85.00) for Copilot, 65.00 (60.00–65.00) for Gemini, and 75.00 (70.00–75.00) for Perplexity. Copilot scored higher than ChatGPT, Gemini, and Perplexity in Holm-adjusted paired comparisons (P=0.023 for each). Other pairwise comparisons were not significant after adjustment (Table 3).
Table 3.
Paired Post Hoc Comparisons of AI Chatbots (Holm-Adjusted P values)
| Pairwise Comparison | EQIP | JAMA | GQS |
|---|---|---|---|
| ChatGPT-Copilot | 0.023 | 0.250 | 1.000 |
| ChatGPT-Gemini | 1.000 | 0.375 | 1.000 |
| ChatGPT-Perplexity | 0.125 | 0.375 | 1.000 |
| Copilot-Gemini | 0.023 | 0.023 | 1.000 |
| Copilot-Perplexity | 0.023 | 1.000 | 0.375 |
| Gemini-Perplexity | 0.094 | 0.039 | 0.625 |
Notes: Cells show Holm-adjusted P values from paired Wilcoxon signed-rank tests. Post hoc testing was performed only after a significant Friedman omnibus test. Statistical significance was defined as P<0.05.
JAMA Benchmark Results
Median JAMA scores were 0.00 (IQR 0.00–1.00) for ChatGPT, 1.00 (1.00–1.00) for Copilot, 0.00 (0.00–0.00) for Gemini, and 1.00 (1.00–1.00) for Perplexity. Copilot scored higher than Gemini (Holm-adjusted P=0.023), and Perplexity scored higher than Gemini (Holm-adjusted P=0.039); no other pairwise comparison remained significant (Table 3).
GQS Results
Median GQS scores were 4.00 (IQR 4.00–4.00) for ChatGPT, Copilot, and Gemini and 3.00 (3.00–4.00) for Perplexity. The overall Friedman test was significant (P=0.015), and no pairwise comparison remained significant after Holm adjustment (Table 3).
Score Distributions
Figure 1 shows the distributions of DISCERN, EQIP, JAMA, and GQS scores across the four chatbots. EQIP showed the clearest separation among models, with Copilot generally displaying higher scores than the other chatbots. JAMA scores were low overall, but Copilot and Perplexity tended to show higher distributions than Gemini. DISCERN scores showed substantial overlap across models, although Copilot and Perplexity tended to occupy somewhat higher score ranges than ChatGPT and Gemini. GQS scores were relatively tightly clustered, with considerable overlap across the four models.
Figure 1.

Distribution of DISCERN, EQIP, GQS, and JAMA scores across the four AI chatbots. Boxplots show the score distributions across the nine included queries for each chatbot.
Readability and Benchmark Comparison
All six readability indices differed significantly across models in Friedman tests, and all six omnibus results remained significant after Holm correction across the six indices, with adjusted P≤.0034 (Table 4). Descriptively, Gemini had the highest median ARI, GFI, FKGL, CLI, and SMOG values and the lowest median FRES. All model medians exceeded 6 for ARI, GFI, FKGL, CLI, and SMOG, and all FRES medians were below 80. The prespecified two-sided one-sample Wilcoxon benchmark comparisons were significant across all 24 model-index comparisons (all P≤0.0078; Supplementary Table S1). Mean readability scores were displayed against the prespecified benchmark lines (Figure 2).
Table 4.
Readability Scores and Matched Comparisons of AI Chatbot Responses
| Index | ChatGPT | Copilot | Gemini | Perplexity | Prespecified Benchmark | Friedman χ2(3) | Kendall’s W | Holm-Adjusted P value |
|---|---|---|---|---|---|---|---|---|
| ARI | 10.76 (10.28, 11.29) | 10.78 (10.67, 10.91) | 13.50 (12.71, 13.73) | 10.88 (9.95, 11.87) | <6 | 16.600 | 0.615 | 0.003 |
| GFI | 10.66 (10.20, 11.16) | 10.26 (9.75, 11.00) | 13.73 (13.30, 14.33) | 10.59 (8.47, 11.04) | <6 | 16.618 | 0.615 | 0.003 |
| FKGL | 8.22 (7.74, 8.80) | 8.53 (8.05, 9.03) | 11.42 (10.99, 11.67) | 8.38 (7.40, 9.36) | <6 | 16.600 | 0.615 | 0.003 |
| CLI | 10.90 (9.85, 11.71) | 12.67 (12.60, 13.39) | 13.34 (12.66, 13.80) | 10.65 (10.33, 12.04) | <6 | 19.933 | 0.738 | <0.001 |
| SMOG | 10.82 (10.25, 11.09) | 10.67 (10.52, 11.04) | 13.26 (12.95, 13.31) | 10.73 (10.02, 10.89) | <6 | 16.600 | 0.615 | 0.003 |
| FRES | 59.61 (57.82, 65.46) | 51.40 (48.07, 57.73) | 45.25 (44.84, 46.33) | 59.04 (58.90, 64.35) | ≥80 | 21.133 | 0.783 | <0.001 |
Notes: Values are median (Q1, Q3). Between-model comparisons used Friedman tests with query as the repeated unit; Kendall’s W is the omnibus effect size. Holm adjustment was applied across the six readability omnibus P values.
Abbreviations: ARI, Automated Readability Index; GFI, Gunning Fog Index; FKGL, Flesch-Kincaid Grade Level; CLI, Coleman-Liau Index; SMOG, Simple Measure of Gobbledygook; FRES, Flesch Reading Ease Score.
Figure 2.

Mean readability scores by chatbot. Dashed lines mark the prespecified benchmarks: 6 for ARI, GFI, FKGL, CLI, and SMOG and 80 for FRES.
Query-Level Variability
The query-level heatmap in Figure 3A showed generally higher EQIP scores for Copilot across multiple queries, while DISCERN scores varied within each model. JAMA scores remained low across most model-query combinations. The readability heatmap also showed variation across models and queries (Figure 3B). Overall, performance varied according to both the chatbot and the specific cancer-pain query.
Figure 3.

Query-level heatmaps for (A) quality and transparency and (B) readability. Rows show the nine included queries; columns show the four chatbots.
Abbreviations: ARI, Automated Readability Index; CLI, Coleman–Liau Index; EQIP, Ensuring Quality Information for Patients; FRES, Flesch Reading Ease Score; FKGL, Flesch–Kincaid Grade Level; GFI, Gunning Fog Index; GQS, Global Quality Score; JAMA, Journal of the American Medical Association; SMOG, Simple Measure of Gobbledygook.
Discussion
The chatbots produced coherent cancer pain information, but their performance varied across the assessed domains. In the matched analysis, DISCERN showed no statistically significant overall difference among models, whereas EQIP, JAMA, and GQS showed significant overall differences. After Holm adjustment, Copilot had higher EQIP scores than ChatGPT, Gemini, and Perplexity, while Copilot and Perplexity had higher JAMA scores than Gemini. GQS differed overall, although no pairwise comparison remained significant after adjustment. All six readability indices also showed significant overall differences across models. Descriptively, Gemini showed the greatest estimated reading difficulty. At the same time, the median outputs of all four models did not meet the prespecified patient-education readability benchmarks. These findings indicate that the observed differences depended on the domain being assessed rather than following a single consistent pattern across all measures.
Information Quality and Transparency
The nonsignificant DISCERN difference should be considered in relation to the construct measured by the instrument. Copilot and Perplexity had fair median scores, whereas ChatGPT and Gemini had poor median scores, but none reached the good range. DISCERN places substantial emphasis on treatment choices, benefits, risks, uncertainty, and shared decision-making.32 Brief searches such as “bone cancer pain” or “does cancer cause pain” do not explicitly request treatment comparisons, and general-purpose chatbots may therefore provide an overview without addressing every treatment-related DISCERN item. Because DISCERN was originally developed to evaluate consumer information on treatment choices, scores for symptom-oriented queries should be interpreted with caution.
Copilot’s EQIP performance indicated that its responses more often contained characteristics valued in patient information, including clear organization, relevant explanation, practical guidance, and an accessible sequence. The query-level results suggested that this pattern occurred across multiple queries rather than being driven by a single response. Search integration may facilitate the organization of symptoms, possible causes, warning signs, and next steps. However, performance on EQIP did not extend uniformly to the other measured domains. Copilot’s DISCERN median remained in the fair range, its JAMA median remained low, and its readability estimates remained above the prespecified grade-level benchmarks.
Perplexity also performed relatively well on DISCERN, EQIP, and JAMA, while showing the lowest median GQS. Citation-rich or retrieval-oriented responses may differ in flow from more conversational answers. Inline citations, source-based qualifications, and repeated caveats may support source checking while also interrupting narrative continuity. Previous studies have reported similar patterns, with Perplexity often performing relatively well on source-related measures while not necessarily producing the simplest or most concise patient-facing explanations.16,17 This variation illustrates the value of evaluating multiple information characteristics rather than relying on a single global score.
JAMA scores remained low across all four systems. The JAMA benchmark reflects visible features such as authorship, attribution, disclosure, and currency.34 Copilot and Perplexity were more likely than Gemini to satisfy some of these criteria, although their median scores remained limited. These differences may partly reflect product and interface behavior because systems with integrated web retrieval are more likely to display citations or source links. JAMA therefore provides information about visible source transparency and attribution in the user-facing response. Factual correctness of individual medical claims requires separate evaluation against appropriate clinical evidence or guidelines.
The present findings are broadly consistent with previous evaluations of medical chatbots. Studies addressing sexually transmitted disease information, maintenance hemodialysis, palliative care, common pain questions, and low back pain have reported variability across models in information quality, source presentation, and readability, with patient-facing outputs frequently exceeding recommended reading levels.16–20 These findings suggest that uneven performance across quality, transparency, and readability domains occurs across multiple health-information settings rather than being specific to cancer pain.
Cancer pain nevertheless represents a clinically complex information domain. A short query such as “cancer back pain” may refer to musculoskeletal pain, vertebral metastasis, treatment-related neuropathic pain, spinal cord compression, or an unrelated benign condition. “Breast cancer pain” may include tumor-related discomfort, postoperative pain, aromatase-inhibitor-associated arthralgia, radiation effects, or lymphedema. “Pancreatic cancer pain” may involve visceral or neuropathic mechanisms, analgesic treatment, interventional options, or assessment of newly developing symptoms. Effective patient information in this setting requires clear communication of uncertainty, clinically relevant possibilities, and indications for professional assessment. The measures used in the present study characterize information quality, presentation, transparency, and readability; claim-level factual accuracy, guideline concordance, treatment appropriateness, and clinical safety were not directly evaluated.
Readability and Patient Education
Readability was a consistent concern across the evaluated systems. All six indices showed significant overall differences among models. Gemini had the highest median ARI, GFI, FKGL, CLI, and SMOG values and the lowest median FRES, representing the greatest estimated reading difficulty descriptively. More importantly for patient education, the median outputs of all four systems did not meet the prespecified readability benchmarks.
Several features of chatbot-generated cancer pain information may contribute to this level of linguistic complexity. Responses commonly contain multisyllabic clinical terms such as “metastasis”, “neuropathic”, “radiotherapy”, “analgesic”, “gastrointestinal”, and “multidisciplinary”. Models may also combine causes, qualifications, and warning information within long sentences or use nominalized wording rather than more direct expressions. Bullet-point formatting can improve visual organization while leaving sentence structure and terminology relatively complex. Visual organization and formula-based readability therefore describe related but distinct aspects of presentation.
The six readability indices are also not independent outcomes. ARI, GFI, FKGL, CLI, SMOG, and FRES share underlying linguistic characteristics, including sentence length, word length, and syllabic complexity. Their generally concordant direction is therefore best understood as a convergent pattern of estimated text complexity rather than six independent confirmations. In addition, readability formulas estimate linguistic difficulty rather than actual patient comprehension or overall suitability of the information.
The potential implications of text complexity are particularly relevant in cancer care. Limited health literacy is more common among older adults, people with fewer educational opportunities, those communicating in a second language, and socioeconomically disadvantaged populations.26–28 Cancer and its treatment may further affect concentration through fatigue, anxiety, sleep disturbance, pain, and treatment-related cognitive symptoms. Complex sentences containing multiple causes, qualifications, and warnings may therefore increase information-processing demands, especially when users are seeking information during periods of symptom burden.
Simplification alone is not sufficient. A short response with a lower estimated reading grade may still be incomplete, while some medical terms are clinically necessary and are better explained than removed. Patient-facing information may benefit from layered presentation, including a brief plain-language summary, clearly prioritized information, definitions of necessary technical terms, optional additional detail, and links or citations to authoritative sources. Studies of rewritten patient information have shown that targeted prompting can improve formula-based readability.23–25,46 Future work should determine whether such changes also improve patient comprehension while preserving clinically important content.
Clinical Implications and Future Directions
For patients and caregivers, general-purpose chatbots may support general education, clarification of medical terminology, and preparation of questions for clinical encounters. Cancer pain, however, is highly dependent on individual context, including cancer type, disease stage, treatment history, medication exposure, symptom onset, and associated features. Chatbot-generated information may therefore be most useful as a source of general information and preparation for communication with healthcare professionals.
For clinicians, increasing use of AI-based health information provides another opportunity to identify patients’ information needs. Asking whether a patient has searched online or consulted a chatbot may reveal areas of uncertainty or misunderstanding and facilitate discussion of appropriate information sources. Clinician-reviewed patient-education materials may also provide useful reference points when patients encounter unfamiliar terminology or recommendations in chatbot outputs.
For developers, the present findings highlight source transparency, information currency, and readable default language as important characteristics of patient-facing responses. Retrieval systems may benefit from prioritizing authoritative oncology, pain-medicine, and palliative-care sources, while visible citations should be relevant to the claims they accompany. Information about uncertainty and circumstances requiring professional assessment can also be integrated into the response structure rather than presented only as a generic disclaimer. Layered presentation may allow brief plain-language information to appear first, with more detailed material available when needed.
Iterative prompting is another relevant feature of conversational AI. The present study intentionally evaluated default single-turn outputs. In practice, users may ask follow-up questions such as “Can you explain that more simply?” or request clarification of unfamiliar terminology. Such interaction may improve understanding and represents an advantage of conversational systems over static webpages. However, users with lower health literacy may be less likely to recognize excessive complexity or request simplification. Accessible default responses therefore remain important before any follow-up interaction. Future studies can compare single-turn and multi-turn workflows and examine whether models retain clinically important information while simplifying language.
Reproducibility also remains important in chatbot research. The same prompt can generate different wording across sessions, and commercial products may change routing, retrieval sources, interface behavior, or other settings over time. Because the present study collected one response for each model-query combination, the findings represent a single-generation snapshot of the observed outputs. Repeated independent generations could characterize generation-to-generation variability more directly. Future studies should also document available model labels, collection dates, access conditions, retrieval settings, and other relevant configuration details to improve reproducibility.
Future evaluation should also incorporate direct patient involvement. Expert ratings and formula-based readability cannot fully determine whether information is understood or whether users can identify appropriate next steps. Think-aloud interviews, comprehension testing, and patient co-design could identify confusing terminology, misunderstanding of uncertainty, or difficulties recognizing warning information. Guideline-based expert review could additionally assess factual accuracy, treatment appropriateness, medication safety, and escalation advice. Together, these approaches would extend the present quality and readability framework toward a more complete evaluation of patient-facing cancer pain information.
Limitations
Several limitations affect interpretation. First, the study included only four chatbots and nine English-language queries, providing a limited matched sample of cancer-pain searches. The queries were derived from Google Trends and therefore reflect relative public search interest rather than a defined sample of questions submitted specifically by patients or caregivers. Google Trends does not provide demographic or clinical information about individual searchers.
Second, each query was submitted once to each chatbot without clarifying dialogue. Generative systems can produce different outputs in response to the same prompt, and the present study therefore represents a single-generation snapshot rather than an estimate of run-to-run variability.
Third, several instruments were developed for static written health information and assess specific constructs. DISCERN was designed for treatment-choice information, whereas some retained queries focused primarily on symptom recognition. JAMA assesses visible source-transparency characteristics, and formula-based readability estimates linguistic complexity rather than patient comprehension. These measures provide structured comparisons of selected information characteristics but do not encompass all aspects of conversational AI.
Fourth, model outputs and user-facing products may change over time. The present findings reflect the systems and access conditions used on July 22, 2026, and subsequent changes in models, retrieval systems, interfaces, or product settings may alter the characteristics of generated responses.
Finally, sentence-level factual accuracy, guideline concordance, treatment appropriateness, and clinical safety were not directly assessed. Future studies can complement information-quality, transparency, and readability measures with guideline-based expert evaluation and direct assessment of patient comprehension.
Conclusion
Among the four chatbots evaluated, performance varied across quality, transparency, educational presentation, and readability. Copilot had higher EQIP scores than ChatGPT, Gemini, and Perplexity, while Copilot and Perplexity had higher JAMA scores than Gemini. DISCERN did not differ significantly across models, and the median readability estimates for all models remained above the prespecified grade-level benchmarks, with FRES medians below 80. Future work should combine information-quality and readability assessment with guideline-based evaluation of factual accuracy and clinical safety.
Acknowledgments
The authors have no acknowledgments to report.
Abbreviations
AI, artificial intelligence; ARI, Automated Readability Index; CLI, Coleman–Liau Index; EQIP, Ensuring Quality Information for Patients; FRES, Flesch Reading Ease Score; FKGL, Flesch–Kincaid Grade Level; GFI, Gunning Fog Index; GQS, Global Quality Score; HHS, United States Department of Health and Human Services; IQR, interquartile range; JAMA, Journal of the American Medical Association; LLM, large language model; MeSH, Medical Subject Headings; NIH, National Institutes of Health; SMOG, Simple Measure of Gobbledygook.
Disclosure
The authors report no conflicts of interest in this work.
References
- 1.Sung H, Ferlay J, Siegel RL, et al. Global cancer statistics 2020: GLOBOCAN estimates of incidence and mortality worldwide for 36 cancers in 185 countries. CA Cancer J Clin. 2021;71(3):209–14. doi: 10.3322/caac.21660 [DOI] [PubMed] [Google Scholar]
- 2.van den Beuken-van Everdingen MHJ, Hochstenbach LMJ, Joosten EAJ, Tjan-Heijnen VCG, Janssen DJA. Update on prevalence of pain in patients with cancer: systematic review and meta-analysis. J Pain Symptom Manage. 2016;51(6):1070–1090.e9. doi: 10.1016/j.jpainsymman.2015.12.340 [DOI] [PubMed] [Google Scholar]
- 3.Snijders RAH, Brom L, Theunissen M, van den Beuken-van Everdingen MHJ. Update on prevalence of pain in patients with cancer 2022: a systematic literature review and meta-analysis. Cancers. 2023;15(3):591. doi: 10.3390/cancers15030591 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 4.Fallon M, Giusti R, Aielli F, et al. Management of cancer pain in adult patients: ESMO clinical practice guidelines. Ann Oncol. 2018;29(Suppl 4):iv166–iv191. doi: 10.1093/annonc/mdy152 [DOI] [PubMed] [Google Scholar]
- 5.Bennett MI, Rayment C, Hjermstad M, Aass N, Caraceni A, Kaasa S. Prevalence and aetiology of neuropathic pain in cancer patients: a systematic review. Pain. 2012;153(2):359–365. doi: 10.1016/j.pain.2011.10.028 [DOI] [PubMed] [Google Scholar]
- 6.Evenepoel M, Haenen V, De Baerdemaecker T, et al. Pain prevalence during cancer treatment: a systematic review and meta-analysis. J Pain Symptom Manage. 2022;63(3):e317–e335. doi: 10.1016/j.jpainsymman.2021.09.011 [DOI] [PubMed] [Google Scholar]
- 7.World Health Organization. WHO Guidelines for the Pharmacological and Radiotherapeutic Management of Cancer Pain in Adults and Adolescents. Geneva: World Health Organization; 2018. [PubMed] [Google Scholar]
- 8.Wu W, Graziano T, Salner A, et al. Acceptability, effectiveness, and roles of mHealth applications in supporting cancer pain self-management: integrative review. JMIR Mhealth Uhealth. 2024;12:e53652. doi: 10.2196/53652 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 9.Sun J, Zhang S, Hou M, et al. Who can help me? Understanding the antecedent and consequence of medical information seeking behavior in the era of big data. Front Public Health. 2023;11:1192405. doi: 10.3389/fpubh.2023.1192405 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 10.Swire-Thompson B, Lazer D. Public health and online misinformation: challenges and recommendations. Annu Rev Public Health. 2020;41(1):433–451. doi: 10.1146/annurev-publhealth-040119-094127 [DOI] [PubMed] [Google Scholar]
- 11.Noorbakhsh-Sabet N, Zand R, Zhang Y, Abedi V. Artificial intelligence transforms the future of health care. Am J Med. 2019;132(7):795–801. doi: 10.1016/j.amjmed.2019.01.017 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 12.Ayers JW, Poliak A, Dredze M, et al. Comparing physician and artificial intelligence chatbot responses to patient questions posted to a public social media forum. JAMA Intern Med. 2023;183(6):589–596. doi: 10.1001/jamainternmed.2023.1838 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 13.Singhal K, Azizi S, Tu T, et al. Large language models encode clinical knowledge. Nature. 2023;620(7972):172–180. doi: 10.1038/s41586-023-06291-2 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 14.Lee P, Bubeck S, Petro J. Benefits, limits, and risks of GPT-4 as an AI chatbot for medicine. N Engl J Med. 2023;388(13):1233–1239. doi: 10.1056/NEJMsr2214184 [DOI] [PubMed] [Google Scholar]
- 15.Musheyev D, Pan A, Loeb S, Kabarriti AE. How well do artificial intelligence chatbots respond to the top search queries about urological malignancies? Eur Urol. 2024;85(1):13–16. doi: 10.1016/j.eururo.2023.07.004 [DOI] [PubMed] [Google Scholar]
- 16.Yıldız HA, Söğütdelen E. AI chatbots as sources of STD information: a study on reliability and readability. J Med Syst. 2025;49(1):43. doi: 10.1007/s10916-025-02178-z [DOI] [PMC free article] [PubMed] [Google Scholar]
- 17.Cao J, Wu Z, Wu Z, et al. The reliability and readability of large language models in answering patient questions on maintenance hemodialysis: a comparative study. Digit Health. 2026;12:1–11. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 18.Hancı V, Ergün B, Gül Ş, et al. Assessment of readability, reliability, and quality of ChatGPT, Bard, Gemini, Copilot, and Perplexity responses on palliative care. Medicine. 2024;103(33):e39305. doi: 10.1097/MD.0000000000039305 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 19.Ozduran E, Hancı V, Erkin Y, et al. Readability, reliability and quality of responses generated by ChatGPT, Gemini, and Perplexity for the most frequently asked questions about pain. Medicine. 2025;104(104):e41780. doi: 10.1097/MD.0000000000041780 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 20.Ozduran E, Hancı V, Erkin Y, et al. Assessing the readability, quality and reliability of responses produced by ChatGPT, Gemini, and Perplexity regarding most frequently asked keywords about low back pain. PeerJ. 2025;13:e18847. doi: 10.7717/peerj.18847 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 21.Meyer MKR, Kandathil CK, Davis SJ, et al. Evaluation of rhinoplasty information from ChatGPT, Gemini, and Claude for readability and accuracy. Aesthetic Plast Surg. 2025;49(7):1868–1873. doi: 10.1007/s00266-024-04343-0 [DOI] [PubMed] [Google Scholar]
- 22.Şahin MF, Ateş H, Keleş A, et al. Responses of five different artificial intelligence chatbots to the top searched queries about erectile dysfunction: a comparative analysis. J Med Syst. 2024;48(1):38. doi: 10.1007/s10916-024-02056-0 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 23.Swisher AR, Wu AW, Liu GC, et al. Enhancing health literacy: evaluating the readability of patient handouts revised by ChatGPT’s large language model. Otolaryngol Head Neck Surg. 2024;171(6):1751–1757. doi: 10.1002/ohn.927 [DOI] [PubMed] [Google Scholar]
- 24.Ellison IE, Oslock WM, Abdullah A, et al. De novo generation of colorectal patient educational materials using large language models: prompt engineering key to improved readability. Surgery. 2025;180:109024. doi: 10.1016/j.surg.2024.109024 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 25.Zaki HA, Mai M, Abdel-Megid H, et al. Using ChatGPT to improve readability of interventional radiology procedure descriptions. Cardiovasc Intervent Radiol. 2024;47(8):1134–1141. doi: 10.1007/s00270-024-03803-z [DOI] [PubMed] [Google Scholar]
- 26.Berkman ND, Sheridan SL, Donahue KE, Halpern DJ, Crotty K. Low health literacy and health outcomes: an updated systematic review. Ann Intern Med. 2011;155(2):97–107. doi: 10.7326/0003-4819-155-2-201107190-00005 [DOI] [PubMed] [Google Scholar]
- 27.DeWalt DA, Berkman ND, Sheridan S, Lohr KN, Pignone MP. Literacy and health outcomes: a systematic review. J Gen Intern Med. 2004;19(12):1228–1239. doi: 10.1111/j.1525-1497.2004.40153.x [DOI] [PMC free article] [PubMed] [Google Scholar]
- 28.Wolf MS, Gazmararian JA, Baker DW. Health literacy and functional health status among older adults. Arch Intern Med. 2005;165(17):1946–1952. doi: 10.1001/archinte.165.17.1946 [DOI] [PubMed] [Google Scholar]
- 29.Weiss BD. Health Literacy and Patient Safety: Help Patients Understand. 2nd ed. Chicago: American Medical Association Foundation; 2007. [Google Scholar]
- 30.Eltorai AEM, Ghanian S, Adams CA, Born CT, Daniels AH. Readability of patient education materials on the American association for surgery of trauma website. Arch Trauma Res. 2014;3(1):e18161. doi: 10.5812/atr.18161 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 31.Rooney MK, Santiago G, Perni S, et al. Readability of patient education materials from high-impact medical journals: a 20-year analysis. J Patient Exp. 2021;8:2374373521998847. doi: 10.1177/2374373521998847 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 32.Charnock D, Shepperd S, Needham G, Gann R. DISCERN: an instrument for judging the quality of written consumer health information on treatment choices. J Epidemiol Community Health. 1999;53(2):105–111. doi: 10.1136/jech.53.2.105 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 33.Moult B, Franck LS, Brady H. Ensuring quality information for patients: development and preliminary validation of a new instrument to improve the quality of written health care information. Health Expect. 2004;7(2):165–175. doi: 10.1111/j.1369-7625.2004.00273.x [DOI] [PMC free article] [PubMed] [Google Scholar]
- 34.Silberg WM, Lundberg GD, Musacchio RA. Assessing, controlling, and assuring the quality of medical information on the Internet: caveant lector et viewor. Jama. 1997;277(15):1244–1245. doi: 10.1001/jama.1997.03540390074039 [DOI] [PubMed] [Google Scholar]
- 35.Bernard A, Langille M, Hughes S, Rose C, Leddin D, Veldhuyzen van Zanten S. A systematic review of patient inflammatory bowel disease information resources on the world wide web. Am J Gastroenterol. 2007;102(9):2070–2077. doi: 10.1111/j.1572-0241.2007.01325.x [DOI] [PubMed] [Google Scholar]
- 36.CHART Collaborative. Reporting guideline for chatbot health advice studies: the Chatbot assessment reporting tool (CHART) statement. BMJ Med. 2025;4(1):e001632. doi: 10.1136/bmjmed-2025-001632 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 37.Mavragani A, Ochoa G. Google trends in infodemiology and infoveillance: methodology framework. JMIR Public Health Surveill. 2019;5(2):e13439. doi: 10.2196/13439 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 38.Rovetta A. Reliability of Google Trends: analysis of the limits and potential of web infoveillance during the COVID-19 pandemic and for future research. Front Res Metr Anal. 2021;6:670226. doi: 10.3389/frma.2021.670226 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 39.Alibudbud R. Google Trends for health research: advantages, applications, methodological considerations, and limitations in psychiatric and mental health infodemiology. Front Big Data. 2023;6:1132764. doi: 10.3389/fdata.2023.1132764 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 40.Smith EA, Senter RJ. Automated readability index. AMRL-TR-66-220. Wright-Patterson air force base. OH: Aerospace Medical Research Laboratories; 1967. [PubMed] [Google Scholar]
- 41.Kincaid JP, Fishburne RP, Rogers RL, Chissom BS. Derivation of new readability formulas for Navy enlisted personnel. Research branch report 8-75. Millington, TN: Naval Technical Training Command; 1975. [Google Scholar]
- 42.Coleman M, Liau TL. A computer readability formula designed for machine scoring. J Appl Psychol. 1975;60(2):283–284. doi: 10.1037/h0076540 [DOI] [Google Scholar]
- 43.Hedman AS. Using the SMOG formula to revise a health-related document. Am J Health Educ. 2008;39(1):61–64. doi: 10.1080/19325037.2008.10599016 [DOI] [Google Scholar]
- 44.Wang LW, Miller MJ, Schmitt MR, Wen FK. Assessing readability formula differences with written health information materials: application, results, and recommendations. Res Social Adm Pharm. 2013;9(5):503–516. doi: 10.1016/j.sapharm.2012.05.009 [DOI] [PubMed] [Google Scholar]
- 45.Garcia Valencia OA, Thongprayoon C, Miao J, et al. Empowering inclusivity: improving readability of living kidney donation information with ChatGPT. Front Digit Health. 2024;6:1366967. doi: 10.3389/fdgth.2024.1366967 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 46.Mondal H, Gupta G, Sarangi PK, et al. Assessing the capability of large language model chatbots in generating plain-language summaries. Cureus. 2025;17(3):e80976. doi: 10.7759/cureus.80976 [DOI] [PMC free article] [PubMed] [Google Scholar]
