Table 2.
Scales used for evaluation of reliability, quality, usefulness, and readability of ChatGPT-5.
| Scale, items, and scoring criteria |
|---|
| Modified DISCERN tool |
| *1 point is given for every yes, with a maximum number of 5 points achievable 1. Are the aims clear and achieved? 2. Are reliable sources of information used? (i.e., publication cited, the responses are from valid studies/sources) 3. Is the information presented by the AI softwares balanced and unbiased? 4. Are additional sources of information listed for patient reference? 5. Does the responses by the AI software address areas of uncertainty? |
| Global Quality Scale |
| *scored based on the following characteristics 1. Poor quality, poor flow of the site, most information missing, not at all useful for patients 2. Generally poor quality and poor flow, some information listed but many important topics missing, of very limited use to patients 3. Moderate quality, suboptimal flow, some important information is adequately discussed but others poorly discussed, somewhat useful for patients 4. Good quality and generally good flow. Most of the relevant information is listed, but some topics not covered, useful for patients 5. Excellent quality and flow, very useful for patients |
| Usefulness Score |
| *scored based on the following characteristics 1. Not useful at all: Unintelligible language, contradictory information and missing important information. Not useful for patients. 2. Very little useful: Partly clear language is used. Some important information is missing or incorrect. For patients limited use possible. 3. Relatively useful: Clear language is used. Most important information is mentioned, but some important information incomplete or incorrect. 4. Partly useful: Clear language is used. Some important information is missing or incorrect, but most important information is addressed. 5. Moderately useful: Clear language is used and most important information is covered, but some important information is still incomplete or incorrect. 6. Very useful: Clear language is used. All important information is mentioned, but some unimportant information or details are also mentioned. 7. Extremely useful: Clear language is used and all important information is mentioned. Extremely useful to patients, additional information and resources are also provided. |
| Flesch Reading Ease Reading Level (Estimated reading grade level) |
| Very easy (US Grade 5 or 11-year-old; 91–100 scores) Easy (US Grade 6; 81–90 scores) Fairly easy (US Grade 7; 71–80 scores) Standard (US Grade 8–9 or 13–15 years; 61–70 scores) Fairly difficult (US Grade 10–12; 51–60 scores) Difficult (US Grade 13–16; 31–50 scores) Very difficult (College graduate; 0–30 scores) |