Skip to main content
Springer logoLink to Springer
. 2026 Apr 9;71(9):4046–4055. doi: 10.1007/s10620-026-09879-6

The Impact of Specific Prompt Engineering Techniques on the Readability of LLM-Generated Patient Materials in Gastroenterology and Hepatology

Husayn F Ramji 1, Aishwarya Gatiganti 1, Anveet Janwadkar 1, Jacob Lampenfeld 1, Ilaria M Simeone 1, Abhijith Atkuru 1, Jason Mathias 1, Stephanie Mrowczynski 1, Sharan Poonja 1, Chandler Gilliard 1, Shaquille Lewis 1, Pooja Arumugam 1, Clara Freedman 1, Luis Morales 1, Everette Martin III 1, Larry Zhou 1, Corinne Zalomek 1, Devika Dixit 2, Saba Abdulsada 2, Matthew Houle 2, Matthew Alias 2, Molly Delk 2,3, Sarah C Glover 2,3, Peng-Sheng Ting 2,3,✉
PMCID: PMC13522015  PMID: 41954673

Abstract

Background and Aims

Health literacy significantly impacts patient outcomes. While the average American reads at an 8th grade reading level, healthcare materials are often written above this, potentially contributing to our nation’s health literacy gap. Large language models (LLM) and prompt engineering may be able to help address this gap by consistently generating materials at the recommended 6th grade reading level. Our study aims to assess how different prompting techniques affect the readability of LLM-generated materials.

Methods

We assessed the effects of five prompt techniques (Zero-Shot, Contextualized, Constrained, Meta, Persona) on the readability of LLM-generated explanations for fifteen common gastroenterology and hepatology conditions across twelve LLMs. Output (n = 2655) readability was assessed with two readability metrics (Simple Measure of Gobbledygook (SMOG) index, Flesch-Kincaid Grade Level (FKGL)), followed by significance testing and post-hoc analysis.

Results

No prompt technique or model consistently produced outputs at or below a 6th grade reading level when assessed by the SMOG index, the preferred metric when assessing healthcare materials (p < 0.001). However, prompts with less constraints yield less readable outputs, while prompts with more constraints yield significantly more readable outputs (p < 0.001).

Conclusions

This study demonstrates the potential of LLMs as a tool in addressing America’s health literacy gap, as we show prompt engineering affects the readability of gastroenterology and hepatology-related explanations. We also found limitations to this technique. Further optimization is necessary before LLMs can consistently generate patient materials without appropriate clinician oversight, but it implies prompt engineering as a tool in addressing our nation’s health literacy gap.

Keywords: Artificial intelligence, LLMs, Health literacy, Patient education materials, Prompt engineering

Introduction

Patient outcomes are influenced by a number of factors. Some are objective and can be measured and quantified. Others are influenced by factors that occur outside the clinic and cannot be influenced by practitioners themselves, such as the social determinants of health (SDOH) [1]. However, a patient’s personal health literacy is a unique modifiable risk factor and SDOH that clinicians can directly influence and address [2, 3]. Health literacy is defined as the “degree to which individuals have the capacity to obtain, process, and understand basic health information and services needed to make appropriate health decisions” [2, 4–6]. More than 80 million Americans demonstrate low levels of health literacy [2]. Low health literacy itself is associated with increased risk of death, hospitalization, and, for patients with gastroenterological conditions, poorer prognoses, longer lengths of stay, and higher readmission rates [2, 7–10]. Lower health literacy also affects the effectiveness of bowel preparations for colonoscopies, potentially affecting the quality of procedures themselves [11–13]. Although patient educational materials (PEMs) are crucial tools in addressing this, they are often written at a grade level higher than what is recommended by the Agency for Healthcare Research and Quality, contributing to a gap between patient and provider, a “health literacy gap” [14].

In order to help address the nation’s health literacy gap, the Agency for Healthcare Research and Quality (AHRQ) and the National Institutes of Health (NIH) both recommend that all patient-facing healthcare materials be written at a 5th to 6th grade reading level [15, 16]. This is intentionally lower than the 8th grade reading level that the average American reads at in order to reach a greater number of individuals [17]. The vast majority of patient education and healthcare materials do not meet this readability recommendation, even when assessed with the appropriate metrics [14, 18–22]. This includes the preferred and validated readability metric for assessing healthcare materials, the Simple Measure of Gobbledygook (SMOG) index, and the more widely used Flesch-Kincaid Grade Level (FKGL) [14, 18–24]. However, with the arrival of new tools like generative artificial intelligence (AI), PEMs may be generated within seconds, demonstrating the potential to help bridge the gap between medical jargon and the differing levels of health literacy among our patients, if the outputs are written at the appropriate reading level [14, 25]. However, there are a number of factors that influence this too.

The most widely available AI tools at this time are large language models (LLMs), or AI-powered tools that can “analyze and generate…textual data” [25, 26]. These include OpenAI’s ChatGPT and Google’s Gemini, and they are already demonstrating numerous applications in the field of medicine [25–28]. As their names suggest, the specific language used in “prompting” these LLMs can greatly influence the quality, tone, and even readability of their outputs [26, 28]. In contrast to search engines, LLMs can present summarized information as a document, but it depends on the prompt. This had led to the emergence of the field of study called “prompt engineering” [26]. Prompt engineering is defined as the “practice of designing, refining, and implementing prompts or instructions that guide the outputs of LLMs to help in various tasks” [26]. There are many types of prompts, such as “Zero-Shot,” “Constrained,” “Contextual”, “Role-based”, and “Meta” [29, 30]. “Zero-Shot” prompts are the simplest queries, such as “What is GERD?” where the user provides no guidance or constraint in how the output should be formatted, written, or what to include in an output [29, 30]. Zero-shot prompts most closely resemble traditional search engine inquiries. This is in comparison to a “Few-Shot” prompt where the user would provide an example of output [29, 30]. Meanwhile, a “Meta” prompt is one written by an LLM itself at a user’s request and often contains a number of constraints [29, 30]. Although there are multiple publications discussing the effect of prompting on the accuracy of outputs, there are a limited number of studies investigating prompt engineering in context of readability of patient educational materials (PEMs) [25, 26, 31–36].

Our study aims to assess how different prompting techniques across several LLMs affect the readability of LLM-generated PEMs in the fields of gastroenterology and hepatology, in order to assess the ability of LLMs to be used as a tool in addressing the health literacy gap through comparisons of prompting techniques.

Methods

Study Design

This cross-sectional study utilized twelve LLMs (OpenAI ChatGPT 4o, ChatGPT 5.0, Anthropic Claude 4 Sonnet, Claude 4.1 Opus, Microsoft Copilot Quick, Copilot Deep, Copilot Smart (GPT-5), Doximity GPT, Google Gemini 2.5 Flash, Gemini 2.5 Pro, Meta AI Llama 4, and OpenEvidence) that were given five different types of prompts—Zero-Shot, Contextualized, Constrained, Meta, and Role-Based/Persona (Table 1). The Meta prompt was created by ChatGPT 4o in response to the request, “Provide a prompt that explains medical conditions to patients at a 6th grade reading level or below.” The Persona prompt was created by ChatGPT 4o as a modified Meta prompt. The five prompts were then utilized for fifteen highly prevalent conditions, with three trials per condition, or a total of 225 outputs per LLM (Table 1) [37–40]. However, OpenEvidence was unable to provide outputs for Persona prompt as it was “outside of scope” for the LLM, leaving a total of 2655 outputs instead of the projected 2700.

Table 1.

Conditions and prompts

Conditions asked about:
GERD, Peptic Ulcer Disease, Diverticulosis, Colorectal Cancer, IBD, Crohn's, Ulcerative Colitis, Cirrhosis, Hepatocellular Carcinoma, Hepatitis, Celiac Disease, Hemorrhoids, IBS, Gallstones, Cholecystitis
Example prompt 1 (Zero-shot prompt):
What is GERD?
Example prompt 2 (Contextualized prompt):
I am a patient that was just diagnosed with GERD. Explain that to me in simple terms
Example prompt 3 (Constrained prompt):
Explain GERD to a patient at a 6th grade reading level or below
Example prompt 4 (Meta prompt):
Explain the medical condition GERD to an adult audience using simple, easy-to-understand language at a 6th grade reading level or lower. Avoid medical jargon, use short sentences, and give clear examples or comparisons to help understanding. Make the explanation friendly and informative, as if you are talking to someone who is new to the topic
Example prompt 5 (Role-based/Persona prompt):
You are a compassionate and knowledgeable gastroenterologist. Your job is to clearly and kindly explain GERD to an adult patient using simple, everyday language—at or below a 6th grade reading level. Avoid medical jargon unless it’s essential, and always explain terms in a way that’s easy-to-understand. Your tone should be friendly, respectful, and reassuring, helping the patient feel informed and supported

Outcomes and Statistical Analysis

Readability metrics (SMOG, FKGL) of the outputs were calculated using the Python textstat library, with lower grade level values indicating improved readability [34–36, 41, 42]. The Python program integrating the open-source textstat Python library was written with assistance from ChatGPT 4o and allowed for scoring of multiple outputs simultaneously. One-way ANOVA was utilized to analyze Across-Prompt variations, with p < 0.05 being deemed statistically significant. Statistical and post-hoc analyses, including pairwise analysis, were conducted in Microsoft Excel with assistance from Google Gemini 2.5 Pro.

Results

Readability of outputs improved for all LLMs across prompts, but no prompt or model was able to consistently deliver information at a 6th grade reading level or below when assessed by the SMOG (Table 2). However, eleven of the twelve LLMs were able to average a reading level within at least one standard deviation of the 6th grade for at least one prompt when scored with the FKGL. The different prompting techniques showed varying degrees of improvement in readability, while LLM performance also varied depending on the prompt.

Table 2.

Average summary statistics by prompt

Prompt SMOG
(Mean ± SD)
Flesch-kincaid
(Mean ± SD)
Avg. # of Words Avg. # of Characters
Zero-shot 15.85 ± 2.74 15.32 ± 3.86 272.36 1888.11
Contextualized 11.8 ± 1.45 10.14 ± 1.95 259.08 1599.69
Constrained 9.38 ± 1.76 7.77 ± 2.95 244.92 1399.21
Meta 8.75 ± 1.36 6.67 ± 1.92 307.43 1764.86
Persona 9.34 ± 1.1 7.27 ± 1.62 379.62 2170.06

The Zero-Shot prompt was the simplest prompt utilized. As expected, it had the highest grade level readability metrics, with an average SMOG of 15.85 ± 2.74, and an average FKGL of 15.32 ± 3.86, suggesting a college-level and even graduate-level education is needed to adequately understand the provided information (Table 2, Figs. 1, 2). The Contextualized prompt saw improved readability, but no output was below a 10th grade reading level by the SMOG index, and an average FKGL of 10.14 ± 1.95. The Constrained prompt saw a dramatic improvement in readability for all metrics. Although the average SMOG was 9.38 ± 1.76, 6 of the 12 tested LLMs were either at or within one standard deviation of 6th grade when assessed with the FKGL, with an average of 7.77 ± 2.95. This pattern followed for the Meta prompt. The Meta prompt had the lowest readability metrics overall, with an average SMOG of 8.75 ± 1.36 and FKGL of 6.67 ± 1.92. In fact, 10 of the 12 LLMs tested were either at or within one standard deviation of a 6th- grade reading level with the Meta prompt when assessed by the FKGL (Table 4). When compared to the Meta prompt, the Persona prompt saw a small but significant increase in the average readability for both SMOG and FKGL, with only 4 of the 11 LLMs at goal with the FKGL.

Fig. 1.

Fig. 1

Box-plots by prompt, SMOG, displaying median, Q1, Q3

Fig. 2.

Fig. 2

Box-plots by prompt, FKGL, displaying median, Q1, Q3

Table 4.

Summary statistics by model

Model Prompt Avg. SMOG Avg. Flesch-Kincaid Grade Avg. # Characters Avg. # Words Avg. # Sentences
ChatGPT 4o Zero-shot 15.48 ± 2.54 14.91 ± 3.20 1528.78 217.36 11.64
Contextualized 11.00 ± 1.31 9.49 ± 2.31 1328.13 219.67 12.71
Constrained 8.61 ± 1.03 6.25 ± 1.19 1146.11 204.09 14.56
Meta 9.00 ± 0.92 6.91 ± 1.20 2085.84 360.22 23.84
Persona 9.26 ± 0.75 7.08 ± 0.95 2247.93 384.96 25.18
ChatGPT 5 Zero-shot 15.30 ± 1.98 14.83 ± 2.81 1366.38 204.73 8.84
Contextualized 11.68 ± 1.26 10.32 ± 1.78 1644.67 264.27 14.42
Constrained 9.54 ± 1.32 7.74 ± 1.94 1392.87 237.47 15.67
Meta 8.91 ± 1.01 6.74 ± 1.61 1862.51 318.69 23.6
Persona 9.94 ± 0.94 8.29 ± 1.69 2137.11 356.69 22.44
Claude 4 Sonnet Zero-shot 15.36 ± 1.41 14.33 ± 1.87 1453.73 211.8 10.58
Contextualized 12.33 ± 1.01 10.65 ± 1.19 1432.98 229.36 12.16
Constrained 8.20 ± 1.02 5.96 ± 1.16 1300.04 230.36 16.69
Meta 8.58 ± 1.39 6.37 ± 1.53 1906.24 326.98 25.2
Persona 9.25 ± 1.06 7.15 ± 1.14 1655.53 286.22 19.07
Claude 4.1 Opus Zero-shot 16.49 ± 1.98 16.16 ± 2.60 1831.64 261.44 11.24
Contextualized 12.63 ± 1.12 11.08 ± 1.40 1581.18 254.22 13
Constrained 7.81 ± 1.12 5.69 ± 1.95 1551.73 277.64 21.47
Meta 7.91 ± 0.83 5.71 ± 0.98 1958.84 343.27 27.89
Persona 9.46 ± 0.97 7.49 ± 1.20 2958 514.58 31.27
Copilot Quick Zero-shot 16.09 ± 3.36 15.98 ± 4.66 1727.71 233.22 10.51
Contextualized 10.59 ± 1.04 8.55 ± 1.38 1532.36 249.33 15.76
Constrained 9.59 ± 1.04 8.16 ± 1.04 1375.78 235.04 12.71
Meta 7.83 ± 0.84 5.40 ± 1.05 1536.84 273.53 22.04
Persona 8.65 ± 0.65 6.46 ± 0.98 1875.11 326.07 21.87
Copilot Deep Zero-shot 18.49 ± 3.60 19.05 ± 5.57 2387.67 320.67 12.49
Contextualized 13.25 ± 1.81 12.24 ± 2.73 2106.42 321.36 16.42
Constrained 12.07 ± 2.44 13.56 ± 4.87 1484.16 256.13 9.4
Meta 8.18 ± 1.60 6.80 ± 3.31 1290.13 225.62 16.36
Persona 9.02 ± 1.18 7.06 ± 1.86 1717.22 299.11 19.8
Copilot Smart Zero-shot 15.12 ± 2.35 14.92 ± 3.53 1494 216.29 9.29
Contextualized 11.04 ± 1.30 9.60 ± 1.68 1422.82 234.16 12.24
Constrained 10.04 ± 1.23 9.14 ± 2.21 1306.44 227.36 12.04
Meta 10.19 ± 1.28 9.24 ± 1.37 1649.4 286.93 14.4
Persona 10.09 ± 0.96 9.23 ± 2.07 1789.2 309.31 15.71
Doximity GPT Zero-shot 15.14 ± 1.81 14.13 ± 2.33 931.36 139.07 7.16
Contextualized 12.43 ± 1.41 10.75 ± 1.40 652.02 104.2 5.98
Constrained 8.88 ± 1.00 6.75 ± 1.08 835.89 150.67 10.22
Meta 8.37 ± 0.89 6.05 ± 1.01 1244.87 221.84 17.13
Persona 9.02 ± 0.97 7.01 ± 0.96 1316.87 231.42 15.4
Gemini 2.5 Pro Zero-shot 14.68 ± 1.06 13.31 ± 1.44 4880.02 719.38 40.58
Contextualized 11.10 ± 0.78 8.69 ± 0.97 3185.09 529.84 34
Constrained 8.09 ± 0.73 5.76 ± 0.69 2367.49 433.22 29.71
Meta 8.52 ± 0.55 5.97 ± 0.60 2508.4 444.22 35.27
Persona 8.53 ± 0.49 5.58 ± 0.42 3913.18 717.58 54.78
Gemini 2.5 Flash Zero-shot 13.72 ± 1.27 12.31 ± 1.73 2624.62 405.22 21.96
Contextualized 11.01 ± 0.73 8.75 ± 0.72 1792.49 301.24 19.07
Constrained 8.67 ± 1.07 6.44 ± 1.25 1534.31 270.02 20.53
Meta 7.88 ± 0.90 5.16 ± 0.93 1733.11 312.47 27.16
Persona 8.82 ± 0.84 6.31 ± 0.76 2613.27 469.42 32.44
Llama 4 Zero-shot 15.17 ± 3.32 14.98 ± 5.49 786.93 118.6 5.49
Contextualized 12.76 ± 1.29 11.26 ± 1.86 819.93 127.89 7.29
Constrained 10.89 ± 1.76 9.36 ± 2.38 981.24 164.36 9.53
Meta 10.56 ± 1.42 8.81 ± 1.74 1276.29 214.84 13.18
Persona 10.66 ± 1.00 8.32 ± 1.38 1647.22 280.42 17.82
OpenEvidence Zero-shot 19.11 ± 1.36 18.92 ± 1.85 1644.44 220.53 9.47
Contextualized 11.73 ± 0.93 10.27 ± 1.00 1698.24 273.47 14.84
Constrained 10.11 ± 0.91 8.40 ± 0.98 1514.44 252.69 15.69
Meta 9.06 ± 0.77 6.92 ± 1.03 2125.8 360.49 26.76

(bold indicates it was either at or within one standard deviation of 6.00 grade reading level)

In Across-Prompt analysis, or comparing prompts to one another, the Meta prompt had significantly more readable outputs, even when compared to the Persona prompt (p < 0.001) (Table 3). The difference in FKGL between all prompt comparisons was statistically significant (p < 0.05). The difference in SMOG between prompt comparisons was statistically significant (p < 0.001), with the exception of the Constrained prompt versus the Persona prompt. There was no statistically significant difference between the output SMOG readability of the Constrained prompt and the Persona prompt (p = 0.996). The largest difference in mean readability existed between the Zero-Shot and Meta prompt.

Table 3.

Prompt pairwise analysis

Prompt 1 (P1) Prompt 2 (P2) SMOG mean difference (P1-P2) SMOG
p-value
FKGL mean difference (P1-P2) FKGL
p-value
Zero-shot Contextualized  − 4.05 p < 0.001  − 5.18 p < 0.001
Zero-shot Constrained  − 6.47 p < 0.001  − 7.55 p < 0.001
Zero-shot Meta  − 7.1 p < 0.001  − 8.65 p < 0.001
Zero-shot Persona  − 6.51 p < 0.001  − 8.05 p < 0.001
Contextualized Constrained  − 2.42 p < 0.001  − 2.37 p < 0.001
Contextualized Meta  − 3.05 p < 0.001  − 3.46 p < 0.001
Contextualized Persona  − 2.46 p < 0.001  − 2.87 p < 0.001
Constrained Meta  − 0.63 p < 0.001  − 1.09 p < 0.001
Constrained Persona  − 0.04 p = 0.996  − 0.49 p < 0.05
Meta Persona 0.58 p < 0.001 0.6 p < 0.05

LLM performance varied, with some performing better than others. Claude 4.1 Opus had the lowest overall average SMOG index for the Constrained prompt, with outputs averaging a SMOG of 7.81 ± 1.12 (Table 4). However, most LLMs had their lowest SMOG for the Meta prompt. Gemini 2.5 Flash had the lowest average FKGL for the Meta prompt with an average grade-level of 5.16 ± 0.93. On average, Doximity GPT had the shortest outputs with 996 characters, while Gemini 2.5 Pro averaged the longest outputs with 3371 characters.

Discussion

To our knowledge, this is one of the first studies to test the effects of prompt engineering techniques on the readability of LLM-generated materials in the fields of gastroenterology and hepatology. We found that certain prompt engineering techniques can lead to enhanced readability, suggesting modern LLMs can become valuable tools in creating personalized PEMs, addressing the health literacy gap, if prompted and optimized appropriately with clinician input and oversight. However, more studies are needed to truly understand the intricacies of prompting and their effects on readability, as not all results seen here were expected. Additionally, this study underscores the importance of prompt engineering not only as a skill and area of further research for clinicians, but also as a topic that may become closely related to a patient’s health literacy with wider adoption of LLMs.

Some results were anticipated. For example, the Zero-Shot prompt provided a clear baseline of what the LLMs will produce when not given any direction or constraints. However, the statistically significant difference in readability metrics between the Persona prompt and the Meta prompt was not expected. Published prompting techniques emphasize that the more specific and constrained prompts, such as specifying a role, can enhance the quality of the output [26]. However, even though the Persona prompt had that additional role-based constraint, the outputs were slightly more difficult to read than that of the Meta prompt. Although in a clinical scenario, this subtle decrease in readability may not be perceived by the patient or provider, it is still worth noting. It implies that LLMs themselves may be the best tool in prompt engineering. It also suggests the act of specifying a role to an LLM can lead to additional effects on the output that alter readability, even in the presence of other constraints. However, as both the readability metrics rely on word and sentence count in their calculation, the significant difference in the readability can likely be attributed to the differences in the length of output. The Persona prompt had the lengthiest outputs of any prompt, with the Meta following closely behind.

Alternatively, when comparing the Constrained and Persona prompts, there was no significant difference in readability when assessed by the SMOG, but there was when assessed by the FKGL. This discrepancy likely comes down to the inherent differences in the metrics themselves, and the differences in length in the outputs between the two prompts (Table 2). The SMOG calculation relies heavily on the number of polysyllabic words, while the FKGL gives more weight to the number of sentences and the average number of syllables per word [19, 23, 24]. Thus, although the prompts had a non-significant difference in SMOG, suggesting they both used a similar number of polysyllabic words, the Persona prompt likely used the same number of polysyllabic words but across multiple sentences, resulting in a lower FKGL. When taken in context of assessing complexity of PEMs, this distinction between what influences the metrics is critical. Most medical terms are polysyllabic and mentioning these terms multiple times without replacing them with a simpler definition will still affect complexity and readability of a PEM, regardless of the length of the materials themselves. This is why the SMOG index is the preferred metric for healthcare materials, as it is not influenced by factors like number of sentences, and is more consistent in the assessment of PEMs than the FKGL [19, 23, 24]. Having longer outputs seems to have diluted the complex words in the explanations, which appears to be a reproducible effect of the Persona prompt across LLMs (Table 4). However, given the “black-box” nature of LLMs, or how it is unclear what operations are occurring within an LLM to produce an output, we cannot definitively say what elements of any prompt influenced the outputs [43]. Therefore, more prompt research and transparency is needed to thoroughly understand how to consistently influence an LLM’s output. Additionally, such research may help our patients who turn to LLMs for inquiries.

If a patient or clinician were to use an LLM like a traditional search engine with a Zero-Shot prompt, the LLM may produce materials too difficult to read for the majority of the patient populations we serve, falling victim to the same criticism of some modern-day PEMs. Therefore, prompt engineering may not only be a topic that crosses disciplines, but one that, if shared and discussed with patients, has the potential to improve health literacy. This could take the shape of pre-made templates for patients or even guides on LLMs in general. At this time, for clinicians seeking to create PEMs with LLMs, we recommend they utilize a Meta prompt technique, where the clinicians can utilize their LLM of choice to generate a prompt in response to the request, “Provide a prompt that explains medical conditions to patients at a 6th grade reading level or below” (Fig. 3). The outputs from the Meta prompt can then be edited by the clinician or expanded upon to create a PEM close or at the recommended level. Additionally, we recommend that clinicians and clinical educators discuss LLMs with patients and students to better understand how they use them. This joint dialogue regarding a new and emerging technology would mirror the same discussions that occurred with the arrival of the internet and later social media [44, 45]. The more clinicians can engage patients on how to use these tools, the more likely both parties will be able to avoid their downsides and dangers. In time, LLMs may be able to bridge the health literacy gap with PEMs generated at the appropriate reading level. But for now, it is still up to healthcare providers to recognize the factors that contribute to our patients’ health literacy, and work toward meeting patients where they are.

Fig. 3.

Fig. 3

How to generate and use a Meta Prompt

This study has many limitations. The Python textstat tool uses punctuation to identify the beginning and end of sentences, meaning outputs that included information in lists or bullet-points without punctuation may have been scored incorrectly. However, given the volume of outputs that were assessed, it is unlikely this would have affected our conclusions [41]. In addition, readability metrics take into account the number of syllables in a document in their calculations. If a medical condition is mentioned multiple times in the output, even after a simplified explanation, this could also artificially increase readability scores, making it appear more difficult to read than what patients may perceive when reading it themselves. We did not assess the readability of patient materials provided by professional societies, including their multimedia materials, given the limitations of the Python program created for this study. Additionally, the grade reading level of patient materials is only one aspect of health literacy, and gauging a patient’s understanding requires clinician assessment. Lastly, we did not assess the accuracy of the outputs, as this was the not the goal of our study. However, a weakness inherent to all LLM research is the risk of LLMs “hallucinating” when providing information, or providing inaccurate or completely made-up information, akin to “AI misinformation” [46]. At this time, there is no way to avoid this, further emphasizing the need for clinician oversight when PEMs are created with LLMs.

Conclusion

In summary, this study demonstrates the potential of prompt engineering in creating readable patient education materials and the importance it can play in the health literacy of our patients. It remains crucial for clinicians to review any LLM-generated materials before distributing them to patients. Even with the most constrained prompts, no LLM was able to consistently generate PEMs at or below a sixth-grade reading level, emphasizing the need for future studies focusing on prompt engineering and patient education. Lastly, we highly recommend clinicians start engaging patients on the proper use of LLMs with the appropriate prompt engineering techniques or templates in order to minimize their potential harms, and ensure effective and safe use of these powerful tools moving forward.

Author Contributions

HFR: conceptualization, methodology, software, data curation, formal analysis, writing – original draft, writing – reviewing & editing; AG: methodology, data curation, writing – reviewing & editing; AJ, JL, IMS, AA: data curation, writing – reviewing & editing; JM, SM, SP, CG, SL, PA, CF, LM, EMIII, LZ, CZ, DD, SA, MH, MA: data curation; MD, SCG: methodology, supervision, writing – reviewing & editing; PST: conceptualization, methodology, supervision, writing – reviewing & editing. All authors reviewed the manuscript.

Data Availability

Data is available on request. The Python program and code we utilized is also available on request.

Declarations

Conflict of interest

The authors declare no competing interests.

Footnotes

Publisher's Note

Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.

References

  • 1.Braveman P, Gottlieb L. The social determinants of health: it’s time to consider the causes of the causes. Public Health Rep 2014;129:19–31. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 2.Berkman ND, Sheridan SL, Donahue KE, Halpern DJ, Crotty K. Low health literacy and health outcomes: an updated systematic review. Ann Intern Med 2011;155:97–107. [DOI] [PubMed] [Google Scholar]
  • 3.Tormey LK, Farraye FA, Paasche-Orlow MK. Understanding Health Literacy and its Impact on Delivering Care to Patients with Inflammatory Bowel Disease. Inflamm Bowel Dis 2016;22:745–751. [DOI] [PubMed] [Google Scholar]
  • 4.Institute of Medicine Committee on Health, L., in Health Literacy: A Prescription to End Confusion, L. Nielsen-Bohlman, A.M. Panzer, and D.A. Kindig, Editors. 2004, National Academies Press (US). Copyright 2004 by the National Academy of Sciences. All rights reserved.: Washington (DC). [PubMed]
  • 5.Coughlin SS, Vernon M, Hatzigeorgiou C, George V. Health Literacy, Social Determinants of Health, and Disease Prevention and Control. J Environ Health Sci 2020;6:3061. [PMC free article] [PubMed] [Google Scholar]
  • 6.Brach C, Harris LM. Healthy People 2030 Health Literacy Definition Tells Organizations: Make Information and Services Easy to Find, Understand, and Use. J Gen Intern Med 2021;36:1084–1085. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 7.Kim, D.H., O’Connor, S.J., Williams, J.H., Opoku-Agyeman, W., Chu, D.I. and Choi, S.W., 2021. The effect of gastrointestinal patients’ health literacy levels on gastrointestinal patients’ health outcomes. Journal of Hospital Management and Health Policy, 5.
  • 8.Fabbri, M., Yost, K., Rutten, L.J.F., Manemann, S.M., Boyd, C.M., Jensen, D.,et al., 2018, Health literacy and outcomes in patients with heart failure: a prospective community study. In: Mayo Clinic Proceedings (Vol. 93, No. 1, pp. 9-15). Elsevier [DOI] [PMC free article] [PubMed]
  • 9.Kaps L, Omogbehin L, Hildebrand K et al. Health literacy in gastrointestinal diseases: a comparative analysis between patients with liver cirrhosis, inflammatory bowel disease and gastrointestinal cancer. Sci Rep 2022;12:21072. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 10.Mercuri C, Nocerino R, Bosco V et al. Health Literacy in Inflammatory Bowel Disease: A Systematic Review of Health Outcomes Predictors and Barriers. J Clin Med 2025;14:8577. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 11.Gwag M, Yoo J. Relationship between Health Literacy and Knowledge, Compliance with Bowel Preparation, and Bowel Cleanliness in Older Patients Undergoing Colonoscopy. Int J Environ Res Public Health 2022;19:2676. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12.Erdoğdu UE, Çaycı HM, Tardu A, Arslan U, Demirci H, Yıldırım Ç. Relationship between health literacy and quality of colonoscopy bowel preparation. Turk J Gastroenterol 2020;31:799. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 13.John GK, Thuluvath AJ, Carrier H, Ahuja NK, Gupta E, Stein E. Poor Health Literacy and Medication Burden Are Significant Predictors for Inadequate Bowel Preparation in an Urban Tertiary Care Setting. J Clin Gastroenterol 2019;53:e382–e386. [DOI] [PubMed] [Google Scholar]
  • 14.Rooney MK, Santiago G, Perni S et al. Readability of Patient Education Materials From High-Impact Medical Journals: A 20-Year Analysis. J Patient Exp 2021;8:2374373521998847. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15.Brega AG, Freedman MA, LeBlanc WG et al. Using the Health Literacy Universal Precautions Toolkit to Improve the Quality of Patient Materials. J Health Commun 2015;20:69–76. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16.AHRQ Health Literacy Universal Precautions Toolkit, C. Brach, Editor. 2024, Agency for Healthcare Research and Quality: Rockville, MD.
  • 17.Eltorai AE, Ghanian S, Adams CA Jr, Born CT, Daniels AH. Readability of patient education materials on the american association for surgery of trauma website. Arch Trauma Res 2014;3:e18161. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 18.Kher A, Johnson S, Griffith R. Readability Assessment of Online Patient Education Material on Congestive Heart Failure. Adv Prev Med 2017;2017:9780317. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19.Wang LW, Miller MJ, Schmitt MR, Wen FK. Assessing readability formula differences with written health information materials: application, results, and recommendations. Res Social Adm Pharm 2013;9:503–516. [DOI] [PubMed] [Google Scholar]
  • 20.Raja H, Lodhi S. Assessing the readability and quality of online information on anosmia. Ann R Coll Surg Engl 2024;106:178–184. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 21.Nash E, Bickerstaff M, Chetwynd AJ, Hawcutt DB, Oni L. The readability of parent information leaflets in paediatric studies. Pediatr Res 2023;94:1166–1171. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22.Hanci V, Otlu B, Biyikoğlu AS. Assessment of the Readability of the Online Patient Education Materials of Intensive and Critical Care Societies. Crit Care Med 2024;52:e47–e57. [DOI] [PubMed] [Google Scholar]
  • 23.McLaughlin GH. SMOG grading-a new readability formula. Journal of reading 1969;12:639–646. [Google Scholar]
  • 24.Kincaid, J.P., Fishburne Jr, R.P., Rogers, R.L. and Chissom, B.S., Derivation of new readability formulas (automated readability index, fog count and flesch reading ease formula) for navy enlisted personnel (No. RBR875) 1975.
  • 25.Aydin S, Karabacak M, Vlachos V, Margetis K. Large language models in patient education: a scoping review of applications in medicine. Front Med (Lausanne) 2024;11:1477898. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 26.Meskó B. Prompt Engineering as an Important Emerging Skill for Medical Professionals: Tutorial. J Med Internet Res 2023;25:e50638. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 27.Rajpurkar P, Chen E, Banerjee O, Topol EJ. AI in health and medicine. Nat Med 2022;28:31–38. [DOI] [PubMed] [Google Scholar]
  • 28.Wang L, Chen X, Deng X et al. Prompt engineering in consistency and reliability with the evidence-based guideline for LLMs. NPJ Digit Med 2024;7:41. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 29.Effective Prompts for AI: The Essentials. 2025 12/1/2025]; Available from: https://mitsloanedtech.mit.edu/ai/basics/effective-prompts/.
  • 30.Liu J, Liu F, Wang C, Liu S. Prompt Engineering in Clinical Practice: Tutorial for Clinicians. J Med Internet Res 2025;27:e72644. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 31.Wu JH, Nishida T, Moghimi S, Weinreb RN. Effects of prompt engineering on large language model performance in response to questions on common ophthalmic conditions. Taiwan J Ophthalmol 2024;14:454–457. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 32.Jung H, Oh J, Stephenson KA. Prompt engineering with ChatGPT3. 5 and GPT4 to improve patient education on retinal diseases. Can J Ophthalmol 2025;60:e375–e381. [DOI] [PubMed] [Google Scholar]
  • 33.Wang L, Li J, Zhuang B et al. Accuracy of Large Language Models When Answering Clinical Research Questions: Systematic Review and Network Meta-Analysis. J Med Internet Res 2025;27:e64486. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 34.Alamleh S, Mavedatnia D, Francis G et al. Readability, Reliability, and Quality Analysis of Internet-Based Patient Education Materials and Large Language Models on Meniere’s Disease. J Otolaryngol Head Neck Surg 2025;54:19160216251360652. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 35.Will J, Gupta M, Zaretsky J, Dowlath A, Testa P, Feldman J. Enhancing the Readability of Online Patient Education Materials Using Large Language Models: Cross-Sectional Study. J Med Internet Res 2025;27:e69955. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 36.Stephenson-Moe CA, Behers BJ, Gibons RM et al. Assessing the quality and readability of patient education materials on chemotherapy cardiotoxicity from artificial intelligence chatbots: An observational cross-sectional study. Medicine (Baltimore) 2025;104:e42135. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 37.Peery AF, Crockett SD, Barritt AS et al. Burden of Gastrointestinal, Liver, and Pancreatic Diseases in the United States. Gastroenterology 2015;149:1731–1741. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 38.Peery AF, Crockett SD, Murphy CC et al. Burden and Cost of Gastrointestinal, Liver, and Pancreatic Diseases in the United States: Update 2021. Gastroenterology 2022;162:621–644. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 39.Digestive Disease. 2025; Available from: https://www.niddk.nih.gov/health-information/digestive-diseases.
  • 40.Peery AF, Murphy CC, Anderson C et al. Burden and Cost of Gastrointestinal, Liver, and Pancreatic Diseases in the United States: Update 2024. Gastroenterology 2025;168:1000–1024. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 41.Textstat. 07/11/2025]; Available from: https://textstat.org.
  • 42.Nawaz MS, McDermott LE, Thor S. The Readability of Patient Education Materials Pertaining to Gastrointestinal Procedures. Can J Gastroenterol Hepatol 2021;2021:7532905. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 43.Yang G, Ye Q, Xia J. Unbox the black-box for the medical explainable AI via multi-modal and multi-centre data fusion: A mini-review, two showcases and beyond. Inf Fusion 2022;77:29–52. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 44.Diaz JA, Griffith RA, Ng JJ et al. Patients’ use of the Internet for medical information. J Gen Intern Med 2002;17:180–185. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 45.Chen J, Wang Y. Social Media Use for Health Purposes: Systematic Review. J Med Internet Res 2021;23:e17917. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 46.Hatem R, Simmons B, Thornton JE. A Call to Address AI “Hallucinations” and How Healthcare Professionals Can Mitigate Their Risks. Cureus 2023;15:e44720. [DOI] [PMC free article] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Data Availability Statement

Data is available on request. The Python program and code we utilized is also available on request.


Articles from Digestive Diseases and Sciences are provided here courtesy of Springer

RESOURCES