Skip to main content
World Journal of Otorhinolaryngology - Head and Neck Surgery logoLink to World Journal of Otorhinolaryngology - Head and Neck Surgery
. 2026 May 25:10.1002/wjo2.70118. Online ahead of print. doi: 10.1002/wjo2.70118

Comprehensive Evaluation of AI Consent Forms in Otolaryngologic Surgery

Sholem Hack 1,✉, Rebecca Attal 1, Armin Farzad 2, Lilia Ann Crew 3, Shmuel Silverstein 4, Jacob E Karni 5, Nir Livneh 6,7, Shibli Alsleibi 6,7, Naseem Saleh 6,7, Ben Gvili 6,7, David Yogev 6,7, Eric Remer 6,7, Eran E Alon 6,7, Gil Siegal 6,7, Masayoshi Takashima 8, Habib G Zalzal 9
PMCID: PMC13399016  PMID: 42500047

ABSTRACT

Introduction

Surgical consent documents are frequently written at reading levels exceeding average health literacy. Large language models (LLMs) may offer a scalable approach to generating clearer, procedure‐specific consent forms. This study evaluated the clarity, clinical accuracy, and acceptability of consent forms generated by GPT‐4 and Claude for common otolaryngologic procedures.

Methods

Twenty AI‐generated consent forms (10 GPT‐4.0, 10 Claude‐2.1) were produced using standardized prompts. In a survey‐based, non‐clinical setting, five board‐certified otolaryngologists independently rated each form for medical accuracy, readability, comprehensibility, legal/ethical sufficiency, and usability using a 4‐point scale. A cross‐sectional cohort of 300 English‐speaking adults (15 raters per form) evaluated perceived clarity and signing comfort on 5‐point Likert scales, and perceived trust using a binary (Yes/No) item, and completed eight binary quality assessments. A blinded subgroup (n = 10) compared AI‐generated and official national health system templates across five Likert domains. Readability was assessed using Flesch–Kincaid Grade Level (FKGL).

Results

Mean lay ratings for clarity across AI‐generated forms were high overall. Claude demonstrated numerically higher scores than GPT‐4 for clarity (4.72 vs. 4.68), perceived trust (reported as proportions), and signing comfort (4.40 vs. 4.27). However, when analyzed at the form level, differences between models were not statistically significant for clarity (mean difference 0.04; t(9) = −0.80; p = 0.44) or signing comfort (mean difference 0.13, t(9) = −1.68, p = 0.13). Across binary domains, ≥ 95% of participants affirmed adequate explanation of risks, benefits, and alternatives. Experts rated GPT‐4 more accurate than Claude (2.2 vs. 1.5, p = 0.034). Mean Flesch–Kincaid Grade Level was lower for AI‐generated forms compared to official templates. Although prompts targeted a 6th–8th grade reading level, achieved readability scores were slightly higher (8.8–9.4).

Conclusions

In a non‐clinical evaluation, AI‐generated consent forms were perceived as clear and clinically complete, with model‐specific trade‐offs between perceived clarity and clinical detail. These perception‐based findings—reflecting participant ratings of clarity, perceived trust, and willingness to sign rather than objective comprehension—are hypothesis‐generating, and prospective clinical and legal validation in more representative patient populations is required.

Keywords: health literacy, informed consent, large language models, otolaryngology, patient communication

1. Introduction

Informed consent is a cornerstone of ethical medical practice, serving as a mechanism to ensure patient autonomy, safety, and trust in healthcare settings [1]. As a structured communication process, informed consent enables patients to make informed decisions by understanding the risks, benefits, and alternatives of proposed procedures [1]. Despite its importance, research consistently demonstrates that traditional surgical consent forms fail to meet the needs of diverse patient populations, often presenting information at readability levels far beyond the average health literacy of the population, thereby leading to compromised shared‐decision making, patient confusion and reduced comprehension [2, 3].

The gap between the complexity of medical documentation and the average patient's reading level is well‐documented. In the United States, the average adult reads at an 8th‐grade level, yet most consent forms are written at a college freshman level or higher [2]. This disconnect disproportionately affects vulnerable groups, including individuals with limited health literacy, non‐native English speakers, and those with lower socioeconomic status [4, 5]. Efforts to simplify consent forms, such as creating procedure‐specific documents and leveraging visual aids, have demonstrated modest improvements. However, these approaches require significant time resources, and continuous updates, making them challenging to sustain in many healthcare settings [6, 7].

Recent advancements in artificial intelligence (AI), particularly Large Language Models (LLMs), offer new avenues for addressing these challenges [8, 9, 10, 11, 12, 13, 14]. LLMs, such as ChatGPT (developed by OpenAI) and Claude (developed by Anthropic), have demonstrated the ability to generate coherent, context‐specific text, with applications ranging from summarizing complex information to creating patient education materials [15, 16, 17]. By harnessing the capabilities of these models, healthcare systems can potentially transform the development of consent forms, making them more accessible and tailored to individual patient needs [18]. Preliminary research has shown that LLMs can produce consent forms that match or exceed traditional forms in readability, accuracy, and completeness [2, 18].

Questions remain about the real‐world acceptability, trustworthiness, and safety of these documents, as few studies have examined the comparative performance of multiple AI models or triangulated expert review with layperson perspectives. This study addresses those gaps by evaluating the quality and acceptability of AI‐generated consent forms for 10 common otolaryngologic procedures. We compare outputs from two leading LLMs, GPT‐4 and Claude, on dimensions of readability, clinical accuracy, and layperson‐perceived clarity, trust, and signing comfort. Our evaluation combines structured reviews by otolaryngologists and feedback from a diverse layperson cohort to assess clarity, trust, and signing comfort. By integrating both expert and patient perspectives, this study offers a multidimensional assessment of AI's potential to improve surgical informed consent. In doing so, it contributes to the growing evidence that LLMs, when appropriately validated, can support equitable, scalable, and patient‐centered communication in surgical care. We hypothesized that AI‐generated consent forms would demonstrate high perceived clarity and acceptability among lay participants, with differences between models reflecting trade‐offs between readability and clinical detail. The primary outcome was layperson‐rated clarity of language in the main cohort.

2. Methods

2.1. Study Design

This was a three‐arm evaluation study of AI‐generated surgical consent forms for 10 common otolaryngologic procedures. We assessed (1) clinical accuracy and usability using an expert otolaryngologist panel, (2) layperson perceptions in a large cross‐sectional cohort, and (3) blinded head‐to‐head comparisons between AI‐generated and official national health system consent templates. The study was approved by the Institutional Review Board (1761‐24‐SMC‐D) and all data were collected anonymously.

2.2. AI‐Generated Consent Forms

Ten otolaryngologic procedures were selected: functional endoscopic sinus surgery, cochlear implantation, neck dissection, parotidectomy, rhinoplasty, septoplasty, stapedectomy, thyroidectomy, tonsillectomy, and myringotomy. For each procedure, one consent form was generated with GPT‐4.0 (OpenAI) and one with Claude 2.1 (Anthropic) using standardized prompts that requested a description of the procedure, risks, benefits, alternatives, post‐operative expectations, and a legal attestation section (Supporting Information: S1). This yielded 20 AI‐generated forms (10 per model) for evaluation.

2.3. Arm 1: Expert Otolaryngologist Evaluation

Five board‐certified otolaryngologists (NL, SA, NS, BG, and DY), none of whom were involved in generating the prompts, independently reviewed all 20 AI‐generated forms. Expert reviewers were blinded to the source model (GPT‐4 vs Claude) for each consent form. Experts rated each form on five domains: (1) medical accuracy and completeness, (2) readability, (3) comprehensibility for patients, (4) legal and ethical sufficiency, and (5) practical usability for clinicians. Each domain was scored on a four‐category ordinal scale ranging from 0 to 3 (0 = inaccurate, 1 = incomplete, 2 = partial, 3 = complete), based on predefined evaluation criteria provided to reviewers. This structured scoring system was chosen to encourage clear categorical judgments and reduce ambiguity between intermediate ratings, facilitating consistent interpretation among reviewers. Mean domain scores were calculated for each model, and paired‐sample t‐tests with Cohen's d effect sizes were used to compare GPT‐4 and Claude. Inter‐rater agreement across experts was assessed using Fleiss' kappa.

2.4. Arm 2: Main Layperson Cohort (n = 300)

Three hundred adults were recruited through personal and professional networks in the United States, Israel, and other countries. Eligibility criteria included age ≥ 18 years, self‐reported English proficiency, and consent to anonymous participation. Demographic characteristics for the main cohort are summarized in Table 1. The mean age was 35 years (range 18–70); 55% identified as male and 45% as female. Participants resided in the United States (n = 150), Israel (n = 80), or other countries (n = 70). English proficiency was reported as fluent (63%), advanced (20%), intermediate (10%), or basic (7%). The highest completed education level was primary (5%), secondary (25%), or university degree (70%).

Table 1.

Demographic characteristics of the main (n = 300) and subgroup (n = 10) lay reviewer cohorts, including age, gender, nationality, English proficiency, and education level.

Characteristic Original cohort (n = 300) Subgroup (n = 10)
Age (years, mean) 35 [range:18–70] 36 [range:18–70]
Gender
Male 165 (55%) 6 (60%)
Female 135 (45%) 4 (40%)
Country of residence
USA 150 (50%) 5 (50%)
Israel 80 (26.7%) 5 (50%)
Other 70 (23.3%) 0 (0%)
English proficiency
Fluent 190 (63.3%) 5 (50%)
Basic 20 (6.7%) 3 (30%)
Intermediate/Conversational 30 (10%) 2 (20%)
Advanced 60 (20%) 0 (0%)
Highest education level
Primary 15 (5%) 0 (0%)
Secondary 75 (25%) 2 (20%)
University degree 210 (70%) 8 (80%)

Note: No data.

Each participant was randomly assigned to review a single AI‐generated consent form, with 15 unique raters per form (10 procedures × 2 models × 15 raters = 300). After reading the form, participants completed two types of ratings. First, they evaluated two Likert‐scale domains on a 5‐point scale (1 = strongly disagree, 5 = strongly agree): clarity of the language and comfort with signing the form (“signing comfort”), while perceived trust was assessed using a binary (Yes/No) item evaluating whether the form built trust in the information provided. Second, they answered eight binary (Yes/No) questions assessing whether the form: (1) clearly explained risks, (2) clearly explained benefits, (3) presented alternatives, (4) avoided difficult medical jargon, (5) used sentences that were not too long, (6) provided sufficient information to feel informed, (7) felt tailored to the specific procedure, and (8) helped build trust in the information presented. Items were phrased positively such that higher scores indicated better perceived quality (Table 3). All consent forms were generated and evaluated in English.

Table 3.

Summary of binary layperson ratings for ai‐generated surgical consent forms (GPT‐4 vs. Claude 2.1).

Domain GPT‐4 Yes (%) Claude Yes (%)
Risks clearly explained 98 96
Benefits clearly explained 100 100
Alternatives presented 98 95
Avoided medical jargon 98 97
Sentences not too long 100 100
Sufficient information 100 98
Form felt tailored 100 100
Form built trust 100 98

Note: This table presents the proportion of layperson respondents (n = 300) who selected “Yes” in response to eight binary evaluation items following review of AI‐generated surgical consent forms. Each form was reviewed by 15 unique participants (10 procedures × 2 models = 20 forms). Items assessed whether forms clearly explained risks and benefits, presented alternatives, avoided medical jargon, used simple sentence structure, provided sufficient information, felt tailored to the procedure, and built trust. All items were phrased positively, such that higher percentages indicate better perceived quality. No missing data were present.

2.5. Arm 3: Blinded Benchmark Subgroup (n = 10)

To benchmark AI‐generated forms against official templates, a separate subgroup of 10 adult lay reviewers (mean age 36 years; 60% male) was recruited, evenly split between Israel and the United States. Five participants reported fluent English proficiency, two conversational, and three basic; eight had a university degree and two had completed secondary school.

Based on initial ratings from the main cohort, three representative procedures—cochlear implantation, rhinoplasty, and thyroidectomy—were chosen for focused comparison. For each procedure, the top‐rated GPT‐4 and Claude forms were paired with an official template obtained from a national health system (e.g., Israeli Medical Association, Queensland Health, or UK National Health Service). Each subgroup participant reviewed all nine forms (three GPT‐4, three Claude, three official) in randomized order, blinded to the source. They rated five Likert domains (1–5 scale): clarity of language, explanation of alternatives, ease of readability, confidence in signing, and overall quality (Table 2).

Table 2.

Subgroup blinded.

Domain Claude GPT‐4 Official p‐value (Claude vs. Official) p‐value (GPT‐4 vs. Official) p‐value (GPT‐4 vs. Claude)
Explanation of alternatives 4.4 4.5 2.0 < 0.001 < 0.001 0.447
Clarity of language 4.6 4.5 3.5 < 0.001 < 0.001 0.054
Ease of readability 4.8 4.8 4.0 < 0.001 0.003 0.094
Confidence in signing 5 5 4.0 <0.001 < 0.001 0.852
Overall 4.7 4.5 4.0 < 0.001 0.021 0.093

Note: Mean layperson ratings (scale: 1 = Strongly disagree to 5 = Strongly agree) across five consent quality domains for Claude, GPT‐4, and official human‐written templates. Each of the 10 participants rated all 9 forms in a blinded, randomized order. Higher scores reflect greater perceived clarity, readability, and acceptability.

Given the small sample size, analyses from this subgroup were considered exploratory and hypothesis‐generating.

2.6. Readability Analysis

Textual complexity of each consent form was quantified using the Flesch–Kincaid Grade Level (FKGL) formula, including both headings and body text. FKGL was calculated for all 20 AI‐generated forms and for official templates from the three national health systems. Mean grade levels were computed by source (GPT‐4, Claude, official templates), and between‐group comparisons used independent‐sample t‐tests.

2.7. Statistical Analysis

Descriptive statistics (means, standard deviations, 95% confidence intervals [CIs], and proportions) were computed for all study arms. For the main lay cohort, model‐level comparisons between GPT‐4 and Claude were conducted at the form level by comparing mean Likert scores across the 10 procedures generated by each model using paired‐sample t‐tests, with each procedure represented once per model to allow within‐procedure comparisons. For Likert‐scale outcomes in the main lay cohort, variability was summarized using standard deviations and 95% CIs calculated using the t‐distribution at the form level (n = 10 per model). Cohen's d was calculated to estimate effect sizes. Binary lay outcomes were summarized as proportions of “Yes” responses without formal hypothesis testing. Expert panel comparisons between models were performed using paired‐sample t‐tests. Because each participant in the benchmark subgroup evaluated all nine forms, within‐subject comparisons between AI‐generated and official human‐written templates were analyzed using paired‐sample t‐tests. Statistical significance was defined as p < 0.05. All analyses were performed using IBM SPSS Statistics version 29.0. Model‐level analyses for the main lay cohort were conducted at the form level (10 procedures per model), using mean scores per form as the unit of analysis to avoid pseudoreplication.

3. Results

3.1. Participant Characteristics

A total of 300 lay participants completed evaluations in the main cohort (Supplementary Figure 1). Each of the 20 AI‐generated consent forms received 15 independent lay ratings. As summarized in Table 1, the mean age was 35 years (range 18–70); 55% of participants identified as male and 45% as female. Participants were primarily from the United States (n = 150) and Israel (n = 80), with the remaining 70 residing in other countries. English proficiency was self‐reported as fluent (63%), advanced (20%), intermediate (10%), or basic (7%). The highest level of education completed was university degree for 70%, secondary school for 25%, and primary school for 5%. Because a substantial proportion of participants had university‐level education, perceived readability and clarity ratings may overestimate the performance of the consent forms compared with more representative surgical patient populations.

3.2. Layperson Evaluation of AI‐Generated Consent Forms (Main Cohort)

Across all 300 participants, AI‐generated consent forms were rated highly on the two Likert domains. GPT‐4 forms received mean ratings of 4.68 (95% CI 4.61–4.76) for clarity and 4.27 (95% CI 4.12–4.43) for signing comfort. Claude forms received higher mean ratings of 4.72 (95% CI 4.64–4.80) for clarity and 4.40 (95% CI 4.25–4.55) for signing comfort (Figure 1). Perceived trust was assessed as a binary outcome and is reported as proportions (Table 3).

Figure 1.

Figure 1

Layperson rating by domain and model. Mean Likert scores (1–5) for clarity and signing comfort across AI‐generated consent forms. Claude demonstrated numerically higher scores than GPT‐4 across both domains; however, differences between models were not statistically significant when analyzed at the form level (n = 10 per model). Bars represent mean values; error bars represent 95% confidence intervals.

Model‐level comparisons using form‐level means (n = 10 per model) demonstrated no statistically significant differences between GPT‐4 and Claude for clarity [mean difference 0.04, t(9) = −0.80, p = 0.44] or signing comfort [mean difference 0.13; t(9) = −1.68, p = 0.13]. CIs were narrow across domains, indicating consistent performance across procedures.

Binary data (Table 3) further supported strong performance of both models. For GPT‐4 forms, 98% of participants affirmed clear explanation of risks, 100% clear explanation of benefits, 98% presentation of alternatives, 98% avoidance of medical jargon, 100% appropriate sentence simplicity, 100% sufficiency of information, 100% procedural tailoring, and 100% trust‐building. Claude forms received similarly high ratings: 96% for risks, 100% for benefits, 95% for alternatives, 97% for avoidance of jargon, 100% sentence simplicity, 98% sufficiency, 100% tailoring, and 98% trust‐building. Several binary outcomes demonstrated ceiling effects, which may reflect limited sensitivity of the binary instrument or social desirability bias among respondents, reducing the ability to discriminate between models.

3.3. Expert Evaluation of AI‐Generated Consent Forms

Five board‐certified otolaryngologists independently reviewed all 20 AI‐generated forms. Inter‐rater agreement across the five evaluation domains was very high (Fleiss' κ = 0.95), indicating strong consistency among experts. GPT‐4 forms were rated significantly more accurate and complete than Claude on the 0–3 expert evaluation scale (mean 2.2 vs. 1.5, t(4) = 3.16, p = 0.034, Cohen's d = 1.0; 95% CI of the difference 0.05–0.95). No significant differences were observed between models for readability, legal/ethical sufficiency, comprehensibility for patients, or practical usability for clinicians. The very high inter‐rater agreement likely reflects the use of predefined structured scoring criteria with explicit domain definitions and ordinal anchors, which reduced interpretive variability among reviewers.

3.4. Readability Analysis

Readability assessment using the Flesch–Kincaid Grade Level formula demonstrated that Claude forms required a mean reading level of 8.8, GPT‐4 forms required 9.4, and official human‐written templates required 11.1 (Figure 2). Across all AI‐generated forms, the combined mean FKGL was 9.1. Official templates were therefore approximately two grade levels more complex than AI‐generated forms. Despite these objective readability differences, lay participants rated AI‐generated forms as easier to read than official templates. Notably, although prompts targeted a 6th–8th grade reading level, achieved FKGL scores remained higher (8.8–9.4), indicating limitations in controlling output readability using LLMs. In the benchmark subgroup, mean readability ratings were 4.6 for Claude, 4.4 for GPT‐4, and 3.9 for official templates (Supporting Information: Figure 2).

Figure 2.

Figure 2

Mean readability scores by form source. Mean Flesch–Kincaid Grade Levels for consent forms generated by GPT‐4, Claude, and official human‐written templates. Claude‐generated forms had the lowest mean grade level (8.8), followed by GPT‐4 (9.4), whereas official templates had the highest mean grade level (11.1). Despite prompts targeting a 6th–8th grade reading level, achieved FKGL scores remained higher (8.8–9.4), highlighting limitations in controlling output readability using LLMs.

3.5. Blinded Benchmark Subgroup Comparison

In the blinded benchmark arm, 10 new lay reviewers evaluated nine forms (three Claude, three GPT‐4, three official templates) for cochlear implantation, rhinoplasty, and thyroidectomy. Across the 5 Likert domains—clarity, explanation of alternatives, readability, confidence in signing, and overall quality—AI‐generated forms generally received higher mean ratings than matched human‐written templates (Table 2). Claude achieved a mean of 4.6 for clarity, 4.8–5.0 for readability and confidence in signing, and 4.7 for overall quality. GPT‐4 showed similar results with slightly lower but still high ratings, whereas official templates scored lower across all domains and were more frequently flagged for excessive medical jargon and limited tailoring. Given the small sample size and exploratory nature of this arm, these differences should be interpreted cautiously as hypothesis‐generating rather than definitive.

4. Discussion

This study is among the first to comprehensively evaluate AI‐generated surgical consent forms across two LLMs, three national health system templates, and dual stakeholder perspectives. Our findings suggest that AI‐generated forms can be medically accurate and acceptable to lay reviewers across diverse educational and linguistic backgrounds. However, while these forms met general standards for completeness as judged by clinical experts, their legal sufficiency was not formally assessed and may vary substantially by country, institution, and local law. Claude‐generated documents were perceived as clearer and more trustworthy in descriptive analyses, although these differences were not statistically significant in the main lay cohort when analyzed at the form level, whereas GPT‐4 forms were rated significantly more accurate by expert otolaryngologists. These findings suggest that LLMs may help generate patient‐facing documents that are credible and usable, with potential to support the informed consent process when combined with appropriate clinical oversight, institutional governance, and legal review.

Our results align with and extend recent studies exploring AI in consent form generation. Decker et al. showed that ChatGPT‐generated RBAs were more complete and readable than those authored by surgeons for six common general surgeries [18]. Similarly, Vaira et al. demonstrated that GPT‐4 outperformed both Gemini and a human resident in producing accurate and complete oral surgery consent forms, as judged by clinical experts [19]. Our study confirms these trends in the otolaryngologic context and expands on them by including layperson evaluations, an element that was absent in the Vaira study but repeatedly cited as a critical next step [19]. We also evaluated forms from two LLMs side‐by‐side, allowing us to highlight distinct trade‐offs in tone, clarity, and clinical robustness that would otherwise be overlooked in single‐model designs.

Layperson responses in our study were overwhelmingly positive. Across Likert‐based domains (clarity and signing comfort), forms were rated highly in the main cohort, and for clarity, readability, explanation of alternatives, confidence to sign, and overall quality in the benchmark subgroup. Perceived trust, assessed as a binary outcome, was also rated highly, with over 95% of participants across both models affirming that the forms clearly explained risks, benefits, and alternatives, avoided confusing jargon, and helped build trust. These findings echo the work of Gan et al., who reported reduced patient anxiety and increased satisfaction after ChatGPT‐enhanced consent discussions [20]. Our findings suggest that LLM‐generated documents may help support perceived clarity and trust, especially in procedures that carry emotional or esthetic significance, such as rhinoplasty or cochlear implantation. However, when analyzed at the form level, differences between models for clarity and signing comfort were not statistically significant, suggesting broadly comparable performance between GPT‐4 and Claude in lay‐facing domains. Importantly, signing comfort should be interpreted cautiously, as willingness to sign does not necessarily reflect true comprehension, clinical accuracy, or legal sufficiency, and may instead reflect perceived trust in the presentation of information rather than informed decision‐making.

Expert evaluations further reinforce these conclusions. GPT‐4 forms were rated significantly more accurate than those generated by Claude (p = 0.034, d = 1.0), with strong inter‐rater reliability (Fleiss' Kappa = 0.95). This is consistent with Pegano et al. and Rydzewski et al., who also found that GPT‐4 preserved more nuanced clinical detail than comparator models [21, 22]. The inclusion of core informed consent components across forms suggests that LLMs can generate documents that are structurally complete, although formal legal sufficiency was not independently validated.

A unique strength of this study lies in its triangulated design. We compared multiple LLMs, integrated both expert and lay feedback, and benchmarked AI‐generated content against national health system templates from Israel, the UK, and Australia. This multidimensional approach may enhance the external relevance of the findings compared to prior studies that relied solely on simulated patients or clinician reviewers.

The addition of a focused comparison arm, although based on a small and non‐representative subgroup, provided exploratory evidence that AI‐generated consent materials received higher lay ratings for clarity, trust, and readability than some existing national health system forms. However, these findings should be considered exploratory, and further research using larger and more representative samples is needed before drawing strong conclusions about the relative advantages or clinical effectiveness of AI‐generated documents compared to established standards.

A notable observation was the discrepancy between objective readability metrics and subjective ratings. Notably, despite prompts targeting a 6th–8th grade reading level, achieved FKGL scores remained slightly higher (8.8–9.4), highlighting current limitations in controlling output readability using LLMs. This suggests that while LLMs may improve readability relative to traditional templates, they may not reliably achieve predefined literacy targets without additional prompting constraints or post‐generation editing.

Although AI‐generated forms had FKGL scores ranging from 8.8 to 9.4, participants rated them highly for readability (mean scores 4.4–4.6), whereas official templates demonstrated a higher FKGL of 11.1 and lower subjective readability ratings (3.9). This highlights the limitations of formula‐based readability assessments, which often fail to capture tone, layout, and conceptual clarity. Similar critiques have been made by Gill et al. and Gao et al., who found that patient‐perceived readability often diverges from automated grade‐level scores [23]. Our findings reinforce the value of human‐centered validation in evaluating educational materials.

The consistent strengths of Claude in lay‐facing domains and GPT‐4 in clinical fidelity suggest that different models may be better suited to different stages of the informed consent process. Claude may be ideal for initial communication or low‐literacy populations, whereas GPT‐4 may better support pre‐operative documentation or specialist review. Hybrid approaches leveraging the strengths of each may offer the most robust solution for clinical integration.

Ethical and equity considerations are central to the deployment of AI‐generated consent materials [24, 25, 26, 27, 28, 29, 30]. Because LLMs are trained on large, non‐transparent corpora, they may reproduce biases in tone, cultural framing, or risk communication that disadvantage specific groups (e.g., by normalizing majority perspectives or minimizing uncertainty) [24, 25, 26, 27, 28, 29, 30]. Overly polished language may also engender unwarranted trust in the document's completeness or legal sufficiency [24, 25, 26, 27, 28, 29, 30]. In addition, responsibility for errors or omissions in AI‐generated consent text remains unclear, raising questions about liability for clinicians, institutions, and vendors. Any clinical deployment of such tools should therefore incorporate explicit institutional governance, local legal review, and mechanisms for ongoing monitoring of bias, harm, and patient feedback [24, 25, 26, 27, 28, 29, 30].

5. Limitations and Future Directions

Several limitations warrant consideration. First, participants were recruited through convenience sampling using personal and professional networks, resulting in a relatively young and highly educated sample that may not reflect the demographic and health‐literacy distribution of typical surgical populations. Additionally, 70% of participants held a university degree, a level of education that substantially exceeds that of many surgical populations and represents a major limitation with respect to health literacy. This demographic skew likely overestimates perceived clarity and comprehension and introduces selection and verification bias, thereby limiting the generalizability of our findings to lower‐literacy or more socioeconomically diverse patient groups. In addition, English proficiency was self‐reported and not a primary study variable; however, a small proportion of participants (7%) reported basic proficiency, which may have introduced variability in perception‐based ratings. Future research should employ more systematic and representative recruitment strategies.

Second, although each of the 20 AI‐generated forms was rated by 15 unique lay participants, model‐level statistical comparisons were based on 10 procedure‐level means per model, which limits power to detect smaller differences between GPT‐4 and Claude. Similarly, the expert panel consisted of five board‐certified otolaryngologists. While inter‐rater reliability was high (Fleiss' κ = 0.95), the small expert sample (n = 5) limits statistical power, and the observed statistically significant difference between models (p = 0.034) should be interpreted cautiously; a larger and more heterogeneous expert group might have provided additional perspectives on clinical nuance and legal sufficiency.

Third, evaluations were conducted outside clinical environments and did not include real patients in decision‐making contexts. Objective comprehension testing, such as knowledge quizzes, recall tasks, or simulated decision‐making, was not performed. Our outcome measures were therefore limited to participant perceptions of clarity, trust, and signing comfort, rather than direct measures of understanding or retention. While such objective testing would be valuable for research purposes, it is worth noting that in routine clinical practice, patients are rarely assessed with formal comprehension tests after reviewing consent forms; instead, patient perception and self‐reported understanding guide the consent process. Nonetheless, further studies using objective comprehension assessments may provide additional insights into the real‐world effectiveness of AI‐generated consent documents.

Fourth, while GPT‐4.0 and Claude 2.1 were leading LLMs at the time of evaluation, both have since been superseded by newer models. They are trained on general corpora and are not optimized for specific legal frameworks, institutional policies, or cultural norms. While core capabilities in structured medical text generation may persist, performance characteristics may differ with newer systems, and future validation is warranted. Clinical oversight and institutional or legal review are essential before real‐world deployment, especially as regulatory and medicolegal frameworks for AI‐generated consent forms are still evolving.

Despite expert validation, LLMs remain susceptible to factual inaccuracies, omissions, or “hallucinations.” Although no major errors were identified in this study, hallucination risk remains an inherent limitation of generative AI. Ongoing expert review, iterative content vetting, and, where feasible, domain‐specific model tuning are critical to ensure these tools remain safe, ethical, and compliant with institutional and legal requirements.

Finally, while national templates were used for human‐written comparisons, these documents varied widely in structure, purpose, and intended audience, limiting the precision of quantitative comparison. This diversity, however, also offers important qualitative benchmarks for evaluating the adaptability and clarity of AI‐generated documents.

Future research should explore adaptive, personalized consent generation tailored to patient education level, emotional needs, and language preference. Real‐world trials assessing patient satisfaction, legal outcomes, and workflow efficiency are necessary to establish clinical and legal feasibility. Integration with institutional review, medicolegal oversight, and patient‐centered feedback will be essential as AI‐generated documents move toward clinical use.

6. Conclusions

AI‐generated surgical consent forms produced by GPT‐4 and Claude were rated as accurate and acceptable by both lay and expert reviewers in this survey‐based, non‐clinical evaluation. Claude excelled in perceived clarity and readability, while GPT‐4 demonstrated greater clinical accuracy and completeness as judged by otolaryngologists. While these results suggest potential value in leveraging large language models to support patient communication, real‐world implementation should proceed cautiously, with appropriate clinical, legal, and institutional oversight. Prospective studies in more diverse, real‐world patient populations are needed to determine the impact of AI‐generated documents on understanding, decision‐making, and outcomes.

Author Contributions

Sholem Hack: conception and design, acquisition of data, analysis of data, resources, methodology, project administration, visualization, drafting of article and/or critical revision, final approval of manuscript. Rebecca Attal, Armin Farzad, Lilia Ann Crew, Jacob E. Karni, Shmuel Silverstein, Nir Livneh, Shibli Alsleibi, Naseem Saleh, Ben Gvili, David Yogev, Eric Remer, Eran E. Alon, and Gil Siegal: acquisition of data, visualization, drafting of article and/or critical revision, final approval of manuscript Masayoshi Takashima: supervision, validation, drafting of article and/or critical revision, final approval of manuscript. Habib G. Zalzal: conception and design, analysis of data, supervision, methodology, project administration, visualization, drafting of article and/or critical revision, final approval of manuscript.

Funding

The authors have nothing to report.

Ethics Statement

This study was approved by the Institutional Review Board (IRB approval number: 1761‐24‐SMC‐D). All data were collected anonymously, and no personal health information was obtained.

Conflicts of Interest

The authors declare no conflicts of interest.

Supporting information

Supporting File 1

WJO2-9999-0-s001.pdf (890.5KB, pdf)

Supporting File 2

WJO2-9999-0-s003.png (807.8KB, png)

Supporting File 3

WJO2-9999-0-s002.docx (16.7KB, docx)

Acknowledgments

The authors would like to thank the laypersons who participated in this study for their time and contribution.

Data Availability Statement

The data that support the findings of this study are available from the corresponding author upon reasonable request.

References

  • 1. Shah P., Thornton I., Kopitnik N. L., and Hipskind J. E., “Informed Consent,” StatPearls [Internet] (StatPearls Publishing, 2025), https://www.ncbi.nlm.nih.gov/books/NBK430827/. [PubMed] [Google Scholar]
  • 2. Ali R., Connolly I. D., Tang O. Y., et al., “Bridging the Literacy Gap for Surgical Consents: An Ai‐Human Expert Collaborative Approach,” NPJ Digital Medicine 7, no. 1 (2024): 63, 10.1038/s41746-024-01039-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 3. Duran A., Cortuk O., and Ok B., “Future Perspective of Risk Prediction in Aesthetic Surgery: Is Artificial Intelligence Reliable,” Aesthetic Surgery Journal 44, no. 11 (2024): NP839–NP849, 10.1093/asj/sjae140. [DOI] [PubMed] [Google Scholar]
  • 4. Berkman N. D., Sheridan S. L., Donahue K. E., Halpern D. J., and Crotty K., “Low Health Literacy and Health Outcomes: An Updated Systematic Review,” Annals of Internal Medicine 155, no. 2 (2011): 97–107, 10.7326/0003-4819-155-2-201107190-00005. [DOI] [PubMed] [Google Scholar]
  • 5. Rudd R. E., “Improving Americans' Health Literacy,” New England Journal of Medicine 363, no. 24 (2010): 2283–2285, 10.1056/NEJMp1008755. [DOI] [PubMed] [Google Scholar]
  • 6. Sudore R. L. and Schillinger D., “Interventions to Improve Care for Patients With Limited Health Literacy,” Journal of Clinical Outcomes Management 16, no. 1 (2009): 20–29. [PMC free article] [PubMed] [Google Scholar]
  • 7. Jin Q., Chen F., Zhou Y., et al., “Hidden Flaws Behind Expert‐Level Accuracy of Multimodal GPT‐4 Vision in Medicine,” NPJ Digital Medicine 7, no. 1 (2024): 190, 10.1038/s41746-024-01185-7. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 8. Iserson K. V., “Informed Consent for Artificial Intelligence in Emergency Medicine: A Practical Guide,” American Journal of Emergency Medicine 76 (2024): 225–230, 10.1016/j.ajem.2023.11.022. [DOI] [PubMed] [Google Scholar]
  • 9. Alter I. L., Chan K., Lechien J., and Rameau A., “An Introduction to Machine Learning and Generative Artificial Intelligence for Otolaryngologists‐Head and Neck Surgeons: A Narrative Review,” European Archives of Oto‐Rhino‐Laryngology 281, no. 5 (2024): 2723–2731, 10.1007/s00405-024-08512-4. [DOI] [PubMed] [Google Scholar]
  • 10. Hack S., Gvili B., Alsleibi S., Yogev D., and Glikson E., “Do Large Language Models Meet Professional Standards in Rhinoplasty?? A Comparative Evaluation With AAO‐HNS Guidelines,” European Journal of Plastic Surgery 48 (2025): 71, 48:1–9, 10.1007/s00238-025-02329-y. [DOI] [Google Scholar]
  • 11. Hack S., Gvili B., Tessler I., Yogev D., Wolfowitz A., and Rozendorn N., “Can Chatbots Please Both Patients and Experts? Benchmarking AI and Clinical Guidelines for Hearing Loss,” Otology & Neurotology 47, no. 1 (2026): 64–69, 10.1097/MAO.0000000000004686. [DOI] [PubMed] [Google Scholar]
  • 12. Hack S., Alsleibi S., Saleh N., Alon E. E., Rabinovics N., and Remer E., “Are Chatbots a Reliable Source for Patient Frequently Asked Questions on Neck Masses,” European Archives of Oto‐Rhino‐Laryngology 282, no. 8 (2025): 4273–4282, 10.1007/s00405-025-09433-6. [DOI] [PubMed] [Google Scholar]
  • 13. Hack S., Attal R., Farzad A., et al., “Performance of Generative AI Across ENT Tasks: A Systematic Review and Meta‐Analysis,” Auris, Nasus, Larynx 52, no. 5 (2025): 585–596, 10.1016/j.anl.2025.08.010. [DOI] [PubMed] [Google Scholar]
  • 14. Hack S., Attal R., Geva K., et al., “Blinded Comparative Evaluation of GPT‐Generated, Online Search‐Derived, and Guideline‐Based Answers for HPV‐Associated Oropharyngeal Cancer,” Oral Oncology 172 (2026): 107813, 10.1016/j.oraloncology.2025.107813. [DOI] [PubMed] [Google Scholar]
  • 15. Albehairi S. A., Alsahli O. M., Bander L. F., Albilasi T. M., and Aljardan E. S., “Postoperative Otoplasty Care With ChatGPT‐4: A Study on Artificial Intelligence (AI)‐Assisted Patient Concern and Education,” Journal of Craniofacial Surgery 36, no. 1 (2025): 296–298, 10.1097/SCS.0000000000010678. [DOI] [PubMed] [Google Scholar]
  • 16. Moise A., Centomo‐Bozzo A., Orishchak O., Alnoury M. K., and Daniel S. J., “Can ChatGPT Replace An Otolaryngologist in Guiding Parents on Tonsillectomy,” Ear, Nose, & Throat Journal, ahead of print, April 2, 2024, 10.1177/01455613241230841. [DOI] [PubMed] [Google Scholar]
  • 17. Pandya S., Alessandri Bonetti M., Liu H. Y., Jeong T., Ziembicki J. A., and Egro F. M., “Concordance of ChatGPT With American Burn Association Guidelines on Acute Burns,” Annals of Plastic Surgery 93, no. 5 (2024): 564–574, 10.1097/SAP.0000000000004128. [DOI] [PubMed] [Google Scholar]
  • 18. Decker H., Trang K., Ramirez J., et al., “Large Language Model‐Based Chatbot vs Surgeon‐Generated Informed Consent Documentation for Common Procedures,” JAMA Network Open 6, no. 10 (2023): e2336997, 10.1001/jamanetworkopen.2023.36997. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19. Vaira L. A., Lechien J. R., Maniaci A., et al., “Evaluating AI‐Generated Informed Consent Documents in Oral Surgery: A Comparative Study of ChatGPT‐4, Bard Gemini Advanced, and Human‐Written Consents,” Journal of Cranio‐Maxillofacial Surgery 53, no. 1 (2025): 18–23, 10.1016/j.jcms.2024.10.002. [DOI] [PubMed] [Google Scholar]
  • 20. Gan W., Ouyang J., She G., et al., “ChatGPT's Role in Alleviating Anxiety in Total Knee Arthroplasty Consent Process: A Randomized Controlled Trial Pilot Study,” International Journal of Surgery 111, no. 3 (2025): 2546–2557, 10.1097/JS9.0000000000002223. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 21. Pagano S., Strumolo L., Michalk K., et al., “Evaluating ChatGPT, Gemini and Other Large Language Models (LLMs) in Orthopaedic Diagnostics: A Prospective Clinical Study,” Computational and Structural Biotechnology Journal 28 (2025): 9–15, 10.1016/j.csbj.2024.12.013. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22. Rydzewski N. R., Dinakaran D., Zhao S. G., et al., “Comparative Evaluation of LLMs in Clinical Oncology,” NEJM AI 1, no. 5 (2024): 2300151, 10.1056/aioa2300151. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 23. Gao M., Varshney A., Chen S., et al., “The Use of Large Language Models to Enhance Cancer Clinical Trial Educational Materials,” JNCI Cancer Spectrum 9, no. 2 (2025): pkaf021, 10.1093/jncics/pkaf021. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 24. Attal R., Farzad A., Geva K., and Hack S., “Can AI Fix the Consent Form? Promise, Pitfalls, and Practical Realities,” World Journal of Surgery 49, no. 6 (2025): 1394–1395, 10.1002/wjs.12629. [DOI] [PubMed] [Google Scholar]
  • 25. Wang Y. and Ma Z., “Ethical and Legal Challenges of Medical AI on Informed Consent: China as An Example,” Developing World Bioethics 25, no. 1 (2025): 46–54, 10.1111/dewb.12442. [DOI] [PubMed] [Google Scholar]
  • 26. Moulaei K., Akhlaghpour S., and Fatehi F., “Patient Consent for the Secondary Use of Health Data in Artificial Intelligence (AI) Models: A Scoping Review,” International Journal of Medical Informatics 198 (2025): 105872, 10.1016/j.ijmedinf.2025.105872. [DOI] [PubMed] [Google Scholar]
  • 27. Sun Q. W., Miller J., and Hull S. C., “Charting the Ethical Landscape of Generative AI‐Augmented Clinical Documentation,” Journal of Medical Ethics 1(2025): jme‐2024‐110656, 10.1136/jme-2024-110656. [DOI] [PubMed] [Google Scholar]
  • 28. Germani F. and Spitale G., “Source Framing Triggers Systematic Bias in Large Language Models,” Science Advances 11, no. 45 (2025): eadz2924, 10.1126/sciadv.adz2924. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 29. Wang C., Liu S., Yang H., Guo J., Wu Y., and Liu J., “Ethical Considerations of Using ChatGPT in Health Care,” Journal of Medical Internet Research 25 (2023): e48009, 10.2196/48009. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 30. Liang M., “Ethical AI in Medical Text Generation: Balancing Innovation With Privacy in Public Health,” Frontiers in Public Health 13 (2025): 1583507, 10.3389/fpubh.2025.1583507. [DOI] [PMC free article] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Supporting File 1

WJO2-9999-0-s001.pdf (890.5KB, pdf)

Supporting File 2

WJO2-9999-0-s003.png (807.8KB, png)

Supporting File 3

WJO2-9999-0-s002.docx (16.7KB, docx)

Data Availability Statement

The data that support the findings of this study are available from the corresponding author upon reasonable request.


Articles from World Journal of Otorhinolaryngology - Head and Neck Surgery are provided here courtesy of Chinese Medical Association

RESOURCES