Abstract
Purpose:
To study the correctness, completeness, language and readability, and real-world applicability of artificial intelligence chatbots-generated informed consent forms for various ophthalmological procedures and interventions.
Methods:
A cross-sectional observational study was performed by ophthalmology faculties of a tertiary care eye hospital. A list of popularly performed ophthalmological interventions in ophthalmological operation theaters was compiled. Questions were created asking for informed consents. Each question was standardized; the age and diagnosis were mentioned, which were eventually fed into two publicly available chatbots, namely, ChatGPT 4o and Deepseek. The answers obtained from these chatbots were then evaluated on the basis of correctness, completeness, language and readability, additional relevant information, irrelevant information, and real-world applicability of the consent (word to word) in Indian Scenario. Chi-square tests were used for performing analysis of categorical data, namely, correctness and completeness, whereas Mann–Whitney U test was performed for numerical data.
Results:
ChatGPT had less words and sentences compared to Deepseek; however, Deepseek offered a higher average readability score on both Flesch Kincaid calculator and Gunning Fog Index. Deepseek required more attempted to obtain the responses. However, 40% of the consents generated by both chatbots were not fit to be used in Indian scenarios.
Conclusion:
Deepseek offered significantly more elaborate readable informed consents than ChatGPT; however, both the chatbots at present failed 40% of the times to create informed consents which can be used in Indian scenarios.
Keywords: AI, Chatgpt, Deepseek, informed consent
The revolution of artificial intelligence (AI) has taken the world by storm. In the recent past, there has been a progressive increase in research and implementation of AI in various spheres of the healthcare system.[1]
AI models are developed using algorithms, to form different neural networks. The primary objective of these neural networks is to learn from a data base and execute certain specific task as prompted.[2]
One of the most frequently used AI tools in day-to-day life by a vast majority of population are the AI chatbots. There are multiple chatbots which are readily available on the Internet, which can be installed on mobile smartphones or be used online to answer questions. These chatbots have added a new dimension of ease, accessibility, and time reduction for getting any form of information.
Obtaining an informed consent is an extremely crucial step for any form of medical intervention. The formulation of the consents considers multiple factors, which includes other treatment options, side effects of the advised treatment, and so on. Each consent is usually read and explained to the patient (in case of adults) or the guardian (in case of minors or individuals who are not fit to provide consents) prior to the signing. The process of signing also involves presence of witness, whose signatures are also required along with the doctors who are going to perform the procedure.
Preparing and framing of consents is a tedious process, and it is considered as a legal document which protects the medical personal in situations when the patient or the family members plans to sue the medical personal.
There are existing Ophthalmology consent templates designed by the All India Ophthalmological Society (AIOS), readily available for download and use from online websites. However, searching for the document and locating the exact procedure consent is a time-taking process. AI chatbots offer extremely fast answers and specific answers to questions prompted to it. We felt that ophthalmology consent templates can be quickly generated and improved further in terms of customization as per patient profile with the help of AI chatbots. However, we noted that there are no studies in the literature where ophthalmology consents have been formulated by the help of AI chatbots.
Therefore, to test customized consents generated by AI chatbots against existing standardized AIOS consent templates, we conducted a study exploring the possibility of using AI chatbots in preparing informed consent forms for ten popular ophthalmological procedures.
Methods
This was a cross-sectional observational study performed at a tertiary care eye hospital. A list of ten most commonly performed ophthalmological interventions in ophthalmological operation theaters was compiled. No separate grouping was done for subspecialties, nor was any attempt made to stratify based on disease complexity or risk. The intent was to understand how the chatbots create consent for any ophthalmological surgical procedure.
Clinical scenarios were formulated asking for informed consents. Each question was standardized; the age and diagnosis were mentioned. A date was decided in the month of January 2025 for obtaining the information from the chatbots. On the decided date, questions were eventually fed into two publicly available chatbots, namely, ChatGPT 4o and Deepseek [Table 1].
Table 1.
List of questions asked as prompts to the chatbots
| Sl number | Questions |
|---|---|
| 1. | Write an informed consent form for a 45-year-old male hypertensive patient about to undergo phacoemulsification of right eye cataract under topical anesthesia. |
| 2. | Write an informed consent form for a 12-year-old male individual with right eye post-traumatic corneal perforation about to undergo corneal perforation repair under general anesthesia. |
| 3. | Write an informed consent form for a 35-year-old female diabetic with right eye bacterial perforated corneal ulcer about to undergo therapeutic penetrating keratoplasty under local anesthesia. |
| 4. | Write an informed consent form for a 56-year-old male with right eye proptosis, about to undergo orbitotomy under general anesthesia. |
| 5. | Write an informed consent form for an 8-year-old child with intermittent divergent squint, about to undergo squint surgery under general anesthesia. |
| 6. | Write an informed consent form for a 21-year-old female with both eye myopia about to undergo both eye implantable Collamer lens implantation under local anesthesia. |
| 7. | Write an informed consent form for a 40-year-old male with both eye myopia and left eye rhegmatogenous retinal detachment about to undergo pars plana vitrectomy with silicon oil injection under local anesthesia. |
| 8. | Write an informed consent form for a 4-week-old child with retinopathy of prematurity planned for intravitreal injection under general anesthesia. |
| 9. | Write an informed consent form for a 42-year-old diabetic male with both eye nonproliferative diabetic retinopathy and left eye clinically significant macular oedema, planned for left eye injection of Anti-VEGF under topical anesthesia. |
| 10. | Write an informed consent for a 52-year-old diagnosed with right eye primary angle closure glaucoma planned for eye trabeculectomy under local anesthesia. |
Chat GPT (Generative Pre-trained Transformer) is an AI chatbot by the company OpenAI, which is publicly available. It is one of the most commonly used AI chatbots used globally.
Deepseek is an AI software company which has gained immense attention in view of its economical nature compared to its peers.
The informed consents obtained from these chatbots were then evaluated on the basis of correctness, completeness, language and readability, additional relevant information, irrelevant information, and real-world applicability of the consent (word to word). Correctness, completeness, additional relevant data, and irrelevant data were based on the standard All India Ophthalmological Society (AIOS) informed consent forms in English language. Language and readability were evaluated on the grounds of total number of words, sentences, and Flesch–Kincaid ease score,[3] Kincaid grade level,[3] and Gunning fog index.[4] The answers were masked and then given to the graders for grading. Ophthalmologists (three senior residents) were used as graders. We limited the study to ophthalmologists as they had received training and experience in the specialty. Ophthalmologists can better understand the nuances, technicalities, and exact steps involved in the surgical procedures. Also, we wanted to understand whether an ophthalmologist deemed the AI-generated informed consents fit for real-world application. The graders were masked regarding the origin of the consent forms (in terms of which chatbot it was generated). No masking was done between standardized AIOS consent templates which were used as reference.
This being a pilot study, the average scores were used. In case of any dispute regarding intergrader variability in categorical datasets, the most common answer was used for evaluation.
Consent forms generated by both the chatbots were first copied in MS word and assessed separately by three different ophthalmologists. Inter-rater agreement assessment was performed using Fliess’ Kappa test for categoric parameter, namely, correctness, completeness, and real-world applicability. Weighted Cohen’s Kappa was performed for ordinal parameter: irrelevant data and additional relevant data. Real-world applicability of the consents was assessed on the basis of the experience of the ophthalmologists. The study also checked the number of attempts required for obtaining the information and additional comments.
Chi-square tests were used for performing analysis of categorical data, namely, correctness and completeness, whereas Mann–Whitney U test was performed for numerical data.
The study included ten ophthalmological interventions which are usually performed in the operation theaters. The procedures were selected as the commonly performed surgeries; however, it was ensured that the procedures, their steps, side effects, complications, alternative options, and postoperative follow-ups are well known and documented in the literature [Tables 2 and 3].
Table 2.
Scoring of Answers by AI model: ChatGPT 4o
| Cataract Surgery | Traumatic Corneal Perforation repair | Penetrating Therapeutic Keratoplasty | Orbitotomy | Squint Surgery | ICL implantation | Pars Plana Vitrectomy | Intravitreal injection for ROP | Intravitreal injection for DME | Trabeculectomy | |
|---|---|---|---|---|---|---|---|---|---|---|
| Correctness (Correct/Incorrect) | Correct | Correct | Incorrect | Correct | Correct | Correct | Correct | Correct | Correct | Correct |
| Completeness (Complete/Incomplete) | Incomplete | Complete | Incomplete | Incomplete | Complete | Incomplete | Incomplete | Incomplete | Complete | Complete |
| Language and readability -Total sentences -Total words -Flesch Reading Ease -Flesch Kincaid Grade level -Gunning Fox Index |
23 438 23.93 14.56 16.02 |
22 446 29.36 14.24 15.97 |
21 473 22.60 14.91 16.91 |
21 405 30.64 13.80 16.62 |
21 394 35.92 12.95 15.66 |
21 425 32.42 13.73 15.29 |
21 431 23.08 15.14 17.56 |
20 385 31.63 13.64 16.29 |
27 412 25.66 14.57 17.19 |
21 421 24.66 14.81 15.81 |
| Irrelevant data (None/Some/Many) | None | None | None | None | None | None | None | None | None | None |
| Additional relevant data (None/Some/Many) | Some | Some | Some | None | Some | None | Some | Some | Some | Some |
| Real world applicability (word to word) (Possible/Not possible) | Not possible | Possible | Not possible | Possible | Possible | Possible | Possible | Not Possible |
Possible | Not Possible |
| Attempts required for getting response | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 |
Table 3.
Scoring of Answers by AI model: Deepseek v3
| Cataract Surgery | Traumatic Corneal Perforation repair | Penetrating Therapeutic Keratoplasty | Orbitotomy | Squint Surgery | ICL Implantation | Pars Plana Vitrectomy | Intravitreal injection for ROP | Intravitreal injection for DME | Trabeculectomy | |
|---|---|---|---|---|---|---|---|---|---|---|
| Correctness (Correct/Incorrect) | Correct | Incorrect | Correct | Correct | Correct | Correct | Correct | Correct | Correct | Correct |
| Completeness (Complete/Incomplete) | Incomplete | Complete | Complete | Incomplete | Complete | Complete | Complete | Incomplete | Complete | Incomplete |
| Language and readability -Total sentences -Total words -Flesch Reading Ease -Flesch–Kincaid grade level - Gunning Fog Index |
49 648 45.24 10.17 12.04 |
47 650 44.50 10.59 12.49 |
53 668 43.55 10.36 13.17 |
66 594 46.45 9.02 12.91 |
43 652 48.13 10.32 12.70 |
51 624 40.09 10.71 12.99 |
43 525 39.01 11.39 13.35 |
27 499 36.50 12.6 12.98 |
43 565 34.78 11.67 14.41 |
54 667 49.14 9.36 10.94 |
| Irrelevant data (None/Some/Many) | Some | none | none | None | Some | None | None | None | None | None |
| Additional relevant data (None/Some/Many) | Some | Some | some | None | Some | None | Some | Some | Some | None |
| Real world applicability (word to word) (Possible/Not possible) | Not possible | Not possible | Possible | Possible | Possible | Possible | Possible | Not Possible |
Possible | Not Possible |
| Attempts required for getting response | 1 | 5 | 5 | 7 | 6 | 4 | 3 | 4 | 3 | 2 |
Results
Correctness, completeness, and additional relevant and irrelevant data
Informed consents in ophthalmic surgeries (English) published by All India Ophthalmological Society in 2021 were used as references for assessing correctness, completeness, and additional relevant and irrelevant data. We did not encounter too many intergrader discrepancy in categorical datasets. On Fleiss Kappa agreement, there was perfect agreement for correctness, and substantial agreement in completeness and real-world applicability among the answers for both Deepseek and ChatGPT4o consents. On Weighted Cohen’s Kappa, perfect agreement was noted among the grader for parameters of irrelevant data and additional relevant data for both Deepseek and ChatGPT4o.
Correctness: 9/10 informed consents were correct as generated by both chatbots (Deepseek and ChatGPT 4o). Chatgpt 4o created an incorrect informed consent form for penetrating therapeutic keratoplasty, whereas Deepseek generated an incorrect informed consent for post-traumatic corneal perforation repair. There was no significant difference between the two chatbots on Chi-square test (P value = 1).
Completeness: While evaluating completeness, we noted that 4/10 Chatgpt 4o generated consents were complete, where 6/10 informed consents generated by Deepseek were complete in terms of points covered. There was no significant difference between the two on Chi-square test (P value = 0.89).
Additional relevant data: In 8 out of 10 informed consents generated by Chatgpt 4o, and in 7 out of 10 Deepseek-created informed consents, there was additional relevant information which was not there in the AIOS consents. However, there was no specific pattern in the type of procedures where additional relevant data were generated. A few specific points were noted: AI-generated consents had subheadings of “Expected benefits” and “Alternate treatment options,” which were not there in certain AIOS consents. There was also subheading of “Patients responsibilities” focusing on postoperative instruction and additional information in cases the patients had a systemic condition. For children, automatically the subheadings were changed to “Guardian responsibilities” and the additional subheading of “special considerations due to age” was added. These were not there in AIOS consent template.
Irrelevant data: None of the informed consents by Chatgpt 4o had any irrelevant data, whereas 2/10 of the informed consents by Deepseek had irrelevant information.
Language and readability
The average total words and sentences in ChatGPT 4o were 423 and 21.8, respectively, whereas for Deepseek, it was 609.2 and 47.6, respectively. There was a significant difference (P value 0.00018 for words, 0.0012 for sentences). Deepseek had more sentences [Fig. 1].
Figure 1.

Graphical representation of average number of words and sentences by ChatGPT 4o and Deepseek v3
Readability was assessed by using Flesch Reading Ease and Flesch–Kincaid grade level, which revealed a significant difference (P value = 0.00025); there was also a significant difference in the Gunning fog index with the informed consent generated by Deepseek AI chatbot being more readable.
There was a significant difference in the number of attempts required by each chatbox in generating the informed consents. Deepseek required a higher number of attempts.
Real-world applicability of these consents in Indian scenario
Three different ophthalmologists evaluated the consent forms generated by the chatbots separately; according to them, 4/10 informed consents generated by both Deepseek and ChatGPT were not suitable for Indian scenarios.
Discussion
AI chatbots are gaining immense popularity with more passage of time.
They are being tested in answering frequently answered questions in various subjects.[5,6,7,8] Rokshad et al. studied the efficacy and empathy of AI chatbots in answering frequently asked questions on oral oncology and found Chat GPT 4 demonstrated the quality of responses. The authors have studied Chat GPT4, GPT3, Bing, Google Bard, and Claude.[1]
Bahir et al.[5] conducted a study wherein the authors compared four AI models. ChatGPT-3.5, ChatGPT-4, Gemini, and Gemini advanced were tested using a dataset of 600 questions which were taken from Israeli Ophthalmology residency exams. The authors noted that Gemini advanced elicited the highest accuracy rate.
Gill et al.[6] performed a comparative study between Gemini advanced and ChatGPT 4. They used 260 randomly generated ophthalmology questions from a question bank in the assessment. The authors concluded that ChatGPT 4.0 performs better in OKAP style exams.
There is enough literature which supports that ChatGPT is one of the most advanced available chatbots currently in use.
ChatGPT uses a transformer model architecture, which uses all of its 1.8 trillion parameters for every task. Its training emphasizes versatility in language generation and creative task. The model has been trained on a massive dataset of text and code.[7]
Deepseek uses a Mixture of Experts architecture, where though it has 671 billion parameters, only a small subset of them is activated for each query. This makes it more efficient models that use all parameters for every task. It uses reinforcement learning to enhance its reasoning abilities.[7] [Fig. 2]
Figure 2.

Pictorial representation of differences between Deepseek v3 and ChatGPT 4o
A few studies have been performed comparing correctness, completely, and real-world application of answers by Deepseek, Chatgpt 4o, and Gemini in context of ocular oncology; however, there have been no studies on preparation of informed consents by the same.[8] Deepseek is a relatively new AI chatbot, with extremely limited studies mentioning its usage. In our study, we noted Chat GPT was able to generate informed consents faster than Deepseek.
Various calculators are used for assessing readability of texts. Kincaid ease score is one of the most widely accepted readability ease calculator. The values obtained from the calculator are directly proportional to the readability of the text. When it comes to readability of informed consents, ideally, they should be as simple and complete as possible. Gunning fog index is another popular readability index where a lower index value indicates a better readability.
The consents generated by Deepseek were more readable than those generated by ChatGPT. The average number of words and sentences in each consents generated by Deepseek was significantly more than that of ChatGPT. Reading a lengthier consent often times can be time-taking and problematic.
The number of attempts required to generate informed consents from Deepseek was almost 5 times more than that needed by ChatGPT 4o.
We also noted that 40% of the consents produced by both the groups were deficient in some degree when compared to that of AIOS informed consents in English language. However, as this is the first of its kind study assessing the credibility of AI chatbot-generated ophthalmology consent forms, we feel that in the current day, the AI chatbots are not completely capable of generating usable informed consents. This begs the question, should we be even using the AI chatbots in generating informed consents or not? However, both the AI chatbots offered the options of editing the already prepared consents by further prompts. So the consents generated can be customized more as per the requirement.
Informed consents are not just legal checkboxes; they are ethical commitment to respect autonomy and promote shared decision-making in patient care. Most legal systems identify them as a legal duty of the healthcare provider, and conducting a procedure without a valid consent can be considered battery, or malpractice. AI-generated consents, though not at present but in future, will be fulfilling all the requisite criteria and possibly be safeguarding the patient’s rights and the physicians’ legal interests.
To assess consistency, prompts were repeated using three different user accounts, consistently yielding similar outputs. While these preliminary observations suggest reproducibility, a larger, more comprehensive study is needed to confirm the robustness of AI-generated responses.
Conclusion
Consents prepared by Deepseek were more readable, longer, and elaborate. However, the number of attempts required to get the response was multiple times more than that of ChatGPT4o. To the best of our knowledge, this is first study comparing the use of Deepseek against ChatGPT in generating ophthalmology informed consents. When compared to existing AIOS informed consents, 40% of the AI-generated consents failed to cover all the aspects, which indicates that the database for information of these AI chatbots needs to be upgraded.
Synopsis
Deepseek v3 generated more elaborate and readable informed consents for ophthalmology procedures than ChatGPT 4.o. However, there are limitations of using AI chatbots in creating informed consents. As per our study, the real-world applicability is a point of concern.
Conflicts of interest:
There are no conflicts of interest.
Funding Statement
Nil.
References
- 1.Rokhshad R, Khoury ZH, Mohammad-Rahimi H, Motie P, Price JB, Tavares T, et al. Efficacy and empathy of AI chatbots in answering frequently asked questions on oral oncology. Oral Surg Oral Med Oral Pathol Oral Radiol. 2025 doi: 10.1016/j.oooo.2024.12.028. S2212-4403(25)00002-1. doi: 10.1016/j.oooo.2024.12.028. [DOI] [PubMed] [Google Scholar]
- 2.Pandey VK, Munshi A, Mohanti BK, Bansal K, Rastogi K. Evaluating ChatGPT to test its robustness as an interactive information database of radiation oncology and to assess its responses to common queries from radiotherapy patients: A single institution investigation. Cancer Radiother. 2024;28(3):258–64. doi: 10.1016/j.canrad.2023.11.005. [DOI] [PubMed] [Google Scholar]
- 3. Available from: https://serpninja.io/tools/flesch-kincaid-calculator/
- 4. Available from: http://gunning-fog-index.com .
- 5.Bahir D, Zur O, Attal L, Nujeidat Z, Knaanie A, Pikkel J, et al. Gemini AI vs. ChatGPT: A comprehensive examination alongside ophthalmology residents in medical knowledge. Graefes Arch Clin Exp Ophthalmol. 2024 doi: 10.1007/s00417-024-06625-4. doi: 10.1007/s00417-024-06625-4. [DOI] [PubMed] [Google Scholar]
- 6.Gill GS, Tsai J, Moxam J, Sanghvi HA, Gupta S. Comparison of Gemini advanced and ChatGPT 4.0’s performances on the ophthalmology resident Ophthalmic Knowledge Assessment Program (OKAP) examination review question banks. Cureus. 2024;16:e69612. doi: 10.7759/cureus.69612. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 7. Available from: https://www.datacamp.com/blog/deepseek-vs-chatgpt .
- 8.Das D, Narayan A, Mishra V, Takia L, Grover S, Bharati A, et al. AI chatbots in answering questions related to ocular oncology: A comparative study between DeepSeek v3, ChatGPT-4o, and Gemini 2.0. Cureus. 2025;17:e90773. doi: 10.7759/cureus.90773. [DOI] [PMC free article] [PubMed] [Google Scholar]
