Skip to main content
NIHPA Author Manuscripts logoLink to NIHPA Author Manuscripts
. Author manuscript; available in PMC: 2026 Apr 27.
Published in final edited form as: Patient Educ Couns. 2026 Jan 17;145:109492. doi: 10.1016/j.pec.2026.109492

Patient and clinician engagement with generative artificial intelligence (GenAI): A scoping review of implications for patient-centered communication

Jessica Hahne 1,*, Brian D Carpenter 1
PMCID: PMC13110415  NIHMSID: NIHMS2163409  PMID: 41570489

Abstract

Objective:

To examine the influence of the emerging use of generative artificial intelligence (GenAI) within electronic health records and among the public on the patient-centeredness of communication in healthcare.

Method:

In this scoping review, we conducted a systematic search for peer-reviewed studies in PubMed and PsycInfo that empirically examined GenAI involvement in clinical communication. We then mapped study findings onto a well-established framework for patient-centered communication.

Results:

Our search yielded 67 studies for analysis. Results suggest that integration of GenAI into healthcare communication has the potential to increase clinician efficiency in interacting with patients, to expand channels for patients to obtain information about their healthcare, and to enhance empathy in clinical communication. However, findings also indicate variability in the quality of information produced by GenAI, the potential for GenAI to recast the clinician as a technical supervisor rather than a humanistic care provider, and several issues of equity and privacy raised by engagement with GenAI.

Conclusion:

As GenAI becomes more prevalent in healthcare, rigorous examination of GenAI is needed to ensure that its development and implementation aids rather than hinders patient-centered communication. We conclude with an agenda for further research on GenAI grounded in the PCC framework underlying our review.

Practice Implications:

Findings from this review highlight the current potential benefits and limitations of GenAI as a third party to clinical communication. Continued efforts toward developing and applying GenAI for effective healthcare communication should focus on protecting patients from potential drawbacks and maximizing nascent benefits for patient-centered communication.

Keywords: Generative artificial intelligence, Patient-centered communication, Electronic health records, Digital health technology

1. Introduction

Technological advancements have historically shaped not only medical interventions but also communication between clinicians and patients. The expansion of the internet in the 1990s and early 2000s accelerated patient access to medical information with varying degrees of credibility outside of clinical visits, a phenomenon often referred to as “Dr. Google.” Likewise, starting in 2009, changes to U.S. federal policy known as “meaningful use criteria” incentivized and accelerated use of electronic health records (EHRs) across the country. As EHRs spread, online platforms known as patient portals began providing patients direct access to information from EHRs. In some countries such as the United States, patient portals and EHRs have become highly interoperable, such that clinical notes, test results, secure messaging between patients and clinicians, and other care information are accessible on the clinician side of care through the EHR and on the patient side as the EHR-linked patient portal. Thus, particularly in countries with high EHR-patient-portal interoperability, EHRs and their linked portals have become platforms for not only passive forms of communication between patients and clinicians (such as patients reading notes written by clinicians) but also active forms of communication (such as secure messaging).

More recently, generative artificial intelligence (GenAI), a category of algorithms that can simulate human intelligence to learn from patterns of data and create new content and ideas [1], has been advancing exponentially in forms such as internet chatbots and AI-powered EHR features [2], raising new implications for healthcare communication. Indeed, Epic, the largest EHR vendor in the United States, announced a partnership with Microsoft in 2023 to integrate GenAI into EHR software [3]. Since then, GenAI has been used by clinicians at health systems piloting the partnership to draft over 150,000 visit notes and millions of responses to patient questions over secure messages [4]. Beyond EHR integration pilot programs, 38 % of U.S. physicians report having used some form of publicly or institutionally available AI in medical practice — most commonly for documentation, translation, or diagnosis [5].

On the other side of clinical communication, patients are starting to seek answers to health questions from ChatGPT and other Large Language Model (LLM) chatbots. In 2022, 58.5 % of American adults reported having used the internet in the past year to look for health information [6]. And in one of the first investigations of how broader internet-based health information seeking involves LLMs, a 2025 survey found that 52 % of American adults had used an internet-based LLM for any reason at least once, and 39 % reported having asked an LLM for information about physical or mental health [7].

What these technological trends mean for patient-provider communication is still an evolving area of investigation. Patient-centered communication (PCC) is often suggested as an ideal approach to healthcare interactions. Although exact definitions vary, PCC generally calls for clinicians to view patients as persons with unique psychosocial contexts, to elicit and attend to patients’ values and preferences, and to offer patients meaningful involvement in decision-making [8–11]. Previous research has examined how use of EHRs influences the patient-centeredness of clinical communication [12]. As technology continues to evolve, there is a need to explore how clinician and patient engagement with GenAI might further change dynamics of healthcare communication. Therefore, the purpose of this study is to review and synthesize current peer-reviewed empirical research on clinician and patient engagement with GenAI in the context of healthcare communication, analyzing how GenAI detracts from or enhances the patient-centeredness of clinical communication. To our knowledge, this review is the first of its kind to apply a structured framework for PCC to examine the specific communication tasks between patients and clinicians in which GenAI may now play a role, as well as the current state of the research on how GenAI influences specific aspects of PCC, including both preliminary findings and remaining gaps concerning each aspect.

1.1. A framework for patient-centered communication

PCC is associated with a number of positive health outcomes, including improved pain control [13], decreased hospital length-of-stay [14], and higher success in rehabilitation and illness self-management [15,16]. Evidence also suggests that PCC increases patient and family satisfaction [14,16] and decreases patient anxiety [17]. Although definitions of PCC have varied across studies, one widely cited framework for PCC was proposed by Epstein & Street [9]. Their framework cited concepts from earlier theories that explored patients’ subjective understanding of the illness experience (such as self-regulation theory [18], explanatory model theory [19], and uncertainty in illness theory [20]), as well as theories emphasizing the role of patient autonomy in medical decision-making (such as shared decision-making [21] and self-determination theory [22]). The Epstein and Street model outlines six specific aspects, or “functions,” of PCC: fostering healing relationships, exchanging information, responding to emotions, managing uncertainty, making decisions, and enabling self-management. In Table 1, we describe each function according to the definitions offered by Epstein and Street [9].

Table 1.

Functions of Patient-Centered Communication.

Function Definition [9]
Fostering healing relationships Establish mutual trust and rapport, as well as mutual understanding about the roles of clinician and patient in care. Take initiative to facilitate open communication and patient/family involvement in care.
Exchanging information Recognize and respond to the patient’s information needs. Provide accurate information about diagnosis, treatment, and prognosis. Address barriers to understanding and integrate clinical information with the patient’s understanding of illness.
Responding to emotions Elicit the patient’s emotional distress and respond with empathy and support.
Managing uncertainty Acknowledge uncertainty and help patient to manage it by providing support and providing information that is available.
Making decisions Facilitate the involvement of the patient/family in care decision-making processes, to the degree that the patient prefers to be involved. Make decisions based on patient’s values and preferences, as well as accurate clinical information.
Enabling self-management Guide the patient in terms of navigating the healthcare system, finding information, developing coping skills, and taking actions to manage and improve their health.

1.2. EHRs and patient-centered communication

At the time that their framework was established in 2007, Epstein and Street [9] encouraged clinician use of information technology, particularly within the function of “exchanging information.” Over the decade that followed, EHRs shifted from being a supplementary information technology to becoming the medium through which much of modern clinical communication occurs. By 2017, 51 % of patient portal users in the United States reported viewing EHR notes as part of their regular use of portals, and 48 % reported using the portal to message clinicians [23].

A 2017 scoping review by Rathert and colleagues [12] examined effects of EHR use on patient-centeredness of clinical communication, using Epstein and Street’s [9] framework. The authors identified several positive outcomes associated with communication through the EHR, including increased patient engagement and collaborative involvement in healthcare, increased clinician attention to the accuracy of patient history, and improved patient health knowledge. The authors also identified negative outcomes for PCC related to the advent of EHRs, including new potential for clinician data entry errors, decreased clinician eye contact with patients due to gaze directed at computer screens, and decreased patient trust toward clinicians due to concerns about privacy being compromised by EHRs.

Since these findings, use of EHRs has continued to increase. In the U. S., policy changes incrementally pushed for wider use of EHRs, beginning with “meaningful use criteria” that evolved over the 2010’s and culminated in the 21st Century Cures Act in 2021, which required all U. S. healthcare systems to share electronic visit notes with patients free of charge [24]. As these policies unfolded, the proportion of U.S. adults accessing online medical records increased from 27 % in 2017 to 57 % in 2022 [25]. The proportion of patient portal users who reported sending secure messages to clinicians increased from 48 % in 2017 to 58 % in 2020 [23].

1.3. Aims of this review

Previous research has evaluated how prior technological changes such as growing use of EHRs in healthcare affected the patient-centeredness of communication. Now, GenAI is introducing a new layer of technology to healthcare communication, as patients and clinicians are increasingly engaging with GenAI during communication tasks conducted through EHRs (such as drafting visit notes or secure messages) as well as communication tasks adjacent to EHRs (such as asking GenAI for advice on diagnostic decisions or illness management).

Despite mounting engagement with GenAI in healthcare, attitudes among both patients and clinicians about AI are mixed. In a 2022 survey conducted by Pew Research Center, seventy-five percent of a nationally representative sample of over 1000 Americans endorsed concern that clinicians would adopt AI too quickly for healthcare tasks before fully understanding risks [26]. And a 2023 survey of 1000 physicians by the American Medical Association showed clinicians have their own concerns. Forty-one percent reported being equally excited and concerned about potential increased use of AI in healthcare [5]. Besides impact to patient privacy, the second-most common concern endorsed by physicians regarding AI in healthcare was the potential impact on the patient-physician relationship [5]. As research and practice in this dynamic area progress, it is important to have a systematic understanding of how GenAI may affect PCC. To our knowledge, no previous effort has been made to evaluate GenAI’s emerging influence on patient-clinician communication using a theory-based framework for PCC. Therefore, our scoping review has the following three aims:

  1. To identify specific communication tasks conducted by clinicians and patients that may now involve GenAI.

  2. To evaluate the patient-centeredness of communication when GenAI is involved in those communication tasks.

  3. To identify current gaps in knowledge on GenAI and PCC to guide future research.

2. Method

With the goal of “mapping” available evidence from the literature onto a theory-based framework, we considered a scoping review to be the most appropriate methodology for this study [27]. As opposed to a systematic review, which would typically focus on answering a narrower, more focused question in a body of literature in which gaps have already been well-identified, we chose a scoping review methodology, which allows for mapping the use of concepts (such as the six functions of patient-centered communication) across an emerging, cross-disciplinary body of literature [28]. Our review followed the five-stage process for scoping reviews outlined by Arksey and O’Malley [27], as well as the reporting guidelines of the Preferred Reporting Items for Systematic Reviews and Meta-Analyses extension for Scoping Reviews (PRISMA-ScR) [29].

2.1. Search strategy

Our strategy began by searching PubMed and PsycINFO with three key concepts, both individually and in combination: “patient-centered communication,” “electronic health records” (in order to include forms of written patient-physician communication through EHRs and patient portals), and “generative AI.” MeSH and keywords from initial searches were reverse-searched to find additional keywords relevant to the three main concepts. In order to maximize the number of relevant articles returned by searches, our finalized list of search terms combined the first of these two main concepts into one inclusive group of terms and was thus centered on the concepts of “patient-centered communication OR electronic health records” AND “generative AI.” The finalized set of search terms used in each of the two databases to search for these concepts is included in Appendix A. The finalized search was restricted to all articles published in English on or before the date of the search, May 12, 2024, which yielded 2218 articles after removal of duplicates.

2.2. Eligibility criteria and screening

In order to be eligible for inclusion in this review, articles were required to meet the following criteria: (1) peer-reviewed, empirical research, (2) pertaining to clinical healthcare communication, and (3) involving some form of GenAI. Common reasons for exclusion of articles included: (1) editorial, review article, or not peer-reviewed; (2) about a topic peripheral to clinical communication such as veterinary medicine, education of healthcare professionals, or medical research not involving patient care; and (3) about artificial intelligence that is not classified as GenAI, such as rule-based chatbots or early forms of dictation software.

During initial abstract screening, the first author reviewed all titles and abstracts to determine whether they met eligibility criteria, using the online review platform Rayyan [30]. All articles that met eligibility criteria or that could not be fully assessed based on the title and abstract were selected for full-text screening. The first author screened the full text of 92 articles against the eligibility criteria and examined additional relevant articles from reference sections, resulting in a final total of 67 articles included for analysis. A PRISMA flow diagram depicting the search and screening process is displayed in Fig. 1.

Fig. 1.

Fig. 1.

PRISMA Flow Diagram of Article Screening Process.

2.3. Data extraction and synthesis

The first author extracted information from each eligible article, including country of origin, author(s), year of publication, study objective, study design, type(s) of GenAI examined in the study, the role that AI played in clinician-patient communication in the study, function (s) of patient-centered communication for which the study had implications, positive or neutral results for patient-centered communication, and negative results for patient-centered communication. We provide a complete table of extracted data in Appendix B. Because none of the included studies explicitly tested the functions of PCC outlined by Epstein and Street [9], communication functions that were addressed in each article were identified and analyzed by the first author, in the process of reading the full text of articles. Finally, we summarized the findings of this review by narrative synthesis, organized according to implications for the six functions of patient-centered communication.

3. Results

3.1. Characteristics of reviewed studies

Of 67 studies reviewed, nearly two-thirds took place in the U.S. (41; 61.2 %). The remaining studies originated from 14 other countries (see Appendix B for extracted characteristics of all articles). Most (86.6 %) were cross-sectional observational studies, five (7.5 %) were longitudinal, and four (6.0 %) tested interventions. The most common type of GenAI examined were GPT-family models (76 % of studies). Frequencies of specific models examined are displayed in Table 2.

Table 2.

GenAI Models Examined by Included Studies.

GenAI Model (Developer) Studies in Review n (%)
Claude-2 (Anthropic) 2 (3.0 %)
ERNIE Bot, version 2.2.3 (Baidu, Inc.) 1 (1.5%)
Bard (Google) 8 (11.9%)
BERT (Google) 2 (3.0 %)
Gemini, formerly Bard (Google) 1 (1.5 %)
Pegasus-xsum (Google) 1 (1.5 %)
T5-family models (Google) 4 (6.0 %)
BART (Meta) 2 (3.0 %)
LLaMA-family models (Meta) 4 (6.0 %)
Bing (Microsoft) 3 (4.5 %)
Copilot, formerly Bing Chat (Microsoft) 2 (3.0 %)
ChatGPT-3.5 (OpenAI) 28 (41.8 %)
ChatGPT-4.0 (OpenAI) 18 (26.9 %)
ChatGPT, version unspecified (OpenAI) 8 (11.9 %)
GPT-family models, other (OpenAI) 5 (7.5 %)
Perplexity (Perplexity AI) 1 (1.5 %)
Other model (developed by study authors) 10 (14.9 %)
Model not specified 4 (6.0 %)

Note. Because many studies examined more than one model, counts do not add up to 67, and percentages do not add up to 100.

3.2. Overview of findings

Relevant to Aim 1, studies in this review depicted ways in which GenAI is introducing an additional technological layer to communication tasks, through roles such as co-writing notes and results with clinicians, transcribing and summarizing clinician-patient dialogue during medical visits, co-writing clinicians’ responses to patients’ secure messages, and providing medical advice for patients outside of dialogue with clinicians. We summarize these findings in Fig. 2.

Fig. 2.

Fig. 2.

Roles of GenAI in Communication Tasks.

In line with Aims 2 and 3, the following sections summarize results relevant to each of the six PCC functions [9]. Nineteen studies (28.4 %) had implications for fostering healing relationships, 54 (80.6 %) for exchanging information, 11 (16.4 %) for responding to emotions, 17 (25.4 %) for managing uncertainty, 12 (18.0 %) for making decisions, and 16 (24.0 %) for enabling self-management.

3.3. Function 1: fostering healing relationships

Based on the Epstein and Street [9] framework, “healing” clinician-patient relationships are characterized by strong rapport, trust, and mutual respect. Overall, findings from articles in this review suggested that GenAI has mostly positive effects on clinician-patient rapport but threatens to erode patient trust in clinicians and has mixed effects for mutual respect.

3.3.1. Rapport

Study findings varied as to whether using GenAI-generated draft notes and messages saved clinicians time and increased bandwidth for rapport with patients. Two studies found significant time saved compared to writing manually [31,32], one study found no difference [33], and one study found increased time spent [34]. More consistently, clinicians reported subjective benefits, including increased perceived efficiency [32,35] and decreased task load [33]. When GenAI scribes took notes during visits, clinicians felt able to “listen instead of type” [32], and patients thought clinicians were more engaged [32,36].

3.3.2. Trust

Evidence suggests that patients may come to distrust clinicians if clinicians become over-reliant on GenAI for diagnostic decision-making [35,37,38]. Patients somewhat trusted GenAI chatbots to provide medical advice but preferred human clinicians more as complexity of medical questions increased [39]. Notably, patients envisioned a new role for clinicians as technical supervisors responsible for critically evaluating suggestions of GenAI [38].

3.3.3. Respect

GenAI may have mixed effects on respect for patients, considering issues of equity and privacy. Some language models have demonstrated efficient and thorough performance at identifying social determinants of health (SDOH), such as housing status and substance use, from large amounts of text within EHR notes [40,41]. Used in this way, GenAI could help facilitate access to health care and social services [40]. However, GenAI also has potential to multiply biases inherent to its training data. For example, one study asked ChatGPT and other LLMs to synthesize text from EHRs into classifying labels that were added to patient records in order to flag specific SDOH that might benefit from social service interventions (such as employment, housing, or transportation issues). Without careful fine-tuning, LLMs in this study sometimes changed the way that they classified SDOH simply because Hispanic and Black descriptors were injected into the EHR text from which they were extracting information, demonstrating assumptions based on racial/ethnic biases [40].

3.4. Function 2: exchanging information

Clinicians communicating in a patient-centered manner exchange accurate information with patients while also addressing barriers to understanding. Findings from studies in our review suggest that when GenAI assists with information exchange, it exhibits widely variable accuracy but if specifically prompted to do so, can help lower barriers to comprehension such as the reading level of information.

3.4.1. Accuracy

Studies on ChatGPT-family models answering patient questions found that their accuracy varied widely from 54.5 % [42] to 93.9 % [43]. Several studies found that GenAI models demonstrated comparable or better accuracy than physicians at generating patient- or clinician-facing summaries of hospital stays [44,45] or radiologic imaging reports [46–48], but another study found roughly 24 errors per ChatGPT-4.0-generated EHR visit note [49]. Notably, multiple studies showed GenAI accuracy improved after either refining model training data to be more specialized [50] or refining instructions to GenAI models (“prompt engineering”) [51,52].

3.4.2. Readability

Jargon-laden language can impede patient understanding of clinical communication. Across studies [53–57], various LLMs produced patient-facing material at reading levels higher than the sixth-grade-or-lower level recommended for health materials facing the American public [58,59]. Health organization websites [55] or top Google searches [57] featured material of lower, more readable levels (although still above sixth-grade level) than LLM (specifically, ChatGPT) output. However, when ChatGPT was asked to revise output to sixth-grade level, it was able to simplify output to ninth-grade level in one study [60,61] and to sixth-grade level in another [62].

3.5. Function 3: responding to emotions

Responding to patients’ emotions in a patient-centered manner necessitates an approach that is both empathetic and personalized. Findings from this review showed that GenAI typically used more empathetic phrasing than human clinicians but lacked human clinicians’ ability to tailor messaging to individual patients.

3.5.1. Empathy

Studies consistently found that chatbot responses to patient questions were rated highly on empathy [63,64], outscoring the empathy of clinician responses when rated by clinicians [65,66] and patients [67]. Notably, however, one such study also measured patient satisfaction and found no association between patient satisfaction and chatbot authorship of responses [67]. Despite a strong pattern of studies indicating that GenAI uses more empathetic language than most human clinicians, it is unclear, based on this latter finding, whether patients always prefer this volume of empathetic language in responses to their healthcare questions – or whether there may be something qualitatively different about GenAI-authored empathetic language compared to human-authored empathetic language.

3.5.2. Personalization

Clinicians frequently revised GenAI-generated draft responses to patient messages to improve and personalize the tone of messages before sending them to patients [33,34]. GenAI tended to draft lengthy responses laden with empathetic phrases, and clinicians varied as to how appropriate they found this language [33,34].

3.6. Function 4: managing uncertainty

Due to the uncertainty that patients often experience related to treatment outlook and information asymmetry with clinicians, PCC calls for clinicians to provide patients with information that is known about their condition in a relevant, consistent, and comprehensive manner. Studies in our review found that GenAI varied widely in the relevance and consistency of information produced but outscored human clinicians on the comprehensiveness of information in some tasks.

3.6.1. Relevance and consistency

In studies evaluating GenAI chatbot responses to patient messages, clinicians described frequently revising GenAI-drafted responses to change content that was overly generic or irrelevant [33,68]. When answering frequently asked medical questions, GenAI sometimes produced information of nearly perfect clinician-rated relevance [69] but at other times was outperformed by the relevance of clinician-authored answers to the same questions [66]. When chatbots were repeatedly asked the same medical questions in some studies, the consistency of information across trials ranged from 93 % [70] to as low as 72 % [71].

3.6.2. Comprehensiveness

Studies found that information produced by GenAI for clinical communication tasks was only moderately comprehensive or complete [63,72,73]. However, the completeness of clinical summaries (summaries of patients’ longitudinal care) written by GenAI were either equally or more complete than those written by clinicians [44,45,74]. Results also showed that GenAI models trained as general-purpose chatbots improved in clinician-rated comprehensiveness at writing clinical summaries when given more specific prompts that included examples [75].

3.7. Function 5: making decisions

Clinicians using PCC facilitate decision-making that involves the patient and the patient’s family or caregivers, aligns with patients’ values and preferences, and represents sound clinical reasoning. No articles in this review measured how engagement with GenAI might influence patients’ or their loved ones’ degree of involvement in healthcare decision-making, or the alignment of healthcare decisions with patients’ values and preferences. However, numerous findings highlighted how engagement with GenAI can influence the appropriateness of clinicians’ reasoning, from diagnosis through treatment planning. Specifically, GenAI excelled at straightforward differential diagnosis tasks but struggled with making diagnostic decisions that required integration of patient-specific contextual details and with making treatment plans.

3.7.1. Diagnosis

When asked to generate differential diagnoses for patient cases, various versions of ChatGPT have been rated highly by physicians [76] and performed better than formerly developed automated methods [77]. Several studies have found that ChatGPT-4.0 performed similarly to both pediatric [78] and internal medicine [79,80] physicians at identifying correct diagnoses. However, ChatGPT struggled with inferring or interpreting context-specific information about patients that was relevant to making accurate diagnoses [78,80].

3.7.2. Treatment planning

GenAI’s recommendations for treatment and management of conditions were fully concordant with guidelines published by professional societies in only 31.3 % of cases for breast surgery [81] and 41.2 %–82.4 % (differing across models) for gastroesophageal reflux disease [82]. In one study, ChatGPT-3.5’s frequency of incorrect recommendations on medication dosing increased from 53.4 % to 67.3 % when it was asked to consider demographic and other contextual variables [83].

3.8. Function 6: enabling self-management

Clinicians who communicate in a patient-centered manner empower patients’ self-management of their own health, including patients’ ability to find and adhere to information about prevention and practical management of illness, prepare for procedures and understand outcomes, and cope with the difficulties of illness experiences. In studies included in this review, GenAI provided moderately high-quality information to patients preparing for or interpreting results from procedures but widely variable quality of information about prevention and management of illness. Uses of GenAI to assist with coping were understudied.

3.8.1. Prevention and management of illness

In studies that evaluated ChatGPT-family models answering patients’ questions about prevention and self-management of illness, clinician-rated quality of ChatGPT’s answers ranged from moderate to high [60, 63]. Results across LLMs differed in one comparative study, with only 45.8 % of Bard’s answers to vascular surgery disease management questions considered appropriate, in comparison to 75 % of ChatGPT-3.5’s answers [54]. While ChatGPT provided significantly more understandable and actionable answers than Google Search on questions about general medical knowledge (average score of 87 % for ChatGPT versus 78 % for Google Search), it fared significantly worse on these metrics when asked to provide advice on managing conditions (68 % for ChatGPT versus 89 % for Google Search) [84].

3.8.2. Preparing for procedures and understanding outcomes

Clinicians considered the quality of ChatGPT’s responses to range from moderate to high when it was asked to answer questions about preparing for knee surgeries [85], simplify radiology reports, or answer patient questions about imaging scans [48,64,86]. One study measuring the quality of answers to patients’ questions about lab test results found that human responses on Yahoo Answers scored worst, whereas ChatGPT-family models scored best in comparison to other LLMs (LLaMA-2, MedAlpaca, and ORCA_mini) [87].

3.8.3. Coping with illness experiences

Only one study examined GenAI’s influence on patients’ ability to cope effectively with illness. Patients were significantly more likely than their control-group counterparts to continue long-term rehabilitation when a cardiac-rehabilitation-oriented GenAI application aided clinicians in tailoring personalized communication to patients according to whether patients were primarily using problem-focused or emotion-focused coping [88].

4. Discussion and conclusion

4.1. Discussion

As the first review of its kind to examine how emerging research on GenAI and healthcare communication map onto a well-established framework for patient-centered communication (PCC), findings from this scoping review suggest that GenAI involvement in healthcare communication has potential to aid some aspects of PCC and hinder others. It is important to consider these findings in the context of the existing influence of EHRs on PCC, a novel lens for summarizing the emerging uses of GenAI by patients and clinicians for healthcare communication, as summarized visually in our results from Aim 1. When EHRs began to proliferate, many authors depicted EHRs as the “third agent” in the clinical encounter [12,89], particularly in countries such as the United States with high EHR-patient-portal interoperability, where patient-accessible notes, results, and secure messaging with clinicians are often all stored on a single platform. One might argue that this triadic conceptualization is even more salient for describing GenAI’s emerging presence within clinical communication. While EHRs represented a significant advancement in information technology controlled by human actors, GenAI is often described as an active “collaborator” [90,91] or “virtual coworker” [92] in healthcare and other spheres, due to its simulation of human intelligence and expression. Here, we discuss the benefits, drawbacks, and unknowns of GenAI for PCC identified in our review, in light of the EHR-embedded context of modern healthcare communication.

4.1.1. Benefits of GenAI for PCC

4.1.1.1. Increasing clinician bandwidth for interactions with patients.

Related to PCC Function 1, “Fostering Healing Relationships,” GenAI shows promise for alleviating both documentation burden and practical barriers to clinician-patient interaction that were introduced by EHRs. Many studies in our review highlighted the benefits that clinicians perceived for the efficiency of their workflow when GenAI assisted with drafting clinical documentation and responses to patients’ secure messages. Particularly in programs that piloted GenAI-embedded ambient scribe technology for automated transcription of visit notes, clinicians and patients perceived improved ability among clinicians to engage with patients and hold eye contact during clinical encounters. In these ways, our review presents novel findings on how GenAI may help to reverse some of the drawbacks for “fostering healing relationships” introduced by EHRs. Such benefits for efficiency of documentation and increased bandwidth for patient interaction may be of particular value in medical specializations involving relatively long-term relationships with patients, such as primary care.

4.1.1.2. Empowering patients to prepare for and understand procedures

Another aspect of communication at which GenAI (specifically, ChatGPT) performed strongly in this review was providing information to patients about how to prepare for and understand outcomes of procedures (PCC Function 6, “Enabling Self-Management”). EHRs open up new channels for information about preparing for procedures (asking clinicians questions through secure messaging), and patients now have the option of treating GenAI as an alternative source of information that allows less direct reliance on clinicians. EHRs also increase access to medical test results, and GenAI can help patients understand them. Studies in this review suggest that GenAI can provide accurate and useful explanations of test results. Because patients sometimes experience anxiety when they view ambiguous test results on EHRs prior to hearing clinicians’ interpretations [93–96], GenAI’s strong performance in this area is particularly promising.

4.1.1.3. Enhancing empathy in care communication.

One of the strongest aspects of GenAI’s performance found by studies in this review was its empathetic communication when answering patient questions, in relation to PCC Function 3, “Responding to Emotions.” This is notable given that patients consider it important for clinicians’ tone in secure messages to be empathetic [97], yet their messages often lack supportive language [98]. While studies in this review found that blinded clinician and patient raters considered GenAI-drafted answers to patient questions more empathetic than clinician-drafted answers, one study also found no association between GenAI authorship and patient satisfaction. This raises the question of whether the “empathy” expressed by GenAI differs in some way from authentic human empathy, given the strong links between clinician expressions of empathy and patient satisfaction [99, 100]. While GenAI shows promise in empathetic communication, our review highlights a novel gap in the literature as to how this empathetic language might be harnessed to enhance patient satisfaction.

4.1.2. Drawbacks of GenAI for PCC

4.1.2.1. Output of variable accuracy, relevance, and consistency.

EHRs created new potential for data entry errors [12] and became platforms for highly redundant information that clinicians have difficulty sifting through efficiently [101]. While GenAI could offer solutions to these drawbacks, our results suggest that findings on accuracy, relevance, and consistency of GenAI output are still variable (PCC Function 2, “Exchanging Information” and Function 4, “Managing Uncertainty”). Measurement of GenAI output quality is a nascent area of research with few validated tools developed for standardized use and comparison across studies [54]. There is a need for further development of quality evaluation measures for GenAI, particularly because models can sound confident even when they are hallucinating information [102].

Several studies in this review showed that prompt engineering could improve the quality of information produced by GenAI. While some studies found that making instructions increasingly specific improved the accuracy of GenAI output, others found that GenAI produced worse recommendations on diagnoses and treatment plans when asked to reason based on the broader context of a patient’s case and not simply to recite straightforward facts. Thus, optimal prompt engineering might differ across communication tasks. Future research should strive to identify best practices for clinicians and patients instructing GenAI on particular tasks.

4.1.2.2. Recasting the clinician as “technical supervisor”.

Patients in one qualitative study in this review envisioned clinicians as “technical supervisors” responsible for monitoring the mistakes of GenAI [38]. Indeed, much of the current ethical commentary on AI in healthcare has highlighted a need to clarify who is accountable for AI-generated mistakes [103,104]. If some or much of that responsibility will lie with clinicians, this new role for clinicians may have implications not only for their workload but also for the overall dynamic between clinicians and patients. If clinicians are “technical supervisors,” this might enable them to focus on their knowledge-based expertise, but perhaps at the cost of the more humanistic aspects of care that they can provide. Future research should monitor whether such a shift occurs and evaluate the effects of this trade-off between functions of PCC such as Function 1, “Fostering Healing Relationships” and Function 2, “Exchanging Information.” Given differences in the nature of clinician-patient interactions across different medical specialties (e.g., primary care often involving more holistic care, versus care in various specialties often already being highly boundaried, technical, and problem-focused), future research should also examine how the added role of “technical supervisor” influences the ability of clinicians to provide patient-centered care.

4.1.2.3. Issues of equity.

Many research teams, including one in this review [35], have alluded to the old computer science adage “garbage-in, garbage-out” when describing risks of GenAI for equity; essentially, if vast sources of data on which GenAI are trained already contain myriad racial, gender, disability, and other biases, then users can only expect GenAI output to replicate those biases. While many prior studies have highlighted this issue, our review makes the novel case that this current downside of GenAI threatens the respect for patients that is inherent to PCC Function 1, “Fostering Healing Relationships.” Curation of training data should aim to curtail amplification of biases and inequities. Further, studies in this review showed that unless GenAI was specifically instructed to translate material to lower reading levels, LLMs typically produced information written at reading levels too high for the average American to understand. National surveys in the U.S. have found that older adults, racial and ethnic minorities, and individuals with lower income or education levels are at particularly high risk of experiencing low health literacy [105,106]. It will be important for GenAI development and prompt engineering efforts to focus on readability of output so that patients with varying levels of health literacy can benefit from engagement with GenAI.

4.1.3. An agenda for further research on the influence of GenAI on PCC

In addition to the research gaps mentioned earlier, findings from our review also highlighted several issues relevant to PCC that have not yet been thoroughly explored in early research on GenAI in clinical communication, including GenAI’s influence on patient trust toward clinicians, questions around GenAI and patient privacy, effects of GenAI on patient adherence to treatment and coping with illness, and effects of GenAI on involvement of caregivers or family members in care decision-making. We conclude with an agenda (Table 3) for future research on GenAI and healthcare communication. As the first of its kind to be organized by the six functions of PCC and informed by findings from a systematic search on the currently known benefits and drawbacks regarding the influence of GenAI on PCC, the purpose of this research agenda is two-fold: to highlight novel gaps in research on GenAI and PCC, and to provide direction on how future research in this area can anchor its design to a well-established framework for PCC.

Table 3.

Agenda for Further Research on the Influence of GenAI on PCC.

PCC Function Areas for Further Research Potential Research Questions
Fostering Healing Relationships Patient trust in healthcare
  • How might patient trust toward both clinicians’ and their own engagement with GenAI change as GenAI becomes more advanced?

  • What demographic and contextual factors correlate with patients’ tendency to trust GenAI more or less for healthcare communication tasks?

  • If clinicians increasingly take on the role of “technical supervisors” who are responsible for checking GenAI output for errors and maintaining the clarity of information that is communicated to patients, how might this affect the clinician-patient relationship?

Healthcare equity
  • What strategies and safeguards for training GenAI for healthcare communication tasks might help to prevent replication of existing inequities in healthcare?

Patient privacy
  • As clinician engagement with GenAI increases, how might new concerns among patients about privacy affect patients’ willingness to share information with clinicians?

  • How can clinicians and other healthcare staff be trained to minimize risks to patient privacy in the ways in which they engage with GenAI?

Exchanging Information Equity in information access
  • How will differences in health and digital literacy affect different patients’ engagement with GenAI in healthcare communication? In light of potential differences, how can clinicians or healthcare systems facilitate equitable engagement with GenAI?

Evaluation and improvement of GenAI output accuracy and readability
  • How can validated tools be designed and optimized for evaluating the quality of healthcare information produced by GenAI?

  • How might prompts be most effectively engineered for different healthcare communication tasks involving GenAI, to obtain more accurate and readable output?

Responding to Emotions Patient preferences regarding empathy
  • Do patients prefer GenAI-produced healthcare information that has more or fewer statements that have the sole function of expressing empathy?

  • Do patients place differing levels of importance on empathetic language in healthcare communication depending on whether they are told that the communication was authored by a human clinician versus a GenAI chatbot?

Managing Uncertainty Evaluation and improvement of GenAI output relevance and comprehensiveness
  • How might prompts be most effectively engineered for different healthcare communication tasks involving GenAI, to obtain more relevant and comprehensive output?

Making Decisions Caregiver or family member involvement
  • What effects might different types of clinician or patient engagement with GenAI have on caregiver or family members’ collaborative involvement in healthcare?

Alignment with patients’ values and preferences
  • How might involvement of GenAI in healthcare communication influence the degree to which care decisions align with patients’ values and preferences?

Enabling Self-Management Patient adherence and collaborative involvement
  • What effects might different types of clinician or patient engagement with GenAI have on patients’ adherence to treatments and degree of collaborative involvement in healthcare?

  • What effects might different types of clinician or patient engagement with GenAI have on patients’ ability to cope effectively with illnesses?

4.1.4. Limitations

This scoping review had several limitations. Due to the nature of our research question, the studies included examined highly heterogenous GenAI models and areas of healthcare, which limits comparability across studies. Many studies did not clearly report important information about GenAI models examined, such as model version or, for GenAI chatbots, whether each prompt for a task was entered into a single, continuous chat window or a new chat window. Most studies also had relatively small sample sizes and used cross-sectional methods, further limiting the generalizability of findings and the ability to make conclusions about effects. Further, most were cross-sectional survey and observational studies. Finally, the use of a scoping review methodology for this study had limitations in comparison to systematic review methodology, including a lack of assessment of study quality and risk of bias in each included study. Future research should expand investigation of effects through experimental methods, and future systematic reviews should assess trends in study quality.

4.2. Conclusion

GenAI is proliferating in healthcare communication. Findings from our scoping review suggest that GenAI introduces both benefits and drawbacks for the patient-centeredness of healthcare communication. On the one hand, GenAI may help to increase clinician bandwidth to foster healing relationships with patients, particularly in an era when clinicians are facing a heavy burden of EHR documentation tasks with which GenAI can assist. On the other hand, patients may expect clinicians to function as “technical supervisors” of GenAI, potentially creating more relational distance between clinicians and patients. While GenAI shows early promise for enabling patients to self-manage their health, the quality of information that GenAI produces varies greatly between GenAI models, communication tasks, and phrasing of prompts. And while GenAI excels in empathetic dialogue one-on-one, growing engagement with GenAI also raises larger questions about healthcare equity. Future research should promote opportunities to secure the benefits of GenAI for PCC and safeguard against its drawbacks.

4.3. Practice implications

GenAI is emerging as a digital health technology with potential to reshape clinical communication into a triadic patient-clinician-AI relationship. Our findings suggest potential for this shift in clinical communication to offer clinicians more efficient methods of documentation while empowering patients with new methods of obtaining health information. However, findings from our review on the variable quality of AI-generated health information, as well as unresolved privacy and equity related issues around GenAI use, suggest a need for clinicians and patients to use GenAI with caution. These findings also highlight the need for policy or regulations around GenAI use to specify ways in which clinicians are accountable for checking information generated by AI that is used by clinicians to shape communication or treatment decisions. This emerging role of the clinician as “technical supervisor” to AI may yield a need for future training of clinicians not only in the logistics of how to use GenAI technologies but also in communication skills for patient care involving GenAI. This may include training for clinicians on how to disclose clearly their own use of GenAI as a matter of obtaining informed consent from patients, as well as training in how to facilitate discussion with patients about patients’ own use of GenAI and its influence on their understanding of their illness and its care. Findings from this review also highlight current gaps in research and provide directions for further research to ensure safe and patient-centered implementation of GenAI in healthcare.

Acknowledgements

The authors would like to thank Dr. Patrick Hill and Dr. Renee Thompson for their feedback on an early draft of this manuscript.

Funding

Jessica Hahne was supported by National Institute of Aging Grant T32 AG000030-47.

Appendix A

Detailed Search Strategy

Search Terms Entered into PubMed on May 12, 2024

(“natural language processing”[Mesh] OR “generative artificial intelligence” OR “large language model*” OR “ChatGPT*” OR “OpenAI” OR “Bard*” OR “Claude*” OR “LLaMa*” OR “generative adversarial network*” OR “variational autoencoder*” OR “vision-language model*”) AND (“Professional-Patient Relations”[Mesh] OR “Patient-Centered Care”[Mesh] OR “Patient communication” OR “Patient-centered communication” OR “Patient centered communication” OR “Patient-provider communication” OR “Patient provider communication” OR “Doctor-patient communication” OR “Doctor patient communication” OR “Patient Portals”[Mesh] OR “patient portal messag*” OR “Electronic Health Records”[Mesh])

Results for initial screening: 1914

Search Terms Entered into PsycInfo on May 12, 2024

DE “Electronic Health Records” OR “patient portal*” OR “patient portal messag*” OR DE “Patient Centered Care” OR DE “Communication” OR “professional-patient relations” OR “patient communication” OR “Patient-centered communication” OR “Patient centered communication” OR “Patient-provider communication” OR “Patient provider communication” OR “Doctor-patient communication” OR “Doctor patient communication”

AND DE “generative artificial intelligence” OR “natural language processing” OR “large language model*” OR “ChatGPT*” OR “OpenAI” OR “Bard*” OR “Claude*” OR “LLaMa*” OR “generative adversarial network*” OR “variational autoencoder*” OR “vision-language model*”

Results for initial screening: 382

Appendix B

Key Data Charted from Included Studies

Study Authors Country of Origin Objective Study Design Type of GenAI Examined GenAI Roles in Communication Relevant Functions of PCC Positive or Neutral Results for PCC Negative Results for PCC
Aharon et al., 2022 [88] Israel To evaluate whether use of an artificial intelligence program to aid personalized communication between healthcare providers and patients increases patient adherence to a cardiac rehabilitation (CR) program Prospective cohort “Well-Beat” system developed by the authors, a “personalized interaction aid […] based on a real-time Machine Learning algorithm” Co-writing educational materials, follow-up instructions, or responses to patient secure messages Responding to Emotions; Enabling Self-Management Patients in the Well-Beat program – which involved a GenAI platform that tailored communication to patients according to whether they were determined by the platform to be using problem-focused or emotion-focused coping – were significantly more likely to continue participating in the CR program at 3 months (76 %) compared to the historical control group (24 %, log-rank p < 0.001). N/A
Almagazzachi et al., 2024 [70] United States To examine the accuracy and reproducibility of AI-generated responses to frequently asked questions about hypertension Cross-sectional ChatGPT, version unspecified (OpenAI) Answering medical questions, as a substitute/adjunct to secure messaging or seeing clinician Exchanging Information; Managing Uncertainty ChatGPT responses were rated by a physician to be appropriate in 92.5 % of cases, inappropriate in 7.5 %, and had a reproducibility score of 93 %. N/A
Amin et al., 2024 [53] United States To assess the readability and word count of information about various medical conditions generated by four publicly available LLMs, over 4 months Repeated cross-sectional (quasi-longitudinal) ChatGPT-3.5 (OpenAI); ChatGPT-4.0 (OpenAI); Bard (Google); Gemini, formerly known as Bard (Google); Bing search assistant, set to precise (Microsoft) Answering medical questions, as a substitute/adjunct to secure messaging or seeing clinician Exchanging Information The average reading grade levels for ChatGPT-4.0 and Bing decreased from the first time point to the second time point four months later.
Word counts between output at the first and second time point increased for ChatGPT-3.5 while decreasing for Bing and ChatGPT-4.0.
The average reading grade levels for Bard (Gemini) increased from the first time point to the second time point four months later.
None of the models at any time point reach the target of eighth-grade reading level for patient facing material.
Ashraf et al., 2024 [107] United States To characterize online pharmacy recommendations by GenAI, including determining the prevalence of illegal online pharmacy recommendations Cross-sectional Search Generative Experience (SGE) using converse mode, which integrates Bard (Google); Bing’s chat feature (Microsoft) Answering medical questions, as a substitute/adjunct to secure messaging or seeing clinician Exchanging Information Bing Chat generated a significantly higher proportion (21/86, 24 %) of links leading to illegal online pharmacies selling prescription medications than Google SGE (6/92, 6%) (P = .001). 47.33 % (124/262) of websites recommended by the GenAI examined belonged to active online pharmacies, with only 31.29 % (82/262) leading to legitimate ones. 19.04 % (24/126) of Bing Chat’s and 13.23 % (18/136) of Google SGE’s recommendations were links to illegal vendors (which are known to market substandard or falsified medicines), including for controlled substances.
Ayers et al., 2023 [65] United States To investigate the quality and empathy of responses by an artificial intelligence chatbot to patient questions from a public social media forum, in comparison to responses written by physicians Cross-sectional ChatGPT, version unspecified (OpenAI) Answering medical questions, as a substitute/adjunct to secure messaging or seeing clinician Exchanging Information; Responding to emotions Chatbot responses were rated by healthcare professionals as significantly higher quality and significantly more empathetic than physician responses. N/A
Ayoub, Ballout et al., 2023 [76] United States To analyze the ability of ChatGPT to triage, synthesize differential diagnoses, and generate treatment plans for clinical scenarios from cardiology, pulmonology, and neurology Cross-sectional ChatGPT, version unspecified (OpenAI) Augmenting clinician input on decision-making (e.g., diagnoses, treatment recommendations) Making Decisions ChatGPT’s responses to various medical scenarios were rated relatively high by physicians in terms of accuracy of differential diagnoses and the appropriateness of initial treatment plans. ChatGPT’ s responses to various medical scenarios were rated relatively low by physicians in terms of completeness of differential diagnoses.
Ayoub, Lee et al., 2023 [84] United States To assess the understandability, actionability, and accuracy of ChatGPT’s responses to questions about various medical conditions, compared to the top nonsponsored result generated by Google Search Cross-sectional ChatGPT, version unspecified (OpenAI) Answering medical questions, as a substitute/adjunct to secure messaging or seeing clinician Exchanging Information; Enabling Self-Management ChatGPT scored higher than Google Search (87 % vs 78 %, p = .012) for answers to patient education questions. The average PEMAT-P score (which evaluates the understandability and actionability of patient education materials) for medical advice was 68.2 % for ChatGPT and 89.4 % for Google Search (p < .001). PEMAT-P scores were significantly different by source (p < .001) but not by urgency of clinical situation (p = .613).
ChatGPT scored higher than Google Search when offering general medical knowledge, but it scored lower when providing medical recommendations.
Barak-Corren et al., 2024 [31] United States To evaluate the quality (time savings, effort reduction, physician attitudes, and potential risks and barriers) of ChatGPT to produce clinical documentation (supervisory attending notes, structured handoff summaries, and “letter to the family” summaries) in pediatric emergency medicine Cross-sectional ChatGPT-4.0 (OpenAI) Co-writing visit notes; Summarizing previous care records (“chart review”) Fostering Healing Relationships; Exchanging Information When ChatGPT generated supervisory notes, it yielded a 40 % reduction in clinician time and a 33 % decrease in clinician effort for intricate cases, with no significant effect for simpler notes. ChatGPT-generated structured handoff summaries and family letter summaries were rated highly by participating physicians, ranging from 7.0 to 9.0 out of 10.0. Most participants favored the inclusion of ChatGPT-generated summaries in clinical practice. Participating physicians had several critical reservations about incorporating GenAI into their clinical documentation workflow, including concerns about patient privacy being compromised by corporate owners of GenAI and concerns about ChatGPT replicating or amplifying existing issues of structural racism.
Baxter et al., 2024 [68] United States To evaluate ChatGPT-generated responses to negative EHR patient messages, in comparison to actual responses sent by healthcare team members Cross-sectional ChatGPT-3.5 (OpenAI) Co-writing educational materials, follow-up instructions, or responses to patient secure messages Fostering Healing Relationships; Exchanging Information; Responding to Emotions; Managing Uncertainty In many cases, ChatGPT provided reasonable starting drafts for responses to patient messages, particularly when provided with a prompt that gave context for the clinical situation and the need to respond from the perspective of a clinician. Chat-GPT-drafted responses often required further editing. Issues included sometimes not speaking from the perspective of a clinician, using overly generic language, inappropriate escalation (eg, instructing the patient to file complaints to the medical board), and inconsistent recommendations for in-person follow-up visits.
Blease et al., 2024 [35] United States To explore psychiatrists’ experiences and opinions regarding generative AI Cross-sectional Various GenAI identified by respondents, including ChatGPT-3.5 (OpenAI); ChatGPT-4.0 (OpenAI); Bard (Google); Bing (Microsoft); Claude-2 (Anthropic) Co-writing visit notes; Summarizing previous care records (“chart review”); Answering medical questions, as a substitute or adjunct to secure messaging with clinician Fostering Healing Relationships; Exchanging Information Participating psychiatrists expected administrative tasks to be a major benefit of GenAI: 70 % somewhat agreed or agreed that “documentation will be / is more efficient.” 75 % of participating psychiatrists somewhat agreed or agreed that “the majority of their patients will consult these tools before first seeing a doctor.” Almost all participating clinicians somewhat agreed or fully agreed that clinicians need more support and training in understanding GenAI.
Cabral et al., 2024 [79] United States To compare an LLM’ s clinical reasoning abilities against human performance using a standard scale developed for physicians (R-IDEA) Cross-sectional ChatGPT-4.0 (OpenAI) Augmenting clinician input on decision-making (e.g., diagnoses, treatment recommendations) Making Decisions Median (IQR) R-IDEA clinical reasoning scores were 10 (9–10) for the chatbot, 9 (6–10) for attending physicians, and 8 (4–9) for resident physicians. In logistic regression analysis, the chatbot had a significantly higher estimated probability of receiving high R-IDEA scores than attendings and residents. Instances of incorrect clinical reasoning were significantly more frequent for the chatbot compared to resident physicians (2.8 %; P = .04) but not compared to attending physicians (12.5 %; P = .89).
Carnino et al., 2024 [63] United States To evaluate the accuracy, comprehensiveness, and safety of ChatGPT’s recommendations in response to real-world otolaryngology patient questions Cross-sectional ChatGPT-3.5 (OpenAI) Answering medical questions, as a substitute/adjunct to secure messaging or seeing clinician Exchanging Information; Responding to Emotions; Managing Uncertainty; Enabling Self-Management According to otolaryngologist raters, the average bedside manner/empathy score for ChatGPT responses was 4.3/5.0. The average accuracy score for ChatGPT responses was 3.76/5, and average comprehensiveness was 3.59/5. A correlation was observed between higher question difficulty and lower comprehensiveness in responses.
Five out of 15 responses were flagged by raters as potentially dangerous, due to providing incorrect information.
Chen et al., 2024 [72] China To develop and validate a bilingual (English-Chinese) indocyanine green angiography (ICGA) report generation and question-answering system Cross-sectional ICGA-GPT (a model developed by the authors specifically for ICGA images, combining a multimodality transformer architecture and an LLM) Simplifying reports on test or scan results; Co-writing educational materials, follow-up instructions, or responses to patient secure messages; Explaining jargon from notes or test results viewed in portal Exchanging Information; Managing Uncertainty In response to an interactive question-and-answer scenario that involved 100 generated answers, the ophthalmologists provided scores of 4.2, 4.2 and 4.1 out of 5.0 (kappa = 0.779). Assessing the quality of 100 reports on 50 images, three ophthalmologists achieved substantial agreement, yielding scores from 3.20 to 3.55.
Chervenak et al., 2023 [43] United States To compare the responses of ChatGPT to three reputable sources, in response to fertility-related clinical prompts Cross-sectional ChatGPT-3.5 (OpenAI) Answering medical questions, as a substitute/adjunct to secure messaging or seeing clinician Exchanging Information; Managing Uncertainty When compared to the CDC’s answers to 17 infertility FAQ’s, ChatGPT produced responses of comparable length, factual content, sentiment polarity, and subjectivity. ChatGPT also achieved high scores on validated fertility knowledge surveys, compared to published population data. ChatGPT also correctly reproduced missing facts for all 7 summary statements from the “Optimizing Natural Fertility” Practice Committee document. 9 (6.12 %) out of of 147 statements prouced by ChatGPT were categorized as incorrect, and only 1 (0.68 %) statement cited a reference.
Chervonski et al., 2024 [54] United States To assess the quality of generative AI responses to common patient questions about vascular surgery disease Cross-sectional ChatGPT-3.5 (OpenAI); Bard (Google) Answering medical questions, as a substitute/adjunct to secure messaging or seeing clinician Exchanging Information; Managing Uncertainty; Enabling Self-Management On average, ChatGPT responses were rated more accurate than Bard responses (3.08 vs 2.82, p < .01) and rated more complete than Bard responses (2.98 vs 2.62, p < .01). Most ChatGPT responses (75.0%, n = 18) and almost half of Bard responses (45.8 %, n = 11) were unanimously rated appropriate. Almost one-third of Bard responses (29.2 %, n = 7) were rated inappropriate by at least two reviewers. Two Bard responses (8.4 %) were considered inappropriate by the majority of reviewers. Mean Flesch Reading Ease, Flesch-Kincaid Grade Level, and Gunning Fog Index of ChatGPT responses were 29.4, 14.5, and 17.7, respectively (responses were readable with a post-secondary education). Bard’s mean readability scores were 58.9, 8.2, and 11.0, respectively (responses were readable with a high school education).
Chuang et al., 2024 [108] United States To test the performance and performance variability of SPeC, a “zero-shot and model-agnostic framework with a soft prompt encoder,” in prompting three different LLMs to summarize clinical notes and radiology reports Cross-sectional ChatGPT-3.5, for generating prompts (OpenAI); Flan T-5 (Google); BART (Meta); Pegasus-xsum (Google) Simplifying reports on test or scan results; Summarizing previous care records (“chart review”) Fostering Healing Relationships; Exchanging Information Using soft prompts with discrete prompts, the SPeC approach demonstrated better performance and lower variance in summarization of clinical notes and radiology reports by all three LLMs tested, as compared to summarization by the base LLMs using ChatGPT-generated prompts. Summarization of clinical notes and radiology reports by base LLMs using ChatGPT-generated prompts demonstrated notable variance in summarization performance, including in information that would be relevant to clinical decision making.
The authors note in the limitations to their findings that some medical reports may reveal more private information about patients than the radiology reports that were tested in this study, which may present challenges to ensuring patient privacy and complying with data security regulations in real-life applications of this technology.
Denecke et al., 2024 [37] Switzerland To explore the perspectives of researchers working in natural language processing in healthcare regarding the benefits, shortcomings, and risks of using transformer models in healthcare Cross-sectional Transformer models (unspecified) (Open-ended / none specified) Fostering Healing Relationships; Exchanging Information; Responding to Emotions; Making Decisions Potential benefits of transformer models in healthcare raised by participants included the following: Transformer models could improve clinical documentation quality and reduce documentation burden for healthcare professionals; transformer models can reduce errors in healthcare communication and tailor information to specific recipients; transformer models could improve evidence-based decision making and accuracy of diagnoses. Potential challenges of transformer models in healthcare raised by participants included the following: The quality of transformer models is affected by biases in training data, such as race and gender bias; there are risks of incorrect information provided by transformer models coupled with a lack of explainability and interpretability; factors such as low literacy, accessibility issues, and socio-economic status could create barriers to patient use of transformer models; transformer models could reduce clinician-patient interaction and de-humanize care; over-reliance on transformer models by both patients and clinicians could undermine patients’ self-management and decision-making input; transformer models increase privacy and security risks.
Dosso et al., 2024 [55] Canada To evaluate the length, readability, and quality of ChatGPT responses to questions about dementia, compared to online patient education material from three North American organizations Cross-sectional ChatGPT-3.5 (OpenAI) Answering medical questions, as a substitute/adjunct to secure messaging or seeing clinician Exchanging Information All organization and ChatGPT responses received full points for (minimizing) Conflict of Interest. Most organization and ChatGPT responses received the one point possible for Complementarity (of the patient-physician relationship). When analyzed as a pool, the set of responses written by organizations received a QUEST score of 21 out of 28 points. ChatGPT responses to the same questions received a score of 16. All organizations but only 1/3 of ChatGPT responses received full points for tone (balanced/cautious support of claims). Organizations but not ChatGPT referenced identifiable research. Responses from organizations had lower Flesch-Kincaid Grade Levels (better readability) than responses from ChatGPT.
Floyd et al., 2024 [73] United States To measure the accuracy and comprehensiveness of ChatGPT attempts at radiation oncology communication tasks, including answering common patient questions, summarizing clinical research studies, and providing literature reviews with standard-of-care references. Cross-sectional ChatGPT-3.5 (OpenAI) (main study); ChatGPT-4.0 (OpenAI) (post-hoc analyses) Augmenting clinician input on decision-making (e.g., diagnoses, treatment recommendations); Co-writing educational materials, follow-up instructions, or responses to patient secure messages Exchanging Information; Managing Uncertainty; Making Decisions In posthoc analyses comparing results from ChatGPT-4.0 to those produced by ChatGPT-3.5, while ChatGPT-4.0 performed similarly to ChatGPT-3.5 in terms of responding to patient-centered questions and summarizing landmark studies related to standard-of-care for prostate cancer, ChatGPT-4.0 performed much better than ChatGPT3.5 at summarizing standard-of-care practice studies radiation oncology (100 % of studies identified by ChatGPT-4.0 were real, versus only 83 % for ChatGPT-3.5). ChatGPT-4.0 correctly summarized 88 % of those real studies, compared with 30 % correctly summarized by ChatGPT-3.5. ChatGPT-3.5 frequently generated incomplete or inaccurate responses. Only 39.7 % of responses to patient questions were rated correct and comprehensive. When summarizing studies, only 35.0 % of ChatGPT-3.5’s responses were accurate and comprehensive (improving to 43.3 % when it was provided full text of studies).
Garcia et al., 2024 [33] United States To evaluate the implementation of an EHR-integrated LLM created to draft clinician replies to patient portal messages Prospective cohort A HIPAA-compliant, EHR-integrated version of ChatGPT-4.0 (OpenAI), created to generate draft replies to patient portal messages for clinicians Co-writing educational materials, follow-up instructions, or responses to patient secure messages Fostering Healing Relationships; Exchanging Information; Responding to Emotions Improvements in task load and emotional exhaustion scores among clinicians suggested that generated draft replies may help reduce cognitive burden and burnout. Similarly, clinicians expressed high expectations about utility and quality of message drafts that were either met or exceeded by AI-drafted messages. No changes in overall reply time, read time, or write time were found when implementing the AI pilot, compared to clinician prepilot times.
Clincians highlighted the need for improvements in tone, brevity, and personalization; some clinicians preferred longer, more empathetic responses but others preferred shorter, more formal responses.
50 % of draft messages received negative feedback from clinicians in terms of the relevance of message content.
Ge et al., 2024 [46] United States To assess the performance of an LLM in comparison to human manual chart review, in extracting 8 data elements from hepatocellular carcinoma CT or MRI imaging reports Cross-sectional Versa GPT-4 (a PHI-compliant version of Microsoft Azure OpenAI GPT-4 LLM implemented at the University of California, San Francisco) Summarizing previous care records (“chart review”) Exchanging Information Using human manual chart review as the standard, ChatGPT’s overall accuracy in extracting data was 0.934, with accuracies for specific elements of data varying between 0.886 (sum of tumor diameters) to 0.989 (extrahepatic metastases). N/A
Ghanem et al., 2024 [56] United States To evaluate the content and quality of medical information on acute appendicitis generated by ChatGPT 3.5, ChatGPT-4.0, Bard, and Claude-2 Cross-sectional ChatGPT-3.5 (OpenAI), ChatGPT-4.0 (OpenAI), Bard (Google), and Claude-2 (Anthropic) Answering medical questions, as a substitute/adjunct to secure messaging or seeing clinician Exchanging Information; Managing Uncertainty AI-generated medical information from all four chatbots provided adequate explanation of the risks of not treating appendicitis, as well as accurate content regarding etiology, symptoms, and treatment with no major factual errors. Bard was the only AI platform that listed verifiable sources, while Claude-2 provided fabricated sources.
Average response reading levels were far more difficult than recommended levels for the public.
Goldstein & Shahar, 2016 [45] Israel To evaluate the completeness, correctness, and overall quality of patient discharge letters generated by the CliniText system compared to those composed by physicians, for cardiothoracic post-surgery care and diabetes care Cross-sectional Clinitext, an automatic summarization system that uses natural language generation Co-writing visit notes; Summarizing previous care records (“chart review”) Exchanging Information In clinician-produced letters, there were at least two important items missed per letter that were included in letters produced by CliniText.
Three clinicians answered questions about details from the discharge letter 40 % faster and answered four out of the five questions equally well or significantly better, when using CliniText-generated letters, compared to using clinician-composed letters.
Clinicians’ letters got significantly better ratings in three out of four quality measures; however, variance in quality was much higher for clinician-composed letters.
Goldstein et al., 2017 [44] Israel To examine the feasibility and potential clincal benefits of using the Clinitext system to generate summaries of longitudinal clinical records in an ICU context Cross-sectional Clinitext, an automatic summarization system that uses natural language generation Co-writing visit notes; Summarizing previous care records (“chart review”) Exchanging Information In 13/31 (42 %) letters, the number of important items missed in Clinitext-generated letters was less than or equal to the number of important items missed in clinician-composed letters. In each clinician-composed letters, at least two important items that were mentioned by Clinitext were missed.
Clinicians answered questions about details in the letters 40 % faster when using Clinitext-generated letters compared to clinician-composed letters. In four out of five questions the clinicians’ correctness was equal or better when using Clinitext-generated letters compared to using clinician-composed letters.
Clinicians’ letters got significantly better ratings in three out of four quality measures; however, variance in quality was much higher for clinician-composed letters.
Gordon et al., 2024 [71] United States To assess ChatGPT’s accuracy, relevance, and readability in answering patients’ common imaging-related questions and examine the effect of a simple prompt Cross-sectional ChatGPT-3.5 (OpenAI) Explaining jargon from notes or test results viewed in portal; Answering medical questions, as a substitute/adjunct to secure messaging or seeing clinician Exchanging Information; Managing Uncertainty; Enabling Self-Management Questions about safety, a radiology report, the related procedure, preparation before imaging, meaning of terms, and medical staff interactions were posed to ChatGPT with and without a short prompt. According to radiologist ratings, 83 % of unprompted responses were accurate (218 of 264), and the proportion did not change significantly for prompted responses (87 % [229 of 264]; P = .2). When prompts were given, the consistency of responses increased from 72 % (63 of 88) to 86 % (76 of 88) (P = .02). Nearly all responses (99 % [261 of 264]) were at least partly relevant for both prompted and unprompted questions. Fewer unprompted responses were rated fully relevant at 67 % (176 of 264); however, this increased significantly to 80 % relevant when prompts were given (210 of 264; P = .001). Mean Flesch Kincaid Grade Level was 13.6, and was unchanged with prompts (13.0, P = .2). No responses reached the readability level recommended for patient-facing materials.
Guevara et al., 2024 [40] United States To identify optimal methods for using LLMs to extract social determinants of health (SDoH) from narrative text in the EHR, including employment, housing, transportation, parental status, relationship status, and social support Cross-sectional Flan-T5 models, including Flan-T5 base, large, XL, and XXL (Google); ChatGPT-family models, including GPT-turbo-0613 and GPT4–0613 (OpenAI) Summarizing previous care records (“chart review”) Fostering Healing Relationships; Exchanging Information Fine-tuned Flan-T5 XL performed better than any other models for identifying any SDoH mentions. Flan-T5 XXL performed better than any other models for identifying adverse SDoH mentions. Adding LLM-generated synthetic data to training improved the performance of smaller Flan-T5 models. Models that were fine-tuned by the authors outperformed zero- and few-shot performance of ChatGPT-family models, with the exception of GPT-4.0 with 10-shot prompting for adverse SDoH. The authors’ fine-tuned models identified 93.8 % of patients with adverse SDoH, compared to only 2.0 % captured by ICD-10 codes. ChatGPT-family models were significantly more likely than models that were fine-tuned by the authors to change their classification for any SDoH when a female gender was injected into the text. Although not significantly different (perhaps due to small proportion within the overall sample), both the fine-tuned models and ChatGPT-family models were likely to change their classifications for any SDoH and adverse SDoH mentions when Hispanic and Black descriptors were injected into the text.
Hanai et al., 2024 [109] Japan To explore the characteristics of responses by generative AI to hypothetical cancer patient questions about sexual health Cross-sectional ChatGPT-3.5 (OpenAI) Answering medical questions, as a substitute/adjunct to secure messaging or seeing clinician Fostering Healing Relationships; Exchanging Information ChatGPT provided helpful health information on sensitive topics such as sexual health. ChatGPT responses were biased toward recommending non-pharmacological interventions.
Haver et al., 2023 [60] United States To evaluate the clinical appropriateness of ChatGPT’s responses to common questions about breast cancer prevention and screening Cross-sectional ChatGPT, version unspecified (OpenAI) Explaining jargon from notes or test results viewed in portal; Co-writing educational materials, follow-up instructions, or responses to patient secure messages; Answering medical questions, as a substitute/adjunct to secure messaging or seeing clinician Exchanging Information; Enabling Self-Management ChatGPT-generated responses were judged by radiologist raters to be appropriate for 22 of 25 (88 %) questions in the imagined contexts of both patient-facing educational material and chatbot responses to patient questions. N/A
Haver et al., 2024 [61] United States To evaluate the ability of ChatGPT to simplify answers to common questions about breast cancer prevention and screening (previously generated by ChatGPT in Haver et al., 2023) to a sixth-grade reading level Cross-sectional ChatGPT-3.5 (OpenAI) Co-writing educational materials, follow-up instructions, or responses to patient secure messages Exchanging Information; Enabling Self-Management When prompted to simplify to a sixth-grade reading level a set of answers to common questions about breast cancer prevention and screening that had previously been generated by ChatGPT, ChatGPT improved the mean reading ease of responses from an average level of “difficult to read” to an average level of “easy to read. It also decreased word count from an average of 193 words in the original responses to an average of 173 words in the simplified responses. 92 % of simplified responses were considered by a radiologist to be clinically appropriate (versus 88 % of the original responses). ChatGPT’s original and simplified responses to common questions about breast cancer prevention and screening were both written at a reading level higher than the average reader from the general public, despite the fact that ChatGPT demonstrated an ability to lower the reading level of responses when prompted to simplify them (original average reading level: Grade-level 13, versus simplified average reading level: Grade-level 8.9).
He, Bhasuran et al., 2024 [87] United States To assess the ability of various LLMs to generate relevant, accurate, and helpful responses to patient questions about lab tests, and to identify ways to mitigate potential issues using augmentation approaches. Cross-sectional ChatGPT-4.0 (OpenAI); LLaMA-family models: LLaMA-2, MedAlpaca, orca-mini (Meta) Explaining jargon from notes or test results viewed in portal; Answering medical questions, as a substitute/adjunct to secure messaging or seeing clinician Exchanging Information; Managing Uncertainty; Enabling Self-Management When GPT-4.0 output (which scored the highest) was used as the reference answer, responses from GPT-3.5 were the most similar, followed by LLaMA-2, ORCA_mini, and MedAlpaca.
Yahoo data answers from humans scored the lowest. The authors identified several ways to improve the quality of LLM responses on both the prompt side and the response side.
LLM responses occasionally failed at interpretation within the patient’s medical context, made incorrect statements, or lacked references.
He, Zhang et al., 2024 [66] China To assess the ability of two LLM chatbots to address Chinese-language inquiries about Autism Spectrum Disorder Cross-sectional ChatGPT-4.0 (OpenAI); ERNIE Bot, version 2.2.3 (Baidu, Inc.) Answering medical questions, as a substitute/adjunct to secure messaging or seeing clinician Exchanging Information; Responding to Emotions; Managing Uncertainty Responses by physicians were rated more relevant than those by ChatGPT or ERNIE Bot (mean of 3.75, versus 3.69 and 3.41, respectively). Responses by physicians and ChatGPT were rated more accurate than those by ERNIE Bot (3.66 and 3.73, versus 3.52, respectively). ChatGPT’s empathy scores significantly outperformed physicians and ERNIE Bot (3.64, versus 3.13 and 3.11 respectively). Fewer than half of physician assessors (46.86 %) preferred responses drafted by physicians to responses drafted by either LLM (with 34.87 % of assessors favoring ChatGPT and 18.27 % of assessors favoring ERNIE Bot). Physician responses received significantly higher usefulness ratings than ChatGPT or ERNIE Bot responses (3.54, versus 3.40 and 3.05, respectively).
Hristidis et al., 2023 [57] United States To compare the quality of responses from ChatGPT and Google (top result) for a set of questions from people living with dementia or other cognitive decline Cross-sectional ChatGPT-3.0 (OpenAI) Answering medical questions, as a substitute/adjunct to secure messaging or seeing clinician Exchanging Information; Managing Uncertainty Per evaluation by domain experts, ChatGPT’s answers received higher ratings on objectivity and relevance than Google’s results. Results from Google were more current than answers from ChatGPT. Readability was low for both ChatGPT and Google results, especially for ChatGPT (mean grade level 12.17), compared to Google (mean grade level 9.86). ChatGPT rarely included sources for its responses. The authors highlight concerns in their Discussion section about ChatGPT being trained on existing content on the internet and therefore posing the risk of providing users with inaccurate or biased information. They also discuss user privacy issues, such as unknowns around how user interactions with technologies such as ChatGPT may be tracked.
Hu et al., 2024 [51] China To evaluate the ability of ChatGPT (in comparison to a formerly developed information extraction system and a rule-based model) to extract information from radiological reports on CT scans for lung cancer patients, both without prior medical information (zero-shot performance) and with prompts engineered to include prior medical information Cross-sectional ChatGPT (various versions from from July 24, 2023 to November 2, 2023) Simplifying reports on test or scan results; Summarizing previous care records (“chart review”) Exchanging Information In the zero-shot learning phase: (1) ChatGPT performed comparatively to MTQA (a formerly developed information extraction sysem) for questions about tumor long and short diameters, tumor lobulation, mediastinal lymph node status, and pleural invasion or indentation; (2) ChatGPT performed comparatively to the rule-based model for questions about tumor location, tumor long and short diameter, mediastinal lymph node status, and pleural invasion or indentation.
When prompted with prior medical knowledge, ChatGPT achieved significant improvements compared to base prompts on questions about tumor spiculation, lobulation, and pleural invasion or indentation.
In the zero-shot learning phase: (1) ChatGPT performed worse than MTQA for questions about tumor density and spiculation; (2) ChatGPT performed worse than the rule-based model for questions about tumor density, tumor spiculation, tumor lobulation, and hilar lymph node status. Even when prompted with prior medical knowledge, ChatGPT did not achieve significant improvements compared to base prompts on questions about tumor density or lymph node status.
In their Discussion section, the authors highlight concerns about how to protect patient privacy if employing ChatGPT in real world hospital settings, due to the need to input real medical data into ChatGPT, which may increase the risk of medical data leakage.
Huang et al., 2023 [47] United States To examine the accuracy and quality of AI-generated interpretations of chest radiographs in the emergency department setting Cross-sectional According to the authors, “a transformer-based encoder-decoder model that takes chest radiograph images as input and generates radiology report text as output” Simplifying reports on test or scan results Exchanging Information The GenAI model produced reports of comparable clinical accuracy and quality to radiologist-produced reports while achieving higher quality ratings than teleradiologist reports. N/A
Huo et al., 2024 [82] Canada To assess the performance of LLM chatbots in providing recommendations for surgical management of gastroesophageal reflux disease (GERD) Cross-sectional ChatGPT-3.5 (OpenAI); ChatGPT-4.0 (OpenAI); Copilot (Microsoft); Bard (Google); Perplexity (Perplexity AI) Augmenting clinician input on decision-making (e.g., diagnoses, treatment recommendations) Making Decisions N/A For adult patients with GERD: (1) Surgeons were given accurate recommendations for 5/7 (71.4 %) of ChatGPT-4.0 responses, 3/7 (42.9 %) Copilot responses, 6/7 (85.7 %) Google Bard responses, and 3/7 (42.9 %) Perplexity responses; (2) Patients were given accurate recommendations in 3/5 (60.0 %) ChatGPT-4.0 responses, 2/5 (40.0 %) Copilot responses, 4/5 (80.0 %) Google Bard responses, and 1/5 (20.0 %) Perplexity responses.
For pediatric patients with GERD: (1) Surgeons were given accurate recommendations in 2/3 (66.7 %) ChatGPT-4.0 responses, 3/3 (100.0 %) Copilot responses, 3/3 (100.0 %) Google Bard responses, and 2/3 (66.7 %) Perplexity responses; (2) Patients were given appropriate guidance in 2/2 (100.0 %) ChatGPT-4.0 responses, 2/2 (100.0 %) Copilot responses, 1/2 (50.0 %) Google Bard responses, and 1/2 (50.0 %) Perplexity responses.
None of the chatbots provided correct guidance for all clinical questions based on SAGES guideline recommendations.
Iscoe et al., 2024 [110] United States To develop and compare the ability of two LLMs to identify UTI symptoms from unstructured emergency department notes Cross-sectional Two task-specific LLMs developed by the authors: SpaCy (a convolutional neural network-based model), and Clinical Longformer (a transformer-based model) Summarizing previous care records (“chart review”); Augmenting clinician input on decision-making (e.g., diagnoses, treatment recommendations) Exchanging Information; Making Decisions The overall F1 (harmonic mean of precision and recall) measure for note-level symptom identification, weighted by frequency of symptom in the data set, was 0.84 for the SpaCy model and 0.88 for the Longformer mode. These results indicated good performance at the task, on average, for both models. Model performance for identification of UTI symptoms was highly variable across individual symptoms, with the lowest F1 measures for both models being for urinary incontinence (.51 for the SpaCy model and 0 for the Longformer model)
Kanjee et al., 2023 [77] United States To assess the accuracy of ChatGPT-4.0 at generating differential diagnoses in response to New England Journal of Medicine clinicopathologic conferences Cross-sectional ChatGPT-4.0 (OpenAI) Augmenting clinician input on decision-making (e.g., diagnoses, treatment recommendations) Making Decisions ChatGPT-4 provided the correct diagnosis in its differential in 64 % of challenging cases and as its top diagnosis in 39 %. These finding compare favorably with existing differential diagnosis generators. N/A
Kernberg et al., 2024 [49] United States To examine the accuracy and quality of Subjective, Objective, Assessment, and Plan (SOAP) notes generated by ChatGPT-4.0, compared to established transcripts of History and Physical Examination, for simulated patient-provider encounters in various ambulatory specialties Cross-sectional ChatGPT-4.0 (OpenAI) Co-writing visit notes; Transcribing patient answers to medical interview during visit (e.g., illness history, care preferences) Fostering Healing Relationships; Exchanging Information ChatGPT-4.0 consistently generated SOAP-style notes. On average, there were 23.6 errors per clinical case, in the ChatGPT-4.0-generated SOAP-style notes. Errors of omission (86 %) were the most common, followed by errors of addition (10.5 %) and incorrect facts (3.2 %). When replicates of the same case were generated, there was significant variance between replicates, with only 52.9 % of data items reported correctly across all 3 replicates. Accuracy of data items varied across cases.
The authors also include the following disclaimer in their Discussion section: “ It is important to note that open AI platforms, such as ChatGPT-4, are not recommended for clinical use due to the many regulatory and privacy issues.”
Khene et al., 2024 [50] France To evaluate the adequacy of answers provided by personalized chatbot Uro_Chat, to questions about clinical management of urologic cancers, compared to answers generated by ChatGPT-3.5 and ChatGPT-4.0 Cross-sectional A personalized chatbot (Uro_Chat) created using Llamalndex (version 0.6.12) (Meta) to connect data from the European Association of Urology guidelines to GPT 3.5-Turbo (OpenAI); as compared to ChatGPT-3.5 (OpenAI) and ChatGPT-4.0 (OpenAI) Augmenting clinician input on decision-making (e.g., diagnoses, treatment recommendations) Exchanging Information Uro_Chat gave an adequate response to all 20 questions about urological advice, including 5 questions drawn from recent literature. ChatGPT-3.5 and ChatGPT-4.0 generally failed to provide accurate answers.
Kienzle et al., 2024 [85] Germany To assess the quality of ChatGPT’s responses to frequently asked questions preceding total knee arthroplasty, in order to assess its potential usefulness in providing preoperative patient education Cross-sectional ChatGPT-4.0 (OpenAI) Co-writing educational materials, follow-up instructions, or responses to patient secure messages; Answering medical questions, as a substitute/adjunct to secure messaging or seeing clinician Fostering Healing Relationships; Exchanging Information; Managing Uncertainty; Enabling Self-Management As judged by orthopedic surgeons, ChatGPT’s responses to questions about diagnosis and non-surgical treatments received an average DISCERN (quality) score of 3.7 out of 5, responses to surgery-related questions received an average score of 4.0, and responses to daily-life-related questions received an average score of 3.7. ChatGPT’s responses to complications-related questions received an average score of 3.3 out of 5.
Out of the 27 references that ChatGPT provided across 50 responses, 10 of those references (37%) appeared to be fabricated.
The authors include the following note of caution in their Discussion section: “privacy concerns may arise if ChatGPT collects personal and health-related information from patients.”
Kim et al., 2024 [75] South Korea To develop and validate a software that uses ChatGPT-3.5 to generate patient-friendly medical discharge summaries Cross-sectional ChatGPT-3.5 (OpenAI) Co-writing visit notes; Summarizing previous care records (“chart review”) Fostering Healing Relationships; Exchanging Information When using one-shot prompts, the mean overall score for outputs 4.11 out of 5 (Factuality: 4.20; Comprehensiveness: 4.08; Usability: 3.93; Reading ease: 4.25; Linguistic fluency: 4.11)
When using few-shot prompts, the mean overall score for outputs was 4.19 out of 5 (Factuality: 4.19; Comprehensiveness: 4.18; Usability: 3.97; Reading ease: 4.39; Linguistic fluency: 4.22)
When using zero-shot prompts, the mean overall score for outputs was was 3.73 (Factuality: 3.82; Comprehensiveness: 3.68; Usability: 3.36; Reading ease: 4.04; Linguistic fluency: 3.77)
The authors included the following disclaimer in their Discussion section: “there is a risk of patients’ private information being leaked to commercial enterprises if LLMs are used in actual hospitals.”
Kiyomiya et al., 2024 [42] Japan To evaluate responses from ChatGPT to consumer inquiries about various over-the-counter (OTC) drugs Cross-sectional ChatGPT-3.5 (OpenAI) Answering medical questions, as a substitute/adjunct to secure messaging or seeing clinician Exchanging Information N/A ChatGPT’s answers were evalatued by the following criteria: 1) coherence between the question and response, 2) scientific correctness, and 3) appropriateness of the instructed actions. The proportions of ChatGPT’ s answers that satisfied criteria 1, 2, and 3 were 79.5 %, 54.5 %, and 49.6 %, respectively. The proportion of responses that satisfied all three was only 20.8%. Only 61.8% of the responses that satsfied all criteria were reproduced when the same question was input again on a different day. Questions using brand names rather than generic names resulted in lower coherence and scientific correctness.
Lybarger, Yetisgen et al., 2023 [41] United States To explore the extraction of social determinants of health (SDOH) from EHR clinical notes, through a national competition between 15 teams utilizing a variety of NLP techniques (n2c2/UW SDOH Challenge) Cross-sectional Pre-trained language models, including BERT (Google) and T5 (Google) Summarizing previous care records (“chart review”) Fostering Healing Relationships; Exchanging Information 15 teams participated in the competition. Pretrained language models performed best, including generalizability and learning transfer.
Extraction performance varied by SDOH, with lower performance for conditions that increased health risks and higher performance for conditions that reduced health risks (protective factors).
N/A
Lysø et al., 2024 [38] Norway To explore men’s expectations of the future involvement of artificial intelligence in prostate cancer diagnostics Cross-sectional (Not specified) Augmenting clinician input on decision-making (e.g., diagnoses, treatment recommendations); Co-writing educational materials, follow-up instructions, or responses to patient secure messages Fostering Healing Relationships; Responding to Emotions; Making Decisions Confidence in large amounts of data led participating men to expect that AI could contribute to early detection of prostate cancer for individuals with few symptoms, undefinable symptoms, or no symptoms. Participating men attributed doctors with accountability for interpreting and quality-controlling AI output and interpretations.
Doctors were expected to take on the role of technical controller or supervisor and to “press the stop button” if AI made mistakes.
Participants expressed concerns that integrating AI into their interpersonal relationship with their doctor would exclude them from decision-making and threaten patient autonomy.
Maida et al., 2024 [67] Italy To assess the preferences, satisfaction, and perceived empathy of people living with Multiple Sclerosis (PwMS), toward responses to four frequently-asked questions when authored by neurologists verses ChatGPT Cross-sectional ChatGPT-3.5 (OpenAI) Answering medical questions, as a substitute/adjunct to secure messaging or seeing clinician Exchanging Information; Responding to Emotions; Enabling Self-Management 1133 PwMS (age, 45.26 ± 11.50 years; females, 68.49 %) participated in the study. ChatGPT’s responses received significantly higher empathy ratings compared to neurologists’ responses. No association was found between ChatGPT’ responses and mean patient satisfaction. College graduates had significantly lower likelihood to prefer ChatGPT responses than participants with high school education. N/A
McMahon & McMahon, 2024 [111] United States To examine the accuracy of ChatGPT responses to common questions about self-managed abortion safety and usage of abortion pills Cross-sectional ChatGPT-3.5 (OpenAI) Answering medical questions, as a substitute/adjunct to secure messaging or seeing clinician Exchanging Information; Enabling Self-Management ChatGPT correctly described clinician-managed medication abortion as both effective and safe. ChatGPT provided responses that overstated the risk of complications associated with self-managed medication abortion in ways that contradicted evidence.
Mizuta et al., 2024 [80] Japan To measure the accuracy of ChatGPT-4.0 compared to medical professionals, in evaluating differential diagnoses Cross-sectional ChatGPT-4.0 (OpenAI) Augmenting clinician input on decision-making (e.g., diagnoses, treatment recommendations) Making Decisions For each of 82 cases, three sets of differential diagnoses were evaluated, resulting in 246 lists total. The proportion of concordant responses between ChatGPT-4.0 and physicians was 236 out of 246 (95.9 %; kappa = 0.86). In order to judge the correctness of a diagnosis, ChatGPT-4.0 only compared the names of the final and differential diagnoses; unlike the human physicians in the study, it demonstrated difficulty with inferring background context relevant to individual cases.
Nayak et al., 2023 [52] United States To examine the ability of ChatGPT to write a history of present illness (HPI) summary, compared to senior internal medicine residents Cross-sectional ChatGPT-3.5 (OpenAI) Summarizing previous care records (“chart review”) Exchanging Information When ChatGPT responses were produced through multiple rounds of prompt engineering to improve the quality of responses, ratings by attending physicians of resident- and chatbot-generated HPIs differed by less than 1 point on the 15-point composite scale used in the study (resident mean 12.18 vs. chatbot mean 11.23; P =.09). Attending physicians correctly classified HPIs as written by residents or ChatGPT with an accuracy of only 61 %, suggesting that chatbot-generated HPIs were not highly distinguishable from resident-generated HPIs. Resident HPIs scored higher than ChatGPT responses produced by prompt engineering, on a 5-point level-of-detail scale (resident mean 4.13 vs. chatbot mean 3.57; P =.006). ChatGPT’s performance was heavily dependent on prompt quality. Without rigorous prompt engineering, ChatGPT frequently hallucinated patient demographic factors or other information that was not part of the source material being summarized.
Nov et al., 2023 [39] United States To examine the ability of average Americans to distinguish between clinician- and ChatGPT-authored responses to medical questions, and to examine participants’ degree of trust toward AI to answer medical questions Cross-sectional ChatGPT-3.5 (OpenAI) Answering medical questions, as a substitute/adjunct to secure messaging or seeing clinician Fostering Healing Relationships On average, ChatGPT authorship was identified correctly in 65.5 % (1284/1960) of cases, and human authorship was identified correctly in 65.1 % (1276/1960) of cases. On average, patients’ trust in ChatGPT to answer medical questions was weakly positive (mean score of 3.4 out of 5), with lower trust as health-related complexity of questions increased.
Owens et al., 2024 [36] United States To assess how the use of DAX™ (a system that uses ambient voice recognition, coupled with natural language processing and artificial intelligence) affects the patient-physician relationship Prospective cohort DAX TM, a system that uses ambient voice recognition, natural language processing, and artificial intelligence to generate notes on medical encounters Transcribing patient answers to medical interview during visit (e.g., illness history, care preferences) Fostering Healing Relationships In the open-label phase, 288 patients participated in the survey. Patients “strongly agreed” that phyicians using DAX TM were more focused on them (75.4 %; P < 0.001), spent less time typing (78.8 %; P < 0.001) and made encounters feel more personable than in previous encounters (80.9 %; P < 0.001). No evidence was found to suggest that routine (masked) utilization of DAX™ was associated with any significant differences in quality of the patient-physician relationship as measured by the PDRQ-9 scale or in how well patients thought physicians listened to them during encounters. N/A
Park et al., 2024 [48] South Korea To assess the efficacy of AI-generated radiology reports for spine MRIs in terms of report summary, patient-friendliness, and recommendations and to evaluate the consistent performance of report quality and accuracy, contributing to the advancement of radiology workflow Cross-sectional GPT-3.5-turbo (OpenAI) Simplifying reports on test or scan results; Augmenting clinician input on decision-making (e.g., diagnoses, treatment recommendations) Exchanging Information; Managing Uncertainty; Enabling Self-Management When radiology reports were generated by ChatGPT in three formats (summary reports, patient-friendly reports, and recommendations), reports across all three formats received high average scores from both physician and non-physician raters. The average comprehension score by non-physician raters on a 5-point Likert scale for the original report was 2.71 ± 0.73, while the average score for “patient-friendly” reports generated by ChatGPT significantly increased to 4.69 ± 0.48 (p < 0.001). Physician raters identified 23 cases (1.12 %) of artificial hallucination with clinically significant implications, out of 2055 total translated spine MRI reports by ChatGPT.
Hallucinations with clinically significant implications were observed by raters only in the patient-friendly reports in 14 cases. In 9 cases, they were observed concurrently in all three materials, including the summary and recommendation.
In addition, other translations that could potentially lead to misunderstandings but were not categorized as full-on hallucinations occurred in 152/2055 (7.40 %) cases.
Rabbani et al., 2024 [112] United States To evaluate the performance of ChatGPT-3.5 in identifying content related to mental health, sexual health, and substance use in adolescents’ EHR progress notes Cross-sectional ChatGPT-3.5 (OpenAI) Summarizing previous care records (“chart review”) Fostering Healing Relationships ChatGPT achieved 97 % sensitivity in detecting confidential adolescent EHR note content. ChatGPT achieved only 18 % specificity and 34 % positive predictive value in detecting confidential adolescent EHR note content.
Of the excerpts that ChatGPT suggested contained confidential information, 87 % contained hallucinations.
Rogasch et al., 2023 [64] Germany To evaluate the adequacy of answers provided by ChatGPT to questions frequently asked by patients before and after [18 F] FDG PET/CT scans and regarding scan reports Cross-sectional ChatGPT-4.0 (OpenAI) Co-writing educational materials, follow-up instructions, or responses to patient secure messages; Answering medical questions, as a substitute/adjunct to secure messaging or seeing clinician Exchanging Information; Responding to Emotions; Managing Uncertainty; Enabling Self-Management 92 % of ChatGPT’s responses to 25 questions about scans, questions about scan reports, and follow-up questions on scan reports were considered by nuclear medicine staff to be “appropriate,” 92 % were considered “useful,” and 83 % were considered “empathetic.” When ChatGPT was asked to answer the same questions a second time, considerable inconsistencies (such as differences in tumor stage described) were observed between regenerated responses for 16 % of questions.
Rouhi et al., 2024 [62] United States To assess whether two freely available generative AI dialogue platforms could rewrite online aortic stenosis (AS) patient education materials (PEMs) to meet recommended reading skill levels for the public Cross-sectional ChatGPT-3.5 (OpenAI); Bard (Google) Co-writing educational materials, follow-up instructions, or responses to patient secure messages Exchanging Information; Enabling Self-Management Readability measures for original online material had difficult readability (10th–12th grade reading level). ChatGPT-3.5 succeeded at rewriting materials at approximately 6th–7th grade reading level across all readability measures (p < 0.001). Bard successfully rewrote materials at approximately 8th–9th grade level across all readability measures improved readability across all measures (p < 0.001) except for SMOGI (p = 0.729). N/A
Sarangi et al., 2023 [86] India To assess the performance of ChatGPT-3.5 in simplifying radiologic reports to improve understanding among clinicians and patients Cross-sectional ChatGPT-3.5 (OpenAI) Simplifying reports on test or scan results Fostering Healing Relationships; Exchanging Information; Making Decisions ChatGPT received high ratings from radiologists on its responses to questions about how to interpret radiology reports, in terms of the responses’ level of detail, coverage of key findings, simplicity of content, simplicity of language, factual correctness, quality of information for patients, quality of conclusions drawn for patients, and further suggestions (all higher than 4 out of 5 possible points). When asked to translate the report into Hindi, Chat GPT produced a translation that was rated low by radiologists in terms of completeness (2.89/5) and grammatical correctness (2.32/5). The Hindi translation was not consdered to be suitable for communication with patients.
Saturno et al., 2023 [81] United States To evaluate the utility of ChatGPT-4.0 as an adjunctive tool for breast surgery decision-making by comparing its responses to clinical questions to three key surgical guidelines Cross-sectional ChatGPT-3.5 (OpenAI) Augmenting clinician input on decision-making (e.g., diagnoses, treatment recommendations) Making Decisions N/A ChatGPT’ s responses to questions about breast surgery decision-making were 31.3 % fully concordant, 40.6 % partially concordant, and 28.1 % nonconcordant with guidelines created by the American Society of Plastic Surgeons (ASPS).
Scott et al., 2013 [113] United Kingdom To assess the efficacy and utility of AI-generated textual summaries of patients’ medical histories Quasi-experimental The “Report Generator,” a natural language generation system that uses NLP to extract information from medical narratives and then aggregates the information into textual summaries. Summarizing previous care records (“chart review”) Exchanging Information Clinicians provided correct answers to questions about patient histories 80 % of the time when using AI-generated summaries, and only 75 % of the time when using full records from EHRs, although this difference was not significant. Clinicians were able to answer questions in significantly less time (over 50 % less) when using AI-generated summaries, compared to full records. Clinicians also provided overwhelmingly positive feedback on the perceived usefulness of the AI-generated summaries. N/A
Searle et al., 2023 [114] United Kingdom To test and describe a range of methods and models for summarization of brief hospital encounters Quasi-experimental An adapted version of the BART model (Meta); T5 (Google); BERT (Google) Summarizing previous care records (“chart review”) Exchanging Information The models tested by the authors achieved 70 % coherence (28/ 40 summaries) vs. 75 % (30/40) achieved by a gold-standard reference summary. The models tested by the authors achieve 60 % fluency (24/40) vs. 70 % (28/ 40) achieved by the reference. Relevance improved slightly among the guidance models (73 % vs.
70 %), compared to 80 % relevance achieved by the reference summary.
The authors’ guidance model achieved 58 % consistency (23/40) vs. 90 % (36/40) achieved by the reference summary.
Shiraishi et al., 2024 [115] Japan To assess the performance of ChatGPT in answering a series of fictional patient questions, emulating a preliminary consultation on blepharoplasty Cross-sectional ChatGPT, version unspecified (OpenAI) Co-writing educational materials, follow-up instructions, or responses to patient secure messages Exchanging Information N/A Nine questions were generated with reference to American Society of Plastic Surgeons websites. All questions were posed to ChatGPT in the context of a blepharoptosis consultation. Average accuracy was rated by board-certified plastic surgeons as 3.08 ± 0.68 and by non-medical personnel as 3.69. For average informativeness, ratings by board-certified plastic surgeons were significantly lower than ratings by non-medical staff members (2.89 vs. 4.41; p = 0.04). The overall accessibility ratings of answers did not differ between the two groups of raters (3.02 vs 1.94; p = 0.11), with the exception of one question (p = 0.04).
Szczesniewski et al., 2024 [116] Spain To evaluate the quality of information produced by three GenAI models regarding common urological pathologies Cross-sectional ChatGPT, version unspecified (OpenAI); Bard (Google); Copilot (Microsoft) Answering medical questions, as a substitute/adjunct to secure messaging or seeing clinician Exchanging Information Overall, the analyzed generative AIs responses had higher DISCERN questionnaire (quality) scores than scores from previous studies analyzing medical information on social networking sites. The quality of answers obtained from GenAI depended on how questions were formulated. The information produced contained significant biases that differed depending on pathology and which GenAI was consulted.
Tai-Seale et al., 2024 [117] United States To measure the association between GenAI assistance with drafting replies to patient messages and physician time spent answering messages, as well as the length of replies. Randomized controlled trial Generative AI integrated into a health system’s EHR (no further details provided on model) Co-writing educational materials, follow-up instructions, or responses to patient secure messages Fostering Healing Relationships; Exchanging Information; Responding to emotions GenAI-drafted replies were associated with no change in reply time and significantly increased reply length. Physician views on benefits and drawbacks of the draft messages were mixed, with some stating that AI infused helpful empathetic phrasing into messages, while others thought the tone of drafted messages was unhelpful. GenAI-drafted replies were associated with significantly increased read time for physicians.
Tierney et al., 2024 [32] United States To evaluate a recently enabled ambient artificial intelligence (AI) scribe technology, in terms of its appeal to clinicians given the option of utilizing it, its effectiveness at reducing documentation burden for clinicians, its enhancement of the physician-patient relationship, and the quality of its doumentation Prospective cohort An ambient AI scribe that converts speech collected from microphones on clinicians’ secure smartphones into text and applies natural language processing to summarize key clinical content (no further model details provided) Co-writing visit notes; Transcribing patient answers to medical interview during visit (e.g., illness history, care preferences) Fostering Healing Relationships; Exchanging Information; Managing Uncertainty A review of 35 AI-generated transcripts yielded an average score of 48 out of 50 on a documentation quality measurement. There was a statistically significant reduction in time spent on notes per appointment, for physicians who used the ambient AI scribe compared to those who did not.
Based on preliminary data in a sample of 21 patient surveys from a single clinic site, 71 % reported they spent more time speaking with their physician, while one said they spent less time. Overall, 81 % of patients reported that their physician spent less time looking at the computer screen than in their previous visits. All patients stated that the AI scribe either had no effect or enhanced their visit.
AI-generated transcripts frequently produced inconsistencies that required physicians’ review and editing. There were a few instances of hallucination or missing key information.
Van Nuland et al., 2024 [83] The Netherlands To examine the performance of ChatGPT in giving advice on clinical rule-guided dose interventions for hospitalized patients with renal impairment. Cross-sectional ChatGPT-3.5 (OpenAI) Augmenting clinician input on decision-making (e.g., diagnoses, treatment recommendations) Making Decisions N/A ChatGPT was presented with prompts that included an alert from the EHR (specifying lab test results and current medication dosage) and a question asking what a healthcare professional would recommend regarding the medication. ChatGPT’s responses were compared to responses from an expert panel consisting of a clincal pharmacologist and a nephrologist.
For alerts presented without patient variables, ChatGPT’s responses were incorrect in 53.4 % of cases. For alerts including patient variables, ChatGPT’s responses were incorrect in 67.3 % of cases.
Van Veen et al., 2024 [74] United States To evaluate the completeness, correctness, and conciseness of clinical summaries generated by LLMs for radiology reports, patient questions, progress notes, and doctor-patient dialogue Cross-sectional FLAN-T5 (Google); FLAN-UL2 (Google); Alpaca (Meta); Med-Alpaca (Meta); Vicuna (Meta); Llama-2 (Meta); GPT-3.5 (OpenAI); GPT-4.0 (OpenAI) Simplifying reports on test or scan results; Summarizing previous care records (“chart review”) Exchanging Information Pooled results across ten physicians showed that summaries produced by the best adapted model (GPT-4 using ICL) were more complete and contained fewer errors compared to summaries written by medical experts. N/A
Wei et al., 2023 [78] China To evaluate the performance of ChatGPT-4.0 compared to pediatricians, at neurodevelopmental differential diagnosis tasks Quasi-experimental ChatGPT-4.0 (OpenAI) Augmenting clinician input on decision-making (e.g., diagnoses, treatment recommendations) Making Decisions The overall diagnostic accuracy of pediatricians was 60 % (senior pediatricians achieved 66.7 % accuracym and junior pediatricians achieved 55.3 %). ChatGPT achieved diagnostic accuracy of 66.7 %, which was comparable to senior pediatricians. Inter-group diagnostic agreement between ChatGPT and pediatricians yielded kappa = 0.43 (P < 0.001). When additional medical vignettes were provided within the tasks, overall accuracy of pediatricians improved to 66.7 % (senior pediatricians achieved 73.3 % accuracy, and junior pediatricians achieved 60.0 %), while ChatGPT’s accuracy decreased to 53.3 %. In this phase of the study, the kappa value between ChatGPT and pediatricians was 0.35 (P < 0.001).
Zaretsky et al., 2024 [118] United States To determine whether an LLM can transform discharge summaries into a format that is more readable and understandable Cross-sectional Microsoft Azure OpenAI LLMs (no further details on models provided) Summarizing previous care records (“chart review”) Exchanging Information Mean Flesch-Kincaid Grade Level was significantly lower in “patient-friendly” discharge summaries generated by LLMs (consistently at 6th or 7th grade reading level), compared to original summaries. Two physician raters reviewed patient-friendly discharge summaries for accuracy on a 6-point scale, and 54 of 100 reviews (54.0 %) achieved the highest possible rating of 6. Summaries were rated entirely complete in 56 out of 100 reviews (56.0 %).
18/100 reviews of summaries noted safety concerns, some of which were omissions and some of which were inaccurate statements (hallucinations).
Zhang et al., 2024 [69] Singapore To assess the accuracy and relevance of ChatGPT responses to frequently asked questions about total knee replacement Cross-sectional ChatGPT-3.5 (OpenAI) Answering medical questions, as a substitute/adjunct to secure messaging or seeing clinician Exchanging Information; Managing Uncertainty; Enabling Self-Management 44/50 (88 %) ChatGPT responses were rated as accurate, achieving a mean rating of 4.6/5 for accuracy. 50/50 (100 %) responses were classified as relevant, achieving a mean rating of 4.9/5 for relevance. 6/50 (12.0 %) of ChatGPT responses were rated as inaccurate.

Footnotes

CRediT authorship contribution statement

Brian D. Carpenter: Writing – review & editing, Resources. Jessica Hahne: Writing – review & editing, Writing – original draft, Methodology, Formal analysis, Conceptualization.

Declaration of Competing Interest

The authors declare the following financial interests/personal relationships which may be considered as potential competing interests: Jessica Hahne reports financial support was provided by National Institute on Aging. If there are other authors, they declare that they have no known competing financial interests or personal relationships that could have appeared to influence the work reported in this paper.

References

  • [1].McKinsey & Company, What is ChatGPT, DALL-E, and generative AI? | McKinsey, (2024). 〈https://www.mckinsey.com/featured-insights/mckinsey-explainers/what-is-generative-ai〉 (accessed August 24, 2024). [Google Scholar]
  • [2].Ashley M AI and IT: How Artificial Intelligence Is Changing Information Technology Jobs. Univ. Cincinnati; 2024. 〈https://online.uc.edu/blog/how-ai-is-changing-it-jobs/〉 (accessed August 24, 2024).
  • [3].Landi H, How Epic is charging ahead to bring generative AI into healthcare, (2023). 〈https://www.fiercehealthcare.com/health-tech/epic-moves-forward-bring-generative-ai-healthcare-heres-why-handful-health-systems-are〉 (accessed August 28, 2024).
  • [4].Landi H, HIMSS24: How Epic is building out AI, ambient tech in EHRs, (2024). 〈https://www.fiercehealthcare.com/ai-and-machine-learning/himss24-how-epic-building-out-ai-ambient-technology-clinicians〉 (accessed August 24, 2024).
  • [5].American Medical Association, AMA Augmented Intelligence Research, (2023). [Google Scholar]
  • [6].Wang X, Cohen RA. Health Information Technology Use Among Adults: United States, July-December 2022. Hyattsville, MD: National Center for Health Statistics (U.S.); 2023. 10.15620/cdc:133700. [DOI] [Google Scholar]
  • [7].Rainie L, Close encounters of the AI kind:, Elon University Imagining the Digital Future Center, 2025. 〈https://imaginingthedigitalfuture.org/wp-content/uploads/2025/03/ITDF-LLM-User-Report-3-12-25.pdf〉 (accessed April 3, 2025).
  • [8].Balint M, Ball DH, Hare ML. Training medical students in patient-centered medicine. Compr Psychiatry 1969;10:249–58. 10.1016/0010-440X(69)90001-7. [DOI] [PubMed] [Google Scholar]
  • [9].Epstein RM, Street RL Jr., Patient-centered communication in cancer care: Promoting healing and reducing suffering, (2007). 10.1037/e481972008-001. [DOI] [Google Scholar]
  • [10].Langberg EM, Dyhr L, Davidsen AS. Development of the concept of patient-centredness – A systematic review. Patient Educ Couns 2019;102:1228–36. 10.1016/j.pec.2019.02.023. [DOI] [PubMed] [Google Scholar]
  • [11].Latimer T, Roscamp J, Papanikitas A. Patient-centredness and consumerism in healthcare: an ideological mess. J R Soc Med 2017;110:425–7. 10.1177/0141076817731905. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [12].Rathert C, Mittler JN, Banerjee S, McDaniel J. Patient-centered communication in the era of electronic health records: What does the evidence say? Patient Educ Couns 2017;100:50–64. 10.1016/j.pec.2016.07.031. [DOI] [PubMed] [Google Scholar]
  • [13].Tavares APDS, Paparelli C, Kishimoto CS, Cortizo SA, Ebina K, Braz MS, Mazutti SRG, Arruda MJC, Antunes B. Implementing a patient-centred outcome measure in daily routine in a specialist palliative care inpatient hospital unit: An observational study. Palliat Med 2017;31:275–82. 10.1177/0269216316655349. [DOI] [PubMed] [Google Scholar]
  • [14].Goldfarb MJ, Bibas L, Bartlett V, Jones H, Khan N. Outcomes of Patient- and Family-Centered Care Interventions in the ICU: A Systematic Review and Meta-Analysis. Crit Care Med 2017;45:1751. 10.1097/CCM.0000000000002624. [DOI] [PubMed] [Google Scholar]
  • [15].Asmat K, Dhamani K, Gul R, Froelicher ES. The effectiveness of patient-centered care vs. usual care in type 2 diabetes self-management: A systematic review and meta-analysis. Front Public Health 2022;10. 10.3389/fpubh.2022.994766. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [16].Plewnia A, Bengel J, Körner M. Patient-centeredness and its impact on patient satisfaction and treatment outcomes in medical rehabilitation. Patient Educ Couns 2016;99:2063–70. 10.1016/j.pec.2016.07.018. [DOI] [PubMed] [Google Scholar]
  • [17].Pereira L, Figueiredo-Braga M, Carvalho IP. Preoperative anxiety in ambulatory surgery: The impact of an empathic patient-centered approach on psychological and clinical outcomes. Patient Educ Couns 2016;99:733–8. 10.1016/j.pec.2015.11.016. [DOI] [PubMed] [Google Scholar]
  • [18].Leventhal H, Carr S. Speculations on the relationship of behavioral therapy to psychosocial research on cancer. In: Baum A, Andersen BL, editors. Psychosoc. Interv. Cancer Washington: American Psychological Association; 2001. p. 375–400. 10.1037/10402-020. [DOI] [Google Scholar]
  • [19].Kleinman A Patients and Healers in the Context of Culture: An Exploration of the Borderland Between Anthropology, Medicine, and Psychiatry. University of California Press; 1980. [Google Scholar]
  • [20].Mishel MH. Uncertainty in chronic illness. Annu Rev Nurs Res 1999;17:269–94. [PubMed] [Google Scholar]
  • [21].Charles C, Gafni A, Whelan T. Decision-making in the physician–patient encounter: revisiting the shared treatment decision-making model. Soc Sci Med 1999;49:651–61. 10.1016/S0277-9536(99)00145-8. [DOI] [PubMed] [Google Scholar]
  • [22].Deci EL, Ryan RM. Intrinsic Motivation and Self-Determination in Human Behavior. Boston, MA: Springer US; 1985. 10.1007/978-1-4899-2271-7. [DOI] [Google Scholar]
  • [23].Johnson C, Richwine C, Patel V, Individuals’ Access and Use of Patient Portals and Smartphone Health Apps, 2020, (2021). [PubMed] [Google Scholar]
  • [24].OpenNotes, Our History: Fifty Years In the Making, Opennotes.Org (2021). 〈https://www.opennotes.org/history/〉 (accessed August 23, 2024).
  • [25].Strawley C, Richwine C, Individuals’ Access and Use of Patient Portals and Smartphone Health Apps, 2022, (2023). [PubMed] [Google Scholar]
  • [26].Funk Giancarlo Pasquini AT. Alison Spencer and Cary, 60% of Americans Would Be Uncomfortable With Provider Relying on AI in Their Own Health Care. Pew Res Cent 2023. 〈https://www.pewresearch.org/science/2023/02/22/60-of-americans-would-be-uncomfortable-with-provider-relying-on-ai-in-their-own-health-care/〉 (accessed July 31, 2024). [Google Scholar]
  • [27].Arksey H, O’Malley L. Scoping studies: towards a methodological framework. Int J Soc Res Method 2005;8:19–32. 10.1080/1364557032000119616. [DOI] [Google Scholar]
  • [28].Munn Z, Peters MDJ, Stern C, Tufanaru C, McArthur A, Aromataris E. Systematic review or scoping review? Guidance for authors when choosing between a systematic or scoping review approach. BMC Med Res Method 2018;18:143. 10.1186/s12874-018-0611-x. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [29].Tricco AC, Lillie E, Zarin W, O’Brien KK, Colquhoun H, Levac D, Moher D, Peters MDJ, Horsley T, Weeks L, Hempel S, Akl EA, Chang C, McGowan J, Stewart L, Hartling L, Aldcroft A, Wilson MG, Garritty C, Lewin S, Godfrey CM, Macdonald MT, Langlois EV, Soares-Weiser K, Moriarty J, Clifford T, Tunçalp Ö, Straus SE. PRISMA Extension for Scoping Reviews (PRISMA-ScR): Checklist and Explanation. Ann Intern Med 2018;169:467–73. 10.7326/M18-0850. [DOI] [PubMed] [Google Scholar]
  • [30].Ouzzani M, Hammady H, Fedorowicz Z, Elmagarmid A. Rayyan—a web and mobile app for systematic reviews. Syst Rev 2016;5:210. 10.1186/s13643-016-0384-4. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [31].Barak-Corren Y, Wolf R, Rozenblum R, Creedon JK, Lipsett SC, Lyons TW, Michelson KA, Miller KA, Shapiro DJ, Reis BY, Fine AM. Harnessing the Power of Generative AI for Clinical Summaries: Perspectives From Emergency Physicians. 00078–7 Ann Emerg Med 2024;S0196-0644(24). 10.1016/j.annemergmed.2024.01.039. [DOI] [PubMed] [Google Scholar]
  • [32].Tierney AA, Gayre G, Hoberman B, Mattern B, Ballesca M, Kipnis P, Liu V, Lee K. Ambient Artificial Intelligence Scribes to Alleviate the Burden of Clinical Documentation. NEJM Catal 2024;5:CAT.23.0404. 10.1056/CAT.23.0404. [DOI] [Google Scholar]
  • [33].Garcia P, Ma SP, Shah S, Smith M, Jeong Y, Devon-Sand A, Tai-Seale M, Takazawa K, Clutter D, Vogt K, Lugtu C, Rojo M, Lin S, Shanafelt T, Pfeffer MA, Sharp C. Artificial Intelligence–Generated Draft Replies to Patient Inbox Messages. JAMA Netw Open 2024;7:e243201. 10.1001/jamanetworkopen.2024.3201. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [34].Tai-Seale M, Dillon EC, Yang Y, Nordgren R, Steinberg RL, Nauenberg T, Lee TC, Meehan A, Li J, Chan AS, Frosch DL. Physicians’ Well-Being Linked To In-Basket Messages Generated By Algorithms In Electronic Health Records. Health Aff Proj Hope 2019;38:1073–8. 10.1377/hlthaff.2018.05509. [DOI] [PubMed] [Google Scholar]
  • [35].Blease C, Worthen A, Torous J. Psychiatrists’ experiences and opinions of generative artificial intelligence in mental healthcare: An online mixed methods survey. Psychiatry Res 2024;333:115724. 10.1016/j.psychres.2024.115724. [DOI] [PubMed] [Google Scholar]
  • [36].Owens LM, Wilda J, Westendorp J, Grifka R, Fletcher J. Effect of Ambient Voice Technology, Natural Language Processing, and Artificial Intelligence on the Patient-Physician Relationship. Appl Clin Inf 2024;15:660–7. 10.1055/a-2337-4739. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [37].Denecke K, May R, Rivera-Romero O. Transformer Models in Healthcare: A Survey and Thematic Analysis of Potentials, Shortcomings and Risks. J Med Syst 2024;48:23. 10.1007/s10916-024-02043-5. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [38].Lysø EH, Hesjedal MB, Skolbekken J-A, Solbjør M. Men’s sociotechnical imaginaries of artificial intelligence for prostate cancer diagnostics A focus group study. 1982 Soc Sci Med 2024;347:116771. 10.1016/j.socscimed.2024.116771. [DOI] [PubMed] [Google Scholar]
  • [39].Nov O, Singh N, Mann D. Putting ChatGPT’s Medical Advice to the (Turing) Test: Survey Study. JMIR Med Educ 2023;9:e46939. 10.2196/46939. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [40].Guevara M, Chen S, Thomas S, Chaunzwa TL, Franco I, Kann BH, Moningi S, Qian JM, Goldstein M, Harper S, Aerts HJWL, Catalano PJ, Savova GK, Mak RH, Bitterman DS. Large language models to identify social determinants of health in electronic health records. Npj Digit Med 2024;7:1–14. 10.1038/s41746-023-00970-0. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [41].Lybarger K, Yetisgen M, Uzuner Ö. The 2022 n2c2/UW shared task on extracting social determinants of health. J Am Med Inform Assoc 2023;30:1367–78. 10.1093/jamia/ocad012. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [42].Kiyomiya K, Aomori T, Ohtani H. Comprehensive analysis of responses from ChatGPT to consumer inquiries regarding over-the-counter medications. Pharm 2024;79:24–8. 10.1691/ph.2024.3628. [DOI] [PubMed] [Google Scholar]
  • [43].Chervenak J, Lieman H, Blanco-Breindel M, Jindal S. The promise and peril of using a large language model to obtain clinical information: ChatGPT performs strongly as a fertility counseling tool with limitations. Fertil Steril 2023;120: 575–83. 10.1016/j.fertnstert.2023.05.151. [DOI] [PubMed] [Google Scholar]
  • [44].Goldstein A, Shahar Y, Orenbuch E, Cohen MJ. Evaluation of an automated knowledge-based textual summarization system for longitudinal clinical data, in the intensive care domain. Artif Intell Med 2017;82:20–33. 10.1016/j.artmed.2017.09.001. [DOI] [PubMed] [Google Scholar]
  • [45].Goldstein A, Shahar Y. An automated knowledge-based textual summarization system for longitudinal, multivariate clinical data. J Biomed Inf 2016;61:159–75. 10.1016/j.jbi.2016.03.022. [DOI] [PubMed] [Google Scholar]
  • [46].Ge J, Li M, Delk MB, Lai JC. A Comparison of a Large Language Model vs Manual Chart Review for the Extraction of Data Elements From the Electronic Health Record. Gastroenterology 2024;166:707–709.e3. 10.1053/j.gastro.2023.12.019. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [47].Huang J, Neill L, Wittbrodt M, Melnick D, Klug M, Thompson M, Bailitz J, Loftus T, Malik S, Phull A, Weston V, Heller JA, Etemadi M. Generative Artificial Intelligence for Chest Radiograph Interpretation in the Emergency Department. JAMA Netw Open 2023;6:e2336100. 10.1001/jamanetworkopen.2023.36100. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [48].Park J, Oh K, Han K, Lee YH. Patient-centered radiology reports with generative artificial intelligence: adding value to radiology reporting. Sci Rep 2024;14: 13218. 10.1038/s41598-024-63824-z. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [49].Kernberg A, Gold JA, Mohan V. Using ChatGPT-4 to Create Structured Medical Notes From Audio Recordings of Physician-Patient Encounters: Comparative Study. J Med Internet Res 2024;26:e54419. 10.2196/54419. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [50].Khene Z-E, Bigot P, Mathieu R, Rouprêt M, Bensalah K. Development of a Personalized Chat Model Based on the European Association of Urology Oncology Guidelines: Harnessing the Power of Generative Artificial Intelligence in Clinical Practice. Eur Urol Oncol 2024;7:160–2. 10.1016/j.euo.2023.06.009. [DOI] [PubMed] [Google Scholar]
  • [51].Hu D, Liu B, Zhu X, Lu X, Wu N. Zero-shot information extraction from radiological reports using ChatGPT. Int J Med Inf 2024;183:105321. 10.1016/j.ijmedinf.2023.105321. [DOI] [PubMed] [Google Scholar]
  • [52].Nayak A, Alkaitis MS, Nayak K, Nikolov M, Weinfurt KP, Schulman K. Comparison of History of Present Illness Summaries Generated by a Chatbot and Senior Internal Medicine Residents. JAMA Intern Med 2023;183:1026–7. 10.1001/jamainternmed.2023.2561. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [53].Amin K, Doshi R, Forman HP. Large language models as a source of health information: Are they patient-centered? A longitudinal analysis. Healthcare 2024; 12:100731. 10.1016/j.hjdsi.2023.100731. [DOI] [PubMed] [Google Scholar]
  • [54].Chervonski E, Harish KB, Rockman CB, Sadek M, Teter KA, Jacobowitz GR, Berland TL, Lohr J, Moore C, Maldonado TS. Generative artificial intelligence chatbots may provide appropriate informational responses to common vascular surgery questions by patients. Vascular 2024:17085381241240550. 10.1177/17085381241240550. [DOI] [PubMed] [Google Scholar]
  • [55].Dosso JA, Kailley JN, Robillard JM. What Does ChatGPT Know About Dementia? A Comparative Analysis of Information Quality. J Alzheimers Dis 2024;97: 559–65. 10.3233/JAD-230573. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [56].Ghanem YK, Rouhi AD, Al-Houssan A, Saleh Z, Moccia MC, Joshi H, Dumon KR, Hong Y, Spitz F, Joshi AR, Kwiatt M. Dr. Google to Dr. ChatGPT: assessing the content and quality of artificial intelligence-generated medical information on appendicitis. Surg Endosc 2024;38:2887–93. 10.1007/s00464-024-10739-5. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [57].Hristidis V, Ruggiano N, Brown EL, Ganta SRR, Stewart S. ChatGPT vs Google for Queries Related to Dementia and Other Cognitive Decline: Comparison of Results. J Med Internet Res 2023;25:e48966. 10.2196/48966. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [58].Brach C (Ed.), AHRQ Health Literacy Universal Precautions Toolkit, 3rd Edition, Rockville, MD, 2024. [Google Scholar]
  • [59].Wasir AS, Volgman AS, Jolly M. Assessing readability and comprehension of web-based patient education materials by American Heart Association (AHA) and CardioSmart online platform by American College of Cardiology (ACC): How useful are these websites for patient understanding? Am Heart J Cardiol Res Pr 2023;32:100308. 10.1016/j.ahjo.2023.100308. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [60].Haver HL, Ambinder EB, Bahl M, Oluyemi ET, Jeudy J, Yi PH. Appropriateness of Breast Cancer Prevention and Screening Recommendations Provided by ChatGPT. Radiology 2023;307:e230424. 10.1148/radiol.230424. [DOI] [PubMed] [Google Scholar]
  • [61].Haver HL, Gupta AK, Ambinder EB, Bahl M, Oluyemi ET, Jeudy J, Yi PH. Evaluating the Use of ChatGPT to Accurately Simplify Patient-centered Information about Breast Cancer Prevention and Screening. Radiol Imaging Cancer 2024;6:e230086. 10.1148/rycan.230086. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [62].Rouhi AD, Ghanem YK, Yolchieva L, Saleh Z, Joshi H, Moccia MC, Suarez-Pierre A, Han JJ. Can Artificial Intelligence Improve the Readability of Patient Education Materials on Aortic Stenosis? A Pilot Study. Cardiol Ther 2024;13: 137–47. 10.1007/s40119-023-00347-0. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [63].Carnino JM, Pellegrini WR, Willis M, Cohen MB, Paz-Lansberg M, Davis EM, Grillone GA, Levi JR. Assessing ChatGPT’s Responses to Otolaryngology Patient Questions. Ann Otol Rhinol Laryngol 2024;133:658–64. 10.1177/00034894241249621. [DOI] [PubMed] [Google Scholar]
  • [64].Rogasch JMM, Metzger G, Preisler M, Galler M, Thiele F, Brenner W, Feldhaus F, Wetz C, Amthauer H, Furth C, Schatka I. ChatGPT: Can You Prepare My Patients for [18F]FDG PET/CT and Explain My Reports? J Nucl Med 2023. 10.2967/jnumed.123.266114. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [65].Ayers JW, Poliak A, Dredze M, Leas EC, Zhu Z, Kelley JB, Faix DJ, Goodman AM, Longhurst CA, Hogarth M, Smith DM. Comparing Physician and Artificial Intelligence Chatbot Responses to Patient Questions Posted to a Public Social Media Forum. JAMA Intern Med 2023;183:589–96. 10.1001/jamainternmed.2023.1838. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [66].He W, Zhang W, Jin Y, Zhou Q, Zhang H, Xia Q. Physician Versus Large Language Model Chatbot Responses to Web-Based Questions From Autistic Patients in Chinese: Cross-Sectional Comparative Analysis. J Med Internet Res 2024;26: e54706. 10.2196/54706. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [67].Maida E, Moccia M, Palladino R, Borriello G, Affinito G, Clerico M, Repice AM, Di Sapio A, Iodice R, Spiezia AL, Sparaco M, Miele G, Bile F, Scandurra C, Ferraro D, Stromillo ML, Docimo R, De Martino A, Mancinelli L, Abbadessa G, Smolik K, Lorusso L, Leone M, Leveraro E, Lauro F, Trojsi F, Streito LM, Gabriele F, Marinelli F, Ianniello A, De Santis F, Foschi M, De Stefano N, Morra VB, Bisecco A, Coghe G, Cocco E, Romoli M, Corea F, Leocani L, Frau J, Sacco S, Inglese M, Carotenuto A, Lanzillo R, Padovani A, Triassi M, Bonavita S, Lavorgna L. ChatGPT vs. neurologists: a cross-sectional study investigating preference, satisfaction ratings and perceived empathy in responses among people living with multiple sclerosis. J Neurol 2024;271:4057–66. 10.1007/s00415-024-12328-x. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [68].Baxter SL, Longhurst CA, Millen M, Sitapati AM, Tai-Seale M. Generative artificial intelligence responses to patient messages in the electronic health record: early lessons learned. JAMIA Open 2024;7:ooae028. 10.1093/jamiaopen/ooae028. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [69].Zhang S, Liau ZQG, Tan KLM, Chua WL. Evaluating the accuracy and relevance of ChatGPT responses to frequently asked questions regarding total knee replacement. Knee Surg Relat Res 2024;36:15. 10.1186/s43019-024-00218-5. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [70].Almagazzachi A, Mustafa A, Eighaei Sedeh A, Vazquez Gonzalez AE, Polianovskaia A, Abood M, Abdelrahman A, Muyolema Arce V, Acob T, Saleem B. Generative Artificial Intelligence in Patient Education: ChatGPT Takes on Hypertension Questions. Cureus 16 2024:e53441. 10.7759/cureus.53441. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [71].Gordon EB, Towbin AJ, Wingrove P, Shafique U, Haas B, Kitts AB, Feldman J, Furlan A. Enhancing Patient Communication With Chat-GPT in Radiology: Evaluating the Efficacy and Readability of Answers to Common Imaging-Related Questions. J Am Coll Radiol JACR 2024;21:353–9. 10.1016/j.jacr.2023.09.011. [DOI] [PubMed] [Google Scholar]
  • [72].Chen X, Zhang W, Zhao Z, Xu P, Zheng Y, Shi D, He M. ICGA-GPT: report generation and question answering for indocyanine green angiography images. Br J Ophthalmol 2024:bjo-2023-324446. 10.1136/bjo-2023324446. [DOI] [PubMed] [Google Scholar]
  • [73].Floyd W, Kleber T, Carpenter DJ, Pasli M, Qazi J, Huang C, Leng J, Ackerson BG, Pierpoint M, Salama JK, Boyer MJ. Current Strengths and Weaknesses of ChatGPT as a Resource for Radiation Oncology Patients and Providers. Int J Radiat Oncol 2024;118:905–15. 10.1016/j.ijrobp.2023.10.020. [DOI] [PubMed] [Google Scholar]
  • [74].Van Veen D, Van Uden C, Blankemeier L, Delbrouck J-B, Aali A, Bluethgen C, Pareek A, Polacin M, Reis EP, Seehofnerová A, Rohatgi N, Hosamani P, Collins W, Ahuja N, Langlotz CP, Hom J, Gatidis S, Pauly J, Chaudhari AS. Adapted large language models can outperform medical experts in clinical text summarization. Nat Med 2024;30:1134–42. 10.1038/s41591-024-02855-5. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [75].Kim H, Jin HM, Jung YB, You SC. Patient-Friendly Discharge Summaries in Korea Based on ChatGPT: Software Development and Validation. J Korean Med Sci 2024;39:e148. 10.3346/jkms.2024.39.e148. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [76].Ayoub M, Ballout AA, Zayek RA, Ayoub NF. Mind + Machine: ChatGPT as a Basic Clinical Decisions Support Tool. Cureus 15 2023:e43690. 10.7759/cureus.43690. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [77].Kanjee Z, Crowe B, Rodman A. Accuracy of a Generative Artificial Intelligence Model in a Complex Diagnostic Challenge. JAMA 2023;330:78–80. 10.1001/jama.2023.8288. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [78].Wei Q, Cui Y, Wei B, Cheng Q, Xu X. Evaluating the performance of ChatGPT in differential diagnosis of neurodevelopmental disorders: A pediatricians-machine comparison. Psychiatry Res 2023;327:115351. 10.1016/j.psychres.2023.115351. [DOI] [PubMed] [Google Scholar]
  • [79].Cabral S, Restrepo D, Kanjee Z, Wilson P, Crowe B, Abdulnour R-E, Rodman A. Clinical Reasoning of a Generative Artificial Intelligence Model Compared With Physicians. JAMA Intern Med 2024;184:581–3. 10.1001/jamainternmed.2024.0295. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [80].Mizuta K, Hirosawa T, Harada Y, Shimizu T. Can ChatGPT-4 evaluate whether a differential diagnosis list contains the correct diagnosis as accurately as a physician? Diagn Berl Ger 2024. 10.1515/dx-2024-0027. [DOI] [PubMed] [Google Scholar]
  • [81].Saturno MP, Mejia MR, Wang A, Kwon D, Oleru O, Seyidova N, Henderson PW. Generative artificial intelligence fails to provide sufficiently accurate recommendations when compared to established breast reconstruction surgery guidelines. J Plast Reconstr Aesthetic Surg JPRAS 2023;86:248–50. 10.1016/j.bjps.2023.09.030. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [82].Huo B, Calabrese E, Sylla P, Kumar S, Ignacio RC, Oviedo R, Hassan I, Slater BJ, Kaiser A, Walsh DS, Vosburg W. The performance of artificial intelligence large language model-linked chatbots in surgical decision-making for gastroesophageal reflux disease. Surg Endosc 2024;38:2320–30. 10.1007/s00464-024-10807-w. [DOI] [PubMed] [Google Scholar]
  • [83].van Nuland M, Snoep JD, Egberts T, Erdogan A, Wassink R, van der Linden PD. Poor performance of ChatGPT in clinical rule-guided dose interventions in hospitalized patients with renal dysfunction. Eur J Clin Pharm 2024;80:1133–40. 10.1007/s00228-024-03687-5. [DOI] [PubMed] [Google Scholar]
  • [84].Ayoub NF, Lee Y-J, Grimm D, Divi V. Head-to-Head Comparison of ChatGPT Versus Google Search for Medical Knowledge Acquisition. Otolaryngol Head Neck Surg J Am Acad Otolaryngol Head Neck Surg 2023. 10.1002/ohn.465. [DOI] [PubMed] [Google Scholar]
  • [85].Kienzle A, Niemann M, Meller S, Gwinner C. ChatGPT May Offer an Adequate Substitute for Informed Consent to Patients Prior to Total Knee Arthroplasty—Yet Caution Is Needed. J Pers Med 2024;14:69. 10.3390/jpm14010069. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [86].Sarangi PK, Lumbani A, Swarup MS, Panda S, Sahoo SS, Hui P, Choudhary A, Mohakud S, Patel RK, Mondal H. Assessing ChatGPT’s Proficiency in Simplifying Radiological Reports for Healthcare Professionals and Patients. Cureus 15 2023: e50881. 10.7759/cureus.50881. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [87].He Z, Bhasuran B, Jin Q, Tian S, Hanna K, Shavor C, Arguello LG, Murray P, Lu Z. Quality of Answers of Generative Large Language Models Versus Peer Users for Interpreting Laboratory Test Results for Lay Patients: Evaluation Study. J Med Internet Res 2024;26:e56655. 10.2196/56655. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [88].Aharon KB, Gershfeld-Litvin A, Amir O, Nabutovsky I, Klempfner R. Improving cardiac rehabilitation patient adherence via personalized interventions. PloS One 2022;17:e0273815. 10.1371/journal.pone.0273815. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [89].Pearce CM, Kumarapeli P, de Lusignan S. “Effects of exam room EHR use on doctor-patient communication: a systematic literature review” triadic and other key terms may have identified additional literature. Inform Prim Care 2013;21: 40–2. 10.14236/jhi.v21i1.39. [DOI] [PubMed] [Google Scholar]
  • [90].Berretta S, Tausch A, Ontrup G, Gilles B, Peifer C, Kluge A. Defining human-AI teaming the human-centered way: a scoping review and network analysis. Front Artif Intell 2023;6:1250725. 10.3389/frai.2023.1250725. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [91].Sharma A, Lin IW, Miner AS, Atkins DC, Althoff T. Human–AI collaboration enables more empathic conversations in text-based peer-to-peer mental health support. Nat Mach Intell 2023;5:46–57. 10.1038/s42256-02200593-2. [DOI] [Google Scholar]
  • [92].Yee L, Chui M, Roger Roberts, Why agents are the next frontier of generative AI, McKinsey Digit. (2024). 〈https://www.mckinsey.com/capabilities/mckinsey-digital/our-insights/why-agents-are-the-next-frontier-of-generative-ai〉 (accessed September 21, 2024). [Google Scholar]
  • [93].Avdagovska M, Menon D, Stafinski T. Capturing the Impact of Patient Portals Based on the Quadruple Aim and Benefits Evaluation Frameworks: Scoping Review. J Med Internet Res 2020;22:e24568. 10.2196/24568. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [94].Hahne J, Carpenter BD, Epstein AS, Prigerson HG, Derry-Vick HM. Communication Skills Training for Oncology Clinicians After the 21st Century Cures Act: The Need to Contextualize Patient Portal–Delivered Test Results. JCO Oncol Pr 2023;19:99–102. 10.1200/OP.22.00567. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [95].Pillemer F, Price RA, Paone S, Martich GD, Albert S, Haidari L, Updike G, Rudin R, Liu D, Mehrotra A. Direct Release of Test Results to Patients Increases Patient Engagement and Utilization of Care. PLOS ONE 2016;11:e0154743. 10.1371/journal.pone.0154743. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [96].Wiljer D, Leonard KJ, Urowitz S, Apatu E, Massey C, Quartey NK, Catton P. The anxious wait: assessing the impact of patient accessible EHRs for breast cancer patients. BMC Med Inform Decis Mak 2010;10:46. 10.1186/1472-6947-10-46. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [97].Alpert J, Morris B, Thomson M, Matin K, Brown R. Health Communication: Identifying how Patient Portals Impact Communication in Oncology. Health Commun 2019;34:1395–403. 10.1080/10410236.2018.1493418. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [98].Alpert JM, Dyer KE, Lafata JE. Patient-centered communication in digital medical encounters. Patient Educ Couns 2017;100:1852–8. 10.1016/j.pec.2017.04.019. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [99].Derksen F, Bensing J, Lagro-Janssen A. Effectiveness of empathy in general practice: a systematic review. Br J Gen Pr 2013;63:e76–84. 10.3399/bjgp13X660814. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [100].Keshtkar L, Madigan CD, Ward A, Ahmed S, Tanna V, Rahman I, Bostock J, Nockels K, Wang W, Gillies CL, Howick J. The Effect of Practitioner Empathy on Patient Satisfaction: A Systematic Review of Randomized Trials. Ann Intern Med 2024;177:196–209. 10.7326/M23-2168. [DOI] [PubMed] [Google Scholar]
  • [101].Zhang R, Pakhomov SVS, Arsoniadis EG, Lee JT, Wang Y, Melton GB. Detecting clinically relevant new information in clinical notes across specialties and settings. BMC Med Inform Decis Mak 2017;17:68. 10.1186/s12911-017-0464-y. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [102].Athaluri SA, Manthena SV, Kesapragada VSRKM, Yarlagadda V, Dave T, Duddumpudi RTS. Exploring the Boundaries of Reality: Investigating the Phenomenon of Artificial Intelligence Hallucination in Scientific Writing Through ChatGPT References. Cureus 15 2023:e37432. 10.7759/cureus.37432. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [103].Naik N, Hameed BMZ, Shetty DK, Swain D, Shah M, Paul R, Aggarwal K, Ibrahim S, Patil V, Smriti K, Shetty S, Rai BP, Chlosta P, Somani BK. Legal and Ethical Consideration in Artificial Intelligence in Healthcare: Who Takes Responsibility? Front Surg 2022;9:862322. 10.3389/fsurg.2022.862322. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [104].Altamimi I, Altamimi A, Alhumimidi AS, Altamimi A, Temsah M-H. Artificial Intelligence (AI) Chatbots in Medicine: A Supplement, Not a Substitute. Cureus 15 2023:e40922. 10.7759/cureus.40922. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [105].Fleary SA, Ettienne R. Social Disparities in Health Literacy in the United States. HLRP Health Lit Res Pr 2019;3:e47–52. 10.3928/2474830720190131-01. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [106].Jackson DN, Trivedi N, Baur C. Re-Prioritizing Digital Health and Health Literacy in Healthy People 2030 to Affect Health Equity. Health Commun 2021;36: 1155–62. 10.1080/10410236.2020.1748828. [DOI] [PubMed] [Google Scholar]
  • [107].Ashraf AR, Mackey TK, Fittler A. Search Engines and Generative Artificial Intelligence Integration: Public Health Risks and Recommendations to Safeguard Consumers Online. JMIR Public Health Surveill 2024;10:e53086. 10.2196/53086. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [108].Chuang Y-N, Tang R, Jiang X, Hu X. SPeC: A Soft Prompt-Based Calibration on Performance Variability of Large Language Model in Clinical Notes Summarization. J Biomed Inf 2024;151:104606. 10.1016/j.jbi.2024.104606. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [109].Hanai A, Ishikawa T, Kawauchi S, Iida Y, Kawakami E. Generative artificial intelligence and non-pharmacological bias: an experimental study on cancer patient sexual health communications. BMJ Health Care Inf 2024;31:e100924. 10.1136/bmjhci-2023-100924. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [110].Iscoe M, Socrates V, Gilson A, Chi L, Li H, Huang T, Kearns T, Perkins R, Khandjian L, Taylor RA. Identifying signs and symptoms of urinary tract infection from emergency department clinical notes using large language models, Acad. Emerg. Med. Off. J. Soc. Acad Emerg Med 2024;31:599–610. 10.1111/acem.14883. [DOI] [PubMed] [Google Scholar]
  • [111].McMahon HV, McMahon BD. Automating untruths: ChatGPT, self-managed medication abortion, and the threat of misinformation in a post-Roe world. Front Digit Health 2024;6:1287186. 10.3389/fdgth.2024.1287186. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [112].Rabbani N, Brown C, Bedgood M, Goldstein RL, Carlson JL, Pageler NM, Morse KE. Evaluation of a Large Language Model to Identify Confidential Content in Adolescent Encounter Notes. JAMA Pedia 2024;178:308–10. 10.1001/jamapediatrics.2023.6032. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [113].Scott D, Hallett C, Fettiplace R. Data-to-text summarisation of patient records: using computer-generated summaries to access patient histories. Patient Educ Couns 2013;92:153–9. 10.1016/j.pec.2013.04.019. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [114].Searle T, Ibrahim Z, Teo J, Dobson RJB. Discharge summary hospital course summarisation of in patient Electronic Health Record text with clinical concept guided deep pre-trained Transformer models. J Biomed Inf 2023;141:104358. 10.1016/j.jbi.2023.104358. [DOI] [PubMed] [Google Scholar]
  • [115].Shiraishi M, Tanigawa K, Tomioka Y, Miyakuni A, Moriwaki Y, Yang R, Oba J, Okazaki M. Blepharoptosis Consultation with Artificial Intelligence: Aesthetic Surgery Advice and Counseling from Chat Generative Pre-Trained Transformer (ChatGPT). Aesthetic Plast Surg 2024;48:2057–63. 10.1007/s00266-024-04002-4. [DOI] [PubMed] [Google Scholar]
  • [116].Szczesniewski JJ, Ramos Alba A, Rodríguez Castro PM, Lorenzo Gómez MF, Sainz Gonźalez J, Llanes González L. Quality of information about urologic pathology in English and Spanish from ChatGPT, BARD, and Copilot. Actas Urol Esp 2024; S2173-5786(24):00016-7. 10.1016/j.acuroe.2024.02.009. [DOI] [PubMed] [Google Scholar]
  • [117].Tai-Seale M, Baxter SL, Vaida F, Walker A, Sitapati AM, Osborne C, Diaz J, Desai N, Webb S, Polston G, Helsten T, Gross E, Thackaberry J, Mandvi A, Lillie D, Li S, Gin G, Achar S, Hofflich H, Sharp C, Millen M, Longhurst CA. AI-Generated Draft Replies Integrated Into Health Records and Physicians’ Electronic Communication. JAMA Netw Open 2024;7:e246565. 10.1001/jamanetworkopen.2024.6565. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [118].Zaretsky J, Kim JM, Baskharoun S, Zhao Y, Austrian J, Aphinyanaphongs Y, Gupta R, Blecker SB, Feldman J. Generative Artificial Intelligence to Transform Inpatient Discharge Summaries to Patient-Friendly Language and Format. JAMA Netw Open 2024;7:e240357. 10.1001/jamanetworkopen.2024.0357. [DOI] [PMC free article] [PubMed] [Google Scholar]

RESOURCES