Skip to main content
HHS Author Manuscripts logoLink to HHS Author Manuscripts
. Author manuscript; available in PMC: 2026 Apr 29.
Published in final edited form as: BMJ Qual Saf. 2026 Jun 18;35(7):498–502. doi: 10.1136/bmjqs-2025-019475

Thinking critically about AI documentation quality in primary care

Gordon D Schiff 1,2, Maram Khazen 3,4
PMCID: PMC13123647  NIHMSID: NIHMS2161856  PMID: 41663242

The introduction of artificial intelligence (AI)-driven clinical documentation has taken healthcare by storm. According to many leaders and clinicians, there has never been a technology that has been so impressive and rapidly adopted by clinicians.1 Arising on the soil of wide-spread clinician burnout and frustrations, including with the time consumed by charting clinical encounters, AI vendors have developed modules that facilitate clinical documentation.2 These tools work by recording the encounter (with patient consent) as a digital audio file, which is then transcribed using speech-recognition technology, and then processed by AI to produce a formatted note for the clinician to review and sign. This perspective piece is written by a primary care clinician and patient safety researcher using AI documentation for nearly 2 years (GS), and a PhD communication specialist (MK) directly observing hundreds of encounters with this physician.

From a patient safety perspective, especially diagnostic safety, this new technology offers a myriad of improvement opportunities for transforming patient-clinician interactions and documentation workflow.3 Given the important role clinical notes play in formulating, recording and communicating key information, we argue that healthcare must go beyond being dazzled by the magic of instant note production to more critically evaluate this leading healthcare AI application. We identify key issues that warrant addressing and suggest potential mitigation strategies (table 1) keeping in mind a historical perspective of past health information technologies shortcomings/disappointments (eg, clinical decision support, problem lists, copy/pasted notes, diagnosis software).4 5 We examine several issues that have arisen in our work with AI-documented notes, as well as touch on broader issues raised by this new technology.

Table 1.

Issues warranting addressing for AI-clinical documentation

Issue Potential strategies to address
Need to better characterise qualities of a good clinical note, particularly related to diagnostic assessment. ► Convening discussions and promoting research among clinicians, educators, patients and quality, safety and informatics experts to leverage AI to better define, standardise and optimise note content, organisation and workflow efficiencies.
Balance note production efficiency versus clinician tailoring of notes in own voice. Reliance on assessment plans generated by AI may contribute to cognitive biases and diagnostic errors. ► Time-saving/efficiencies of AI notes should be partially reallocated to enable dedicated reflection of clinical encounters and crafting/editing time.
► Use of direct dictation of assessment (eg, by clinician using voice recognition after visit) rather than relying on AI capture during visit. User-friendly ways to customise AI note templates to match the clinician’s narrative style and practice needs.
Errors in AI notes: how to minimise, detect and mitigate. ► While generally rare, occasionally serious; requires vigilance to ensure careful MD review of notes before signing.
► Potential of advanced AI review tools to help detect and correct.
► Patients can play a role in identifying errors in notes, thereby improving note accuracy and patient safety.22
How will trainees learn to write good notes? Risk of atrophy of note-writing skills. ► Standardised curriculum and training requirements (including minimum number of supervised notes) on clinical documentation best practices before learners access AI to generate notes.23
► Developing and teaching new paradigms for how to write notes in the AI era.
Omission of important elements (eg, richer social history (‘chit chat’), physical examination general description, diagnostic assessments) that may not be verbalised during encounters. ► Vendor attention to this design shortcoming.
► Workflow during and after encounters to ensure missing elements are captured.
Retraining clinicians to better verbalise for patients during encounters and ways to add post encounters.24
Loss of information continuity from prior notes if starting with blank AI notes each visit. ► Listening to and understanding clinicians who copy-pasted prior notes.
► Engineering AI note encounter workflow (including prepopulating information) to ensure prior problems are not lost.
Limitations of ‘one-size fits all’ AI notes ► Deployment of new AI features that allow customisable templates, editing, pausing to insert dictated corrections.
► Making such customisation easier for users to unlock such features.
Balance/optimise note comprehensiveness and succinctness to achieve readability and usefulness. Achieve patient-friendly language while preserving medical precision/accuracy. ► Collapsible note synoptic sections to view more versus less detail.
► Leverage AI capabilities to produce note summaries and versions tailored to patient’s health literacy level.
Patient privacy and trust (eg, risk of privacy breaches, erosion of patient trust with robotised AI notes) ► Enhanced data privacy protections (regulatory, institutional, monitoring).
► Attention to AI note language and careful clinician editing of the AI-drafted notes before signing off is especially important in the era of open notes.
► User-friendly and responsive systems for patient feedback on any concerns/questions they may have regarding the content of their notes.
Dominance of for-profit commercial products limiting cooperative learning. ► Public reporting of errors and transparency about shortcomings, including collating experience and understanding why some clinicians who have tried (often briefly) but abandon AI documentation.
► Regulation and oversight of standards.
Climate impacts of AI electricity and water demand. ► While not unique to AI documentation, quality and safety informed healthcare need to be mindful of environmental impacts, thus work with broader efforts to minimise harms.25
AI, artificial intelligence.

In 2021 two key developments threw the doors wide open for rethinking outpatient clinical notes in the USA and beyond. The first was a revision of coding requirements by the US Centres for Medicare and Medicaid which previously required documentation of a specific number of history and physical examination components. The new regulations simplified note-writing, making notes less about check-box documentation, thereby providing flexibility for more meaningful narrative and recording of medical decision-making. This liberated notes from prior constraints and provided an opportunity to rethink and improve recording of diagnostic assessments and plans.6

A second change in the USA, with similar movements in countries in the Americas, Europe, Africa and Asia, mandated opening notes to patients— giving patients access to their medical record and clinical notes.7 8 No longer were notes solely for communication among healthcare staff or documentation as a ‘cover’ against malpractice liability or for physician reimbursement justification. Rather, each note also serves as an important communication channel with the patient. Thus, verbal explanations and instructions, often easily forgotten in the stress of an encounter, can now be later reviewed by patients and families.9 Notes need to be more sensitive to language (especially stigmatising language), patients’ health literacy and anxieties, the nuances of diagnostic uncertainty and patient-centred follow-up instructions.

IMPACT ON CLINICIAN REFLECTION AND COGNITIVE PROCESSES

Although this transformed role of clinical notes afforded an opportunity to rethink note content and documentation workflow, clinicians are mainly being driven by feeling overburdened by electronic medical records (EMRs) and overall productivity demands. From a diagnosis safety standpoint, we have an opportunity to rethink and redesign notes—especially the assessment section—to make it richer in terms of reflecting differential diagnosis, weighing uncertainties and probabilities, grappling with urgency and crafting and operationalising follow-up plans. Instead, clinicians have sought relief and shortcuts through approaches such as copy/pasting older notes, use of templates and human scribes to offload documentation burden.10 Now, commercial vendors have seized on advances in speech recognition and large language model AI to create AI-documentation tools to streamline note-writing worldwide.11

Beyond overlooking the opportunity to collectively, rigorously and intelligently redesign notes for safety, quality, efficiency and better communication, we now have a new set of challenges. While AI documentation competently compiles what is said aloud in the examination room (with some exceptions below), something important is also lost—the time for the clinician to reflect on and craft their assessment. Previously, this typically occurred after the patient left the examination room, often later in the day or at night, perhaps after test results had returned. While the miracle of having a completed note written within a few seconds after the encounter is an undeniable time-saver for the clinician and a gift for more family time at home, it might also be viewed as a theft of reflection time from the patient. Although there is nothing preventing clinicians from taking the time to carefully edit the AI notes to bring back such reflection moments, we suspect that the temptation to quickly sign the note will be irresistible.1

Relatedly, a more profound problem that we have observed in the notes the author (GS) has been ‘writing’ is the complete absence of an assessment—generally the most important component of the note.6 The AI tool capably summarises the history and the plan, but many of the elements of an actual diagnostic assessment may not be spoken aloud (or in some cases the AI discards them). A quick sign-off of the AI note can overlook that the ‘assessment/plan’ of an AI note does not contain any real ‘assessment’. We illustrate two examples (figure 1) of AI notes lacking any assessments that the author (GS) almost quickly signed. These AI-generated notes that initially lacked any assessment; instead, they had to be added manually. Once this assessment reflection space disappears, how will trainees and clinicians learn to craft and maintain good assessments—skills that were historically honed by repeatedly practising writing thoughtful notes.

Figure 1.

Figure 1

Illustration of lack of meaningful assessment in ‘assessment/plan’ section of AI-generated note requiring manual addition. Current AI-generated notes basic templates have three main sections: history, physical examination, assessment/plan. The left-hand column is a screenshot of the assessment/plan section from two AI-Generated Notes illustrating that there is no actual assessment of the patients’ problems. The right-hand column illustrates the revised notes with the added text (in bold) manually revised by the clinician to add meaningful assessment of the problem. AI, artificial intelligence.

OTHER CONSIDERATIONS OF AI-DOCUMENTED NOTES

We have also noticed omissions of other important elements of a clinical encounter. For example, aspects of the physical examination such as general description (eg, patient is disoriented, confused, has pressure of speech, was anxious or inappropriate, appears chronically ill, in mild respiratory distress) are unlikely to be spoken aloud in front of the patient and risk disappearing from examination documentation. And as discussed elsewhere,12 there is important psychosocial history information that the AI bot has decided is unimportant ‘chit-chat’. Instead, the AI tool often eliminates such ‘chit-chat’, despite this being information primary care clinicians would (and should) record given its importance for relationship building.13 Examples of such information eliminated from the author’s own notes include: patient reporting their daughter has overdosed on drugs, planned a trip to Haiti, or having a stressful job situation—topics important to ask about at the next visit inquiring how their daughter was doing, how was their trip, or were things any better on the job.12

Another unresolved issue with AI-documented notes is the problem of how to incorporate older historical data. Currently, many clinicians copy and paste their prior notes to maintain continuity from past visits. While some disparage this copy/pasting (claiming it is a lazy shortcut, brings forward out-of-date information, adds length/‘note bloat’), it serves an important function in creating a thread to carry forward important prior information to ensure it is not overlooked.14 With AI documentation, while copy/paste could still work, the recommended best practice is to start with a blank slate and let the AI produce a clean note each visit. Otherwise, the AI notes become cluttered and redundant, and the AI tool we use has trouble figuring out where to insert the new encounter information to make it clear what is new versus old. Thus, these AI notes might be characterised as a ‘note without a memory’. This has led many of our specialist colleagues, for example, oncologists, who wish to have a running thread of patients’ cancer history and treatments, to prefer their current manual charting approaches over AI note-taking.

Two final issues warrant additional consideration for safety and quality. The first is the readability and usefulness of clinical notes. Unfortunately, even the most carefully crafted (by human or machine) notes are useless if not used. Here we refer to being read by both other clinicians and patients. Studies show that many notes are never read. For example, when patients present to the emergency department (ED), although primary care clinicians’ notes may contain important information, busy ED doctors often do not review these notes.15 And patients, despite now having access in many jurisdictions, may not open the notes or have problems understanding the language. We should not lose the opportunity to explore what AI documentation can do to help address these problems, perhaps by deploying summarisation tools (for the ED doctors) or customising note language to align with patients’ health literacy or spoken language to minimise disparities.

Finally, there is the problem of patient trust. Trust is an essential element in healthcare.16 If notes contain errors (which AI notes occasionally do, with added risk when clinicians skip careful proofreading before signing), it is easy to see how already declining levels of trust might be further diminished. More subtly— and more importantly—when patients read notes that lack the personalisation of the language their own clinicians might use to craft the note, how will they respond? Deference to AI-generated phrasing risks patients perceiving a disregard for nuances or subtleties a physician might write based on patients’ individual preferences and understandings.

LACK OF TRANSPARENCY LIMITING LEARNING IN COMMERCIALISED AI PRODUCTS

Addressing the questions we raise here calls for concerted efforts to study and learn from how this AI tool can be most safely and wisely deployed. At the very least, the metrics evaluating these tools need to get beyond how many minutes it saves clinicians or clinician satisfaction or burnout ratings.1 17 Much of the knowledge about how these systems work, the errors they make and opportunities for improvement currently reside exclusively with the commercial vendors.18 There needs to be much more transparency to compare the various products and test their claims and performance in the real world. Going beyond proprietary black boxes, there needs to be transparent learning and feedback at multiple levels, from the clinicians and patients, but also what the vendors learn as they compare their notes to the encounters’ voice files, speech-generated transcripts, the AI-drafted notes and clinician edits to each note. These companies are now engaged in mining and analysing this data, but our health systems, clinicians and researchers currently lack access to this data and analyses. We need to be able to compare performance of the competing products in standardised ways, lest the least expensive, though perhaps most flawed, product be foisted on financially strapped institutions and clinicians.

WAY FORWARD

Independent researchers need to be funded and given access to data to evaluate how AI-documentation products are performing. Clinicians will require training on verbalisation practices during encounters to facilitate conveying and capture of important information.6 We need to apply systems engineering, human factors, ethnographic and qualitative research to study work flows and outputs. Even simply studying the clinicians that have tried but abandoned using AI documentation,19 or patients who either do or do not like AI-generated notes, would be extremely valuable.20

Notes are a critical cognitive and communication component of healthcare safety. To what extent have clinicians lost control over the means of production of their notes and how can we take them back to ensure they are used for better, safer care, improved thinking and better communication? Recognising shortcomings in ‘old fashioned’ EMR notes21 cannot be an excuse to not proactively take responsibility for improving AI notes. The next generation of AI documentation will be orders more powerful. This future will likely include vendors not only being able to listen to our conversations but also reading our prior notes to further enhance their products’ deep learning.18 We need to be deeply learning as well.

Funding

This study was funded in part by the Agency for Healthcare Research and Quality (AHRQ) (R01HS030232).

Footnotes

Competing interests None declared.

Patient consent for publication Not applicable.

Ethics approval Not applicable.

Provenance and peer review Not commissioned; externally peer reviewed.

REFERENCES

  • 1.Duggan MJ, Gervase J, Schoenbaum A, et al. Clinician Experiences With Ambient Scribe Technology to Assist With Documentation Burden and Efficiency. JAMA Netw Open 2025;8:e2460637. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 2.You JG, Dbouk RH, Landman A, et al. Ambient Documentation Technology in Clinician Experience of Documentation Burden and Burnout. JAMA Netw Open 2025;8:e2528056. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 3.Ng JJW, Wang E, Zhou X, et al. Evaluating the performance of artificial intelligence-based speech recognition for clinical documentation: a systematic review. BMC Med Inform Decis Mak 2025;25:236. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 4.Doctor WR. Hope, hype and harm at dawn. New York: McGraw-Hill, 2017. [Google Scholar]
  • 5.Shah SN, Amato MG, Garlo KG, et al. Renal medication related clinical decision support (CDS) alerts and overrides in the inpatient setting following implementation of a commercial electronic health record: implications for designing more effective alerts. J Am Med Inform Assoc 2021;28:1081–7. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6.Khazen M, Mirica M, Carlile N, et al. Developing a Framework and Electronic Tool for Communicating Diagnostic Uncertainty in Primary Care: A Qualitative Study. JAMA Netw Open 2023;6:e232218. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 7.Coordinator OoN. About onc’s cures act final rule. The Office of the National Coordinator for Health Information Technology. n.d Available: https://www.healthit.gov/curesrule/overview/about-oncs-cures-act-final-rule [Google Scholar]
  • 8.Howard S. Patients’ access to medical records around the world. BMJ 2024;386:q1481. [DOI] [PubMed] [Google Scholar]
  • 9.McCarthy DM, Waite KR, Curtis LM, et al. What did the doctor say? Health literacy and recall of medical instructions. Med Care 2012;50:277–82. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 10.Schiff GD, Zucker L. Medical Scribes: Salvation for Primary Care or Workaround for Poor EMR Usability? J Gen Intern Med 2016;31:979–81. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 11.Shemtob L, Majeed A, Beaney T. Regulation of AI scribes in clinical practice. BMJ 2025;389:r1248. [DOI] [PubMed] [Google Scholar]
  • 12.Schiff GD. AI-Driven Clinical Documentation - Driving Out the Chitchat? N Engl J Med 2025;392:1877–9. [DOI] [PubMed] [Google Scholar]
  • 13.Sinsky CA, Shanafelt TD, Ristow AM. Radical reorientation of the US health care system around relationships: rebalancing the transactional model. Elsevier; 2022;2022:2194–205. [Google Scholar]
  • 14.Tsou AY, Lehmann CU, Michel J, et al. Safe Practices for Copy and Paste in the EHR. Appl Clin Inform 2017;26:12–34. [Google Scholar]
  • 15.Hripcsak G, Sengupta S, Wilcox A, et al. Emergency department access to a longitudinal medical record. J Am Med Inform Assoc 2007;14:235–8. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16.Birkhäuer J, Gaab J, Kossowsky J, et al. Trust in the health care professional and health outcome: A meta-analysis. PLoS One 2017;12:e0170988. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 17.Tierney AA, Gayre G, Hoberman B, et al. Ambient Artificial Intelligence Scribes: Learnings after 1 Year and over 2.5 Million Uses. NEJM Catalyst 2025;6:25. [Google Scholar]
  • 18.Goodman KE, Morgan DJ. Digital Exhaust or Digital Gold? The Value of AI-Generated Clinical Visit Transcripts. N Engl J Med 2026;394:110–3. [DOI] [PubMed] [Google Scholar]
  • 19.Haberle T, Cleveland C, Snow GL, et al. The impact of nuance DAX ambient listening AI documentation: a cohort study. J Am Med Inform Assoc 2024;31:975–9. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 20.Nair R, Hashmi MM, Kassim SS, et al. The Current State of Digital Scribes in Primary Care: A Scoping Review. J Med Syst 2026;50:2. [DOI] [PubMed] [Google Scholar]
  • 21.Egerton-Warburton D, Lim A, Tan YH, et al. Initial Clinical Impressions Are Absent in Around a Quarter of Adult Emergency Department Patient Consultations: Let’s Get Back to Basics. Emerg Med Australas 2025;37:e70086. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22.Bell SK, Bourgeois F, Dong J, et al. Patient Identification of Diagnostic Safety Blindspots and Participation in “Good Catches” Through Shared Visit Notes. Milbank Q 2022;100:1121–65. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 23.Abdulnour R-E, Gin B, Boscardin CK. Educational Strategies for Clinical Supervision of Artificial Intelligence Use. N Engl J Med 2025;393:786–97. [DOI] [PubMed] [Google Scholar]
  • 24.Cain CH, Davis AC, Broder B, et al. Quality Assurance during the Rapid Implementation of an AI-Assisted Clinical Documentation Support Tool. NEJM AI 2025;2:AIcs2400977. [Google Scholar]
  • 25.Bogmans C, Ganpurev G, Gomez-Gonzalez P, et al. Power hungry: how ai will drive energy demand. SSRN [Preprint] 2025. [Google Scholar]

RESOURCES