Abstract
Objectives
To quantify utilization and impact on documentation time of a large language model-powered ambient artificial intelligence (AI) scribe.
Materials and Methods
This prospective quality improvement study was conducted at a large academic medical center with 45 physicians from 8 ambulatory disciplines over 3 months. Utilization and documentation times were derived from electronic health record (EHR) use measures.
Results
The ambient AI scribe was utilized in 9629 of 17 428 encounters (55.25%) with significant interuser heterogeneity. Compared to baseline, median time per note reduced significantly by 0.57 minutes. Median daily documentation, afterhours, and total EHR time also decreased significantly by 6.89, 5.17, and 19.95 minutes/day, respectively.
Discussion
An early pilot of an ambient AI scribe demonstrated robust utilization and reduced time spent on documentation and in the EHR. There was notable individual-level heterogeneity.
Conclusion
Large language model-powered ambient AI scribes may reduce documentation burden. Further studies are needed to identify which users benefit most from current technology and how future iterations can support a broader audience.
Keywords: artificial intelligence, ambient intelligence, documentation, ambient scribes, informatics
Background and significance
Electronic health record (EHR) documentation burden is a significant contributor to physician burnout.1–8 Strategies to reduce documentation burden, including team documentation,9–13 automatic speech recognition technologies,14 and 2021 E&M coding changes15 have been implemented with variable impact.16 The development of generative AI17 has given rise to a new wave of ambient AI scribes powered by large language models (LLMs). Initial versions of this technology utilizing human scribes to review AI output and manual copy-paste workflows to integrate content into the EHR were promising.18 A subsequent study of a fully autonomous LLM-powered ambient AI scribe without EHR integration reported relative time savings; however, uptake was modest, the time analysis was restricted to primary care physicians, and absolute time savings were not reported.19 Newer fully autonomous and EHR integrated versions that can generate drafts of clinical notes within seconds have even greater potential to effect large-scale change but require further evaluation.
Objectives
To better understand the utility and quantify the time impact of an ambient AI scribe, DAX Copilot (Nuance Communications, Inc.) was deployed in select ambulatory disciplines at an academic medical center. Primary outcome measures included utilization and documentation time per note. Secondary measures included daily documentation time, afterhours EHR time (defined as time spent in the EHR outside of weekdays between 7 AM and 5:30 PM) and total EHR time.
Materials and methods
Setting
The intervention was piloted during a 3-month period from October 2023 to January 2024 in ambulatory settings at a single academic medical center (Stanford Health Care).
Technology integration
Stanford Health Care collaborated with EHR developer Epic (Epic Systems) and Nuance (Microsoft Corporation) to integrate DAX Copilot into clinical documentation workflows. The EHR’s mobile app is used to record the physician-patient interaction. The resulting transcript is then processed by an LLM, which generates drafts of 4 note sections that can be accessed via predefined shortcut phrases (SmartSections): history of present illness (HPI), physical exam (PE), results, and assessment and plan (A&P). Physicians insert one or more of these SmartSections into their existing note templates, which then autopopulate once the recording is complete (Figure S1).20
Study recruitment
Physicians were recruited through convenience and purposive sampling based on interest and suspected documentation burden between August 2023 and December 2024. Exclusion criteria included existing in-person or virtual scribe support, working exclusively with trainees, or lacking access to an iPhone (Apple, Inc.). Users were onboarded in a rolling fashion up to a 50-license limit. Users who perceived that the tool had low utility were allowed to voluntarily relinquish their license for reallocation to a nonstudy participant and were excluded from further analysis. Onboarded users who did not have ambulatory patient encounters during the baseline prepilot period or the pilot period were also excluded from analysis.
Physician training
Pilot users underwent required training prior to gaining access to DAX Copilot. Training consisted of a live onboarding support session where SmartSections were added to existing note templates and mobile app workflows were reviewed. This 30-minute live training was subsequently transitioned to a shorter, asynchronous, self-guided video lesson. Multimodal support materials, including knowledge base articles, demo videos, and drop-in group sessions were also made available to pilot users.
Study analyses
All ambulatory encounters (including procedural and virtual encounters) performed by pilot physicians who had access to the tool for at least 30 days between October 29, 2023 and January 27, 2024 were included in the analysis. The initial version of DAX at the beginning of the study period was 10.6.4 and the final version at the end of the study period was 10.7. Encounters where the patient’s preferred language was not English were excluded. The analysis period dates were selected based on audit reporting periods predefined by the EHR. The tool was available to 21 pilot users prior to the analysis start date for technical troubleshooting. All remaining pilot users were enrolled after October 29, 2023, and the entirety of their pilot experience is represented in the analysis period. Where relevant, outcomes from the pilot period were compared against a 2-month baseline prepilot period from May 28, 2023 to July 29, 2023.
Study measures
All outcomes were based on EHR use measures sourced directly from Epic. All measures were summarized as both means and SD as well as medians and interquartile ranges (IQRs).
Statistical analysis
Physician demographics were described with counts and proportions.
Utilization was calculated as the fraction of encounters seen by the physician for which DAX was utilized and was similarly calculated for the use of individual SmartSections.
Electronic health record audit logs (Epic Signal) were only available as averages for each physician across a reporting period rather than at the level of individual encounters. To calculate the average note time per physician in the prepilot period, the average note time for that physician across all prepilot reporting periods was weighted by the number of notes written during the associated reporting period. Weighted averages were similarly calculated for the pilot period, and the difference in average times between the prepilot and pilot periods was calculated for each physician. A one-sample t-test was used to compare the observed differences against a null hypothesis of no change between the 2 periods. A P-value<.05 was considered significant. A similar analysis was performed for average daily documentation time, afterhours EHR time, and total EHR time, but weighting was performed by the number of days in the study period.
Analysis was performed using the statsmodels package in Python programming language version 3.11.4 (Python Software Foundation).
Ethics approval
The Stanford University Institutional Review Board determined that the project was exempt from formal institutional review board oversight on the basis of quality improvement.
Results
Cohort
Of the 50 onboarded users, 3 relinquished their license due to low perceived utility and 2 were excluded for lack of encounters in either the pilot or baseline periods; 45 were included in the study analysis (Figure S2) representing 8 specialties (Table 1). The majority of pilot users were female, 10 or more years beyond training, and specialized in primary care. The full range of ambulatory practice time, from 1 to 10 half days of clinic per week, was represented in the pilot population (Table 1), with a mean of 6.2 half days per week. Each physician saw a median of 298 encounters in the prepilot period (IQR 202 to 411) and 338 encounters in the pilot period (IQR 200 to 517) (Figure S3 and Table S1).
Table 1.
Pilot user demographics.
| Characteristic | Counts (Percentages) |
|---|---|
| Overall cohort | 45 (100%) |
| Gender category | |
| Female | 28 (62%) |
| Male | 17 (38%) |
| Nonbinary | 0 (0%) |
| Years after training | |
| 0-4 | 11 (24%) |
| 5-9 | 3 (7%) |
| 10-14 | 14 (31%) |
| ≥15 | 17 (38%) |
| Half days of ambulatory practice per week | |
| 1-2.5 | 4 (9%) |
| >2.5-5 | 11 (24%) |
| >5-7.5 | 12 (27%) |
| >7.5-10 | 18 (40%) |
| Specialty | |
| Primary Care | 25 (56%) |
| Rheumatology | 7 (16%) |
| Otolaryngology | 4 (9%) |
| Cardiology | 3 (7%) |
| Gastroenterology | 3 (7%) |
| Geriatrics | 1 (2%) |
| Ophthalmology | 1 (2%) |
| Orthopedics | 1 (2%) |
Utilization
Of the 17 428 encounters in the analysis period, DAX Copilot was utilized for 9629 encounters (55.25%). At the level of individual physicians, median utilization was 52.5% (IQR 17.86% to 80.97%), with higher utilization of the HPI and A&P SmartSections (median 39.7% and 14.8%, respectively) than the Results and PE SmartSections (median 2.6% and 1.6%, respectively) (Figure 1 and Table S2). There was significant interuser heterogeneity in utilization, with some physicians utilizing DAX for almost all encounters and others having minimal utilization (Figure 1 and Table S3).
Figure 1.
Utilization of DAX Copilot and individual SmartSections across physicians.
EHR time
The majority of physicians spent less time on clinical documentation during the pilot period compared to the prepilot period as measured using Epic Signal (Figure 2), with a median change in time per note of −0.57 minutes (IQR −1.3 to −0.13) (Table S4). This change was statistically significant (T-statistic −5.19, P<.001). The median user average time per note during the prepilot period was 4.86 minutes (IQR 2.89 to 6.47), which dropped to 3.64 minutes in the pilot period (IQR 2.45 to 5.46). To enable additional comparisons below, we also analyzed daily documentation time (Figure 3A and Table S5) and identified a median decrease of −6.89 minutes (IQR −22.37 to −0.65), which was also statistically significant (T-statistic −4.48, P<.001).
Figure 2.

Change in average note time during ambient AI scribe pilot across physicians.
Figure 3.
Change in (A) average daily documentation time, (B) average daily afterhours time, and (C) average daily total EHR time during ambient AI scribe pilot across physicians. Abbreviation: EHR, electronic health record.
There were also statistically significant decreases in daily afterhours EHR time (T-statistic −2.65, P = .01) and daily total EHR time (T-statistic −5.85, P<.001) with median changes of −5.17 minutes (IQR −21.32 to 3.82) and −19.95 minutes (IQR −39.34 to −3.64), respectively (Figure 3B and C, and Tables S6 and S7).
Discussion
In one of the first pilot studies to quantify utilization and time savings across multiple specialties, ambient AI scribes were widely used and associated with reduced documentation and EHR times. These effects were seen in physicians from a variety of clinical specialties, highlighting the broad applicability of this technology. That said, there was notable individual variability in how physicians used the tool, with some preferring to return to preexisting documentation workflows. In terms of time savings, the absolute average time saved was modest, and there were a small number of users who spent more time per note compared to baseline. This heterogeneity implies that certain user phenotypes may derive more utility from current ambient AI technology than others.
Limitations of this study include small sample size, volunteer and selection bias, potential impact from secular trends, inability to analyze time outcomes at the level of individual patient encounters, lack of comparison to alternate strategies, limitations to English-speaking patients only, and predominance of primary care physicians in our cohort. Additionally, EHR use measures have known limitations when quantifying absolute changes.21 In particular, afterhours EHR time and total EHR time were calculated per calendar day without normalization to the number of encounters seen, so they may be underestimating the impact of this technology particularly given the prevalence of clinic half-days in our organization. The use of a 5-second inactivity timer likely also underestimates both total documentation time as well as the absolute magnitude of change in documentation time.
Future opportunities for research in this area include evaluating the quality of notes written with ambient AI scribe assistance particularly with respect to medical synthesis and clinical communication, understanding how physicians review and edit generated text, assessing physicians’ overall perceptions about the costs and benefits of AI scribes, understanding the impact on the patient-physician relationship, and optimizing tool performance and workflows to maximize benefits.
Conclusion
In this pilot study of an LLM-powered ambient AI scribe, there was robust adoption across multiple specialties with modest reductions in documentation and EHR times. These findings highlight the potential of ambient AI scribes to reduce EHR documentation burden and mitigate physician burnout. However, the technology is not yet a one-size-fits-all solution. To reach its full potential, ambient AI scribe technology will need to evolve to effectively support the diverse needs and practices across the medical landscape.
Supplementary Material
Acknowledgments
Individuals from multiple groups contributed to this work: Stanford Technology and Digital Solutions, Stanford Informatics Education, Stanford Healthcare AI Applied Research Team, Microsoft, and Epic Systems. Chris Lieven, BS (Epic Systems) and Kevin Carlyle, BS (Epic Systems) were critical for technical implementation and data reporting. They were not compensated for this work.
Contributor Information
Stephen P Ma, Department of Medicine, Stanford University School of Medicine, Stanford, CA 94305, United States.
April S Liang, Department of Medicine, Stanford University School of Medicine, Stanford, CA 94305, United States.
Shreya J Shah, Department of Medicine, Stanford University School of Medicine, Stanford, CA 94305, United States; Stanford Healthcare AI Applied Research Team, Division of Primary Care and Population Health, Stanford University School of Medicine, Stanford, CA 94305, United States.
Margaret Smith, Stanford Healthcare AI Applied Research Team, Division of Primary Care and Population Health, Stanford University School of Medicine, Stanford, CA 94305, United States.
Yejin Jeong, Stanford Healthcare AI Applied Research Team, Division of Primary Care and Population Health, Stanford University School of Medicine, Stanford, CA 94305, United States.
Anna Devon-Sand, Stanford Healthcare AI Applied Research Team, Division of Primary Care and Population Health, Stanford University School of Medicine, Stanford, CA 94305, United States.
Trevor Crowell, Stanford Healthcare AI Applied Research Team, Division of Primary Care and Population Health, Stanford University School of Medicine, Stanford, CA 94305, United States.
Clarissa Delahaie, Technology and Digital Solutions, Stanford Medicine, Stanford, CA 94305, United States.
Caroline Hsia, Technology and Digital Solutions, Stanford Medicine, Stanford, CA 94305, United States.
Steven Lin, Department of Medicine, Stanford University School of Medicine, Stanford, CA 94305, United States; Stanford Healthcare AI Applied Research Team, Division of Primary Care and Population Health, Stanford University School of Medicine, Stanford, CA 94305, United States.
Tait Shanafelt, Department of Medicine, Stanford University School of Medicine, Stanford, CA 94305, United States; WellMD Center, Stanford University School of Medicine, Stanford, CA 94305, United States.
Michael A Pfeffer, Department of Medicine, Stanford University School of Medicine, Stanford, CA 94305, United States; Technology and Digital Solutions, Stanford Medicine, Stanford, CA 94305, United States.
Christopher Sharp, Department of Medicine, Stanford University School of Medicine, Stanford, CA 94305, United States.
Patricia Garcia, Department of Medicine, Stanford University School of Medicine, Stanford, CA 94305, United States.
Author contributions
All authors contributed to the design of the study, analysis of the results, drafting and revising the article, and the final approval of the submitted version.
Supplementary material
Supplementary material is available at Journal of the American Medical Informatics Association online.
Funding
None declared.
Conflicts of interest
None declared.
Data availability
The data underlying this article will be shared on reasonable request to the corresponding author.
References
- 1. Shanafelt TD, Dyrbye LN, Sinsky C, et al. Relationship between clerical burden and characteristics of the electronic environment with physician burnout and professional satisfaction. Mayo Clin Proc. 2016;91:836-848. [DOI] [PubMed] [Google Scholar]
- 2. Gardner RL, Cooper E, Haskell J, et al. Physician stress and burnout: the impact of health information technology. J Am Med Inform Assoc. 2019;26:106-114. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 3. Apathy NC, Rotenstein L, Bates DW, Holmgren AJ.. Documentation dynamics: note composition, burden, and physician efficiency. Health Serv Res. 2023;58:674-685. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 4. McPeek-Hinz E, Boazak M, Sexton JB, et al. Clinician burnout associated with sex, clinician type, work culture, and use of electronic health records. JAMA Netw Open. 2021;4:e215686. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 5. Tajirian T, Stergiopoulos V, Strudwick G, et al. The influence of electronic health record use on physician burnout: cross-sectional survey. J Med Internet Res. 2020;22:e19274. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 6. Gaffney A, Woolhandler S, Cai C, et al. Medical documentation burden among US office-based physicians in 2019. JAMA Intern Med. 2022;182:564-566. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 7. Li C, Parpia C, Sriharan A, Keefe DT.. Electronic medical record-related burnout in healthcare providers: a scoping review of outcomes and interventions. BMJ Open. 2022;12:e060865. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 8. Arndt BG, Beasley JW, Watkinson MD, et al. Tethered to the EHR: primary care physician workload assessment using EHR event log data and time-motion observations. Ann Fam Med. 2017;15:419-426. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 9. Gidwani R, Nguyen C, Kofoed A, et al. Impact of scribes on physician satisfaction, patient satisfaction, and charting efficiency: a randomized controlled trial. Ann Fam Med. 2017;15:427-433. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 10. Lin S, Duong A, Nguyen C, Teng V.. Five years’ experience with a medical scribe fellowship: shaping future health professions students while addressing provider burnout. Acad Med. 2021;96:671-679. [DOI] [PubMed] [Google Scholar]
- 11. Apathy NC, Holmgren AJ, Cross DA.. Physician EHR time and visit volume following adoption of team-based documentation support. JAMA Intern Med. 2024;184:1212-1221. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 12. Shaw JG, Winget M, Brown-Johnson C, et al. Primary care 2.0: a prospective evaluation of a novel model of advanced team care with expanded medical assistant support. Ann Fam Med. 2021;19:411-418. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 13. Smith PC, Lyon C, English AF, Conry C.. Practice transformation under the university of Colorado’s primary care redesign model. Ann Fam Med. 2019;17:S24-S32. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 14. Johnson M, Lapkin S, Long V, et al. A systematic review of speech recognition technology in health care. BMC Med Inform Decis Mak. 2014;14:94. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 15.CPT® Evaluation and Management (E/M) Office or Other Outpatient (99202-99215) and Prolonged Services (99354, 99355, 99356, 99417) Code and Guideline Changes. American Medical Association; 2021. https://www.ama-assn.org/system/files/2019-06/cpt-office-prolonged-svs-code-changes.pdf
- 16. Arndt BG, Micek MA, Rule A, et al. More tethered to the EHR: EHR workload trends among academic primary care physicians, 2019-2023. Ann Fam Med. 2024;22:12-18. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 17. Wachter RM, Brynjolfsson E.. Will generative artificial intelligence deliver on its promise in health care? JAMA. 2024;331:65-69. [DOI] [PubMed] [Google Scholar]
- 18. Haberle T, Cleveland C, Snow GL, et al. The impact of nuance DAX ambient listening AI documentation: a cohort study. J Am Med Inform Assoc. 2024;31:975-979. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 19. Tierney AA, Gayre G, Hoberman B, et al. Ambient artificial intelligence scribes to alleviate the burden of clinical documentation. NEJM Catal Innov Care Deliv. 2024;5. [Google Scholar]
- 20. Rule A, Hribar MR.. Frequent but fragmented: use of note templates to document outpatient visits at an academic health center. J Am Med Inform Assoc. 2021;29:137-141. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 21. Magon HS, Helkey D, Shanafelt T, Tawfik D.. Creating conversion factors from EHR event log data: a comparison of investigator-derived and vendor-derived metrics for primary care physicians. AMIA Annu Symp Proc. 2023;2023:1115-1124. [PMC free article] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Data Availability Statement
The data underlying this article will be shared on reasonable request to the corresponding author.


