Abstract
Background
Ambient artificial intelligence scribes have become widespread commercial products in the era of generative artificial intelligence. While studies have examined the effect of these tools on the experience of attending physicians, little evidence is available regarding their use by resident physician trainees.
Objectives
To assess trainee experience with an ambient artificial intelligence scribe using measures of usability, acceptability, and documentation burden.
Methods
This prospective observational study enrolled 47 trainees in a 2-month pilot. Pre/postsurveys were conducted with the NASA Task Load Index (NASA-TLX, raw unweighted form, pre/post, for cognitive load during the documentation), the System Usability Scale (post; general usability), the Net Promoter Score (post; acceptability), and the AMIA TrendBurden Survey (pre/post; documentation burden). Electronic health record utilization metrics were obtained from Epic Signal for both the pilot period and a 6-month baseline.
Results
In total, 43/47 (91.5%) of participants adopted the intervention in practice. NASA-TLX scores improved from 56.3 to 43.3 ( p < 0.001), and multiple items on the TrendBurden survey improved with high measures of acceptability. No significant difference in time spent on notes activity per note written was observed, with a median increase of 0.4 minutes ( p = 0.568).
Conclusion
Trainee use of an ambient artificial intelligence scribe was associated with improvements in documentation burden. Additional research on the effect of this technology on trainee learning and expertise development is needed.
Keywords: generative artificial intelligence, artificial intelligence, graduate medical education, electronic health record, health information technology
Background and Significance
Ambient artificial intelligence scribes (henceforth referred to as “ambient AI scribes”) have rapidly become widespread as the use of generative artificial intelligence has increased. 1 Studies of these tools have so far focused mainly on attending physicians or other nontraining clinicians caring for patients primarily in nonteaching roles. 2 3 4 5 6 A gap exists in the literature regarding how these tools will be used by trainees. This study was designed to elicit trainee experiences of cognitive load and documentation burden while using this technology in the ambulatory setting, with a specific focus on resident physicians.
Documentation burden is widely recognized as a concern among physicians. 7 Recent work aims to develop generalizable and rigorous tools to measure documentation burden, such as the American Medical Informatics Association 25 × 5 Initiative. 8 In their recent scoping review, Levy et al found that few studies of interventions aiming to mitigate trainee experiences of documentation burden measured the trainee experience directly. 9 This is important as the documentation burden has been connected to physician burnout. 10 Therefore, we sought to investigate the potential ramifications of ambient AI scribes in the trainee experience, focusing both on self-reported measures of documentation burden as well as measures of electronic health record (EHR) documentation activity obtained from both the ambient AI scribe vendor and our EHR vendor.
Objectives
To assess the experiences of trainees introduced to an ambient AI scribe in ambulatory practice, by measuring outcomes including usability, acceptability, and documentation burden.
Methods
Study Design and Setting
This prospective observational study was a pilot evaluation conducted from November to December 2024 at two urban academic medical centers. During the pilot, ambient AI technology (Abridge, Inc., Pittsburgh, Pennsylvania) was introduced into routine practice for a cohort of postgraduate medical trainees (resident physicians). The AI scribe had previously been deployed for attending physician use at both sites before the study period.
Ethical Approval
The study was determined to be exempt from review by the Yale University Institutional Review Board (approval no.: HIC 2000038118) before participant recruitment.
Recruitment
The Graduate Medical Education office invited participation by residency programs at both sites in the pilot. Participating residency programs nominated individual residents at postgraduate year (PGY) two or higher who were subsequently invited to enroll. All invited residents were required to be on at least one 2-week ambulatory block during the pilot period. Trainees were only permitted to use the technology in the ambulatory or emergency department setting based on institutional policies at the time of this study.
Survey Design and Data Collection
A survey was developed and distributed via the Research Electronic Data Capture (REDCap) platform, which captured demographic information and standardized assessments of baseline task load and documentation burden. 11 12 Instruments administered included the NASA Task Load Index (NASA-TLX, raw unweighted form, pre/post, for cognitive load), the System Usability Scale (post; general usability), the Net Promoter Score (post; acceptability), and the AMIA TrendBurden Survey (pre/post; documentation burden). 13 14 15 16 17 Subsequently, residents who completed the presurvey were provisioned with access to the tool and received the same educational resources provided to onboarding attending physicians, including an introductory video, policies for use, and tip sheets.
AI scribe utilization metrics, including the number of notes recorded and total minutes recorded, were obtained from the vendor. EHR utilization metrics, including ambulatory note-related activity, were collected from Epic Signal (Epic Systems, Verona, Wisconsin, United States). Time in notes per day was defined as the number of minutes the participant was active in the notes activity, divided by the number of days they logged into the system. Time in notes per note was defined as the number of minutes the participant was active in the notes activity, divided by the number of notes written. Baseline EHR activity metrics were derived from a 6-month prepilot period, restricted to months in which the trainee had at least one ambulatory encounter. Pilot-period metrics were averaged across the 2 months of the pilot. Participants were excluded from the EHR utilization metrics analysis if they did not have at least 1 month of baseline data or if their practice was only in the emergency department, as the Signal metrics available for the ambulatory environment and emergency department use different definitions related to note activity and could not be reliably compared.
Statistical Analysis
All statistical analysis was conducted in R (v4.2.2). 18 Participant characteristics were reported using descriptive statistics. In paired analyses, normality of distributions was assessed visually and with the Shapiro–Wilk test to guide use of either the paired t -test or the Wilcoxon signed rank test. Two-sided tests were used unless otherwise specified, with a statistical significance threshold of p < 0.05. A Bonferroni correction was included for multiple comparisons when evaluating the TrendBurden instrument. Visualizations were generated in ggplot2. 19
Results
In total, 55 trainees were selected by their respective residency program directors to be invited to participate in the pilot. Of these, 47 (85.5%) responded to the presurvey and were provisioned with access to the ambient AI scribe. Participant demographics are summarized in Table 1 .
Table 1. Participant trainee demographics ( n = 47) .
| Completed both pre- and postsurveys ( n = 37) | Completed only presurvey ( n = 10) | Total ( n = 47) | |
|---|---|---|---|
| Age (y) | |||
| 25–34 | 33 (89.2%) | 10 (100%) | 43 (91.5%) |
| 35–44 | 4 (10.8%) | 0 (0%) | 4 (8.5%) |
| Gender | |||
| Woman | 21 (56.8%) | 6 (60.0%) | 27 (57.4%) |
| Man | 16 (43.2%) | 3 (30.0%) | 19 (40.4%) |
| Other gender identity | 0 (0%) | 0 (0%) | 0 (0%) |
| Prefer not to answer | 0 (0%) | 1 (10.0%) | 1 (2.1%) |
| PGY level | |||
| 2 | 15 (40.5%) | 5 (50.0%) | 20 (42.6%) |
| 3 | 18 (48.6%) | 4 (40.0%) | 22 (46.8%) |
| 4 | 4 (10.8%) | 1 (10.0%) | 5 (10.6%) |
| Training specialty | |||
| Internal medicine | 24 (64.9%) | 4 (40.0%) | 28 (59.6%) |
| Emergency medicine | 6 (16.2%) | 0 (0%) | 6 (12.8%) |
| OB/GYN | 3 (8.1%) | 5 (50.0%) | 8 (17.0%) |
| Pediatrics | 4 (10.8%) | 1 (10.0%) | 5 (10.6%) |
| Tool utilization mean (SD) | |||
| Total notes generated | 26.6 (17.6) | 7.00 (8.16) | 23.3 (17.9) |
| Total minutes recorded | 491 (334) | 159 (157) | 449 (355) |
Forty-three of the 47 trainees provisioned with access to the tool (91.5%) completed at least one AI-generated note during the study period. Participants who used the tool recorded a mean of 440 minutes of patient encounters during the encounter and completed a mean of 23.0 notes. Distributions of tool utilization are shown in Fig. 1 . In total, 37/47 (78.7%) of participants completed both the pre- and postsurveys and had paired data available to analyze effects on task load and other self-report metrics. Signal data were only available for 31/47 (66.0%) of participants and were not available for any emergency medicine trainees.
Fig. 1.
Upper panels: trainee utilization of ambient AI scribe during pilot period, provided by vendor ( n = 43). Lower panels: electronic health record note activity time, baseline and pilot, collected from Epic Signal ( n = 31). All shown stratified by resident physician specialty.
Ambulatory note activity time metrics are reported from Epic Signal for trainees participating in the pilot ( Fig. 1 ). Participants spent less time on notes per day in the pilot period, median difference 10.1 minutes, 95% CI: 5.9 to 15.2 minutes, p < 0.001, Wilcoxon's text. No difference was noted in time in notes per note, median difference 0.4 minutes, 95% CI: −2.2 to 0.8 minutes, p = 0.568, Wilcoxon's text. Participants wrote fewer notes during the pilot, with a mean of 71.0, 95% CI: 57.5 to 84.5 notes per month at baseline compared to a mean of 44.9, 95% CI: 32.4 to 57.4 notes per month during the pilot. Additional Signal-reported metrics, such as time outside scheduled hours, time on unscheduled days, and pajama time, were also evaluated; however, they were compromised by a high number of outliers likely related to the scheduling variability among trainees and inconsistency of documented clinic assignments.
In the unweighted NASA-TLX, task load related to documentation tasks significantly improved from a mean of 56.3 (SD: 11.0) to 43.3 (SD: 18.1), p < 0.001, paired t-test. Additionally, changes in individual items on the TrendBurden Survey were noted, as shown in Fig. 2 , with improvements in measures of burden noted in item 1, p = 0.031, Wilcoxon (the amount of time and effort I spend documenting patient care is appropriate) and item 2, p = 0.001, Wilcoxon (I finish work later than desired or need to do work at home because of excessive documentation tasks), though only the latter would be significant if adjusted by the Bonferroni correction All remaining survey items were not statistically different at p < 0.05. As of the time of this publication, an aggregate measure for this assessment has not been validated.
Fig. 2.
TrendBurden documentation burden survey responses, pre/post ( n = 37).
In the postsurvey, the mean System Usability Scale was 78.8 (SD: 13.1), which is generally interpreted as good to excellent. 14 20 The overall Net Promoter Score was 51.4, reflecting a high user acceptability of the technology relative to common benchmarks. 15 Visualizations of these responses are included in Supplementary Figs. S1 to S3 (available in the online version only).
Discussion
In this novel study among trainees, participants used an ambient AI scribe, and metrics related to usability, acceptability, and documentation burden were measured. We observed lower TLX and perceived documentation burden following implementation, with good to excellent usability and acceptability. Additionally, the overall time spent on the notes activity per day was lower during the pilot, though a difference in time per note completed was not observed. This may have been mediated by a reduced number of notes performed during the pilot period due to other temporal trends, such as holiday scheduling, which coincided with part of the pilot, limiting the interpretability of this finding.
This work builds on several recent studies by other health systems implementing different ambient AI scribes among their attending physicians. The differences in perceived task load and documentation burden in this study were similar to those reported among attending physicians by Shah et al in their pilot implementation study among 48 ambulatory attending physicians. 2 Similarly, our decrease in note activity time was similar to that seen by Ma et al in their prospective quality improvement study of 45 ambulatory attending physicians, though we were not able to reliably measure some of their other encounter-level findings due to differences in trainee scheduling and documentation patterns. 3 High attending physician acceptability was also noted by Haberle et al in their matched prospective cohort study of 99 independently practicing ambulatory providers. 4 In a mixed methods study of 46 independently practicing ambulatory clinicians, Duggan et al also saw similar improvements in documentation efficiency, documentation task burden, and self-reported engagement with patients during the clinical encounters. 5
While our study demonstrates similar measures of task load reduction, acceptability, and effects on documentation burden relative to studies on attending physicians, significant unanswered questions about how these technologies will affect trainees remain. The trainee experience in the clinic setting differs fundamentally from that of independently practicing physicians, as trainees are still learning their specialty, establishing their practice patterns, and developing clinical expertise—all while being actively supervised and precepting patients alongside an attending physician. Similarly, the patient encounter has a dual aim for these trainees: they are providing clinical care while also functioning as adult learners in an academic setting. The duality of roles poses unique opportunities and hazards. On one hand, a reduction in documentation burden could facilitate increased engagement with patients as well as other educational opportunities and resources. This was noted anecdotally in free-text comments received from participants. However, it is also possible that the scribe could serve as a crutch. Expertise development in the clinical setting includes synthesizing clinical data. In its default operating mode, the AI scribe software used in our study generates medical decision-making text that may circumvent the formative diagnostic process for trainees, for example, translating descriptors of disease processes into more explicit diagnoses. However, the scribe does not generally document any treatment decisions not explicitly verbally articulated by the recording provider. For this reason, individual GME programs will need to determine how best to educate and supervise trainees in using these tools to optimize safety, service, and clinical skill development. This is an area that would particularly benefit from further qualitative investigation, as high-quality evidence is limited, and learners may interact differently with the technology.
Limitations of this study include its short duration, the limited number of residency programs involved, and nonrandom selection of pilot participants by their program directors. The design did not control for seasonality or temporal trends, including potential improvements related to general skills development unrelated to the intervention. Differential nonresponse was noted on the postsurvey, with a particularly low response rate among obstetrics and gynecology trainees and those with lower utilization of the tool. Additionally, as we were not able to capture EHR use metrics for some participants, there may be differential effects on documentation practices in settings such as the emergency department that were not observed. As this was a subset of trainees who were intentionally selected by their program directors for inclusion, it is possible that they may have been more eager to adopt new technologies than their peers. They were also aware that their EHR activity was being monitored during the pilot, which may have altered their documentation behaviors. Longitudinal data collection after wider implementation, which is currently ongoing at the institutions involved, may help address some of these limitations. This study also did not explore the factuality of note content generated by the AI scribes, including the presence of any hallucinations.
Conclusion
Trainee use of an ambient AI scribe in the ambulatory setting was associated with improvements in experienced documentation burden and high usability/acceptability. EHR use metrics were difficult to interpret due to contemporaneous seasonal changes in clinical activity. Additional research on the effect of these tools on the educational development of trainees is needed.
Clinical Relevance Statement
As ambient AI digital scribes are becoming increasingly widespread in clinical practice, understanding their effects on specific groups of practitioners is important. Physician trainees make up a significant proportion of the workforce at academic medical centers, and as such, understanding the effect of these tools on their documentation burden is important to efforts to address excess documentation burden and optimize learning.
Multiple-Choice Questions
-
What was the NASA-Task Load index designed to assess?
The weight of all objects required to perform a physical task.
The subjective difficulty of a task.
The financial cost of a task within an organizational framework.
The number of people needed to perform the task.
Correct Answer: The correct answer is option b. The subjective difficulty of a task. The NASA-Task Load Index (NASA-TLX) is a human factors research tool designed to measure the perceived workload associated with a task. The tool captures multiple aspects of the difficulty of a task, including the mental, physical, and temporal demands on the individual performing the task. The other responses are incorrect as the tool was not designed to measure only the physical difficulty of a task, does not measure cost, and does not necessarily predict the number of people needed for a task.
-
What is the current purpose of an ambient AI digital scribe in clinical practice?
To review clinician notes and detect potential errors.
To call patients, schedule appointments, and answer medical questions.
To perform medication reconciliation prior to an encounter.
To transcribe an encounter and generate a clinical note.
Correct Answer: The correct answer is option d. To transcribe an encounter and generate a clinical note. Ambient AI digital scribes typically function by recording audio from the interaction between a clinician and a patient. This audio is then transcribed to text and passed to a large language model, which generates a clinical note. While the other tasks may be performed by different artificial intelligence tools, none of the other described functions are consistent with the task performed by a scribe.
Acknowledgment
The authors would like to thank the graduate medical education office and the residency program directors who granted us access to resident physician pilot testers.
Funding Statement
Funding None.
Conflict of Interest D.S.W. reported receiving funding from the National Institutes of Health for unrelated work. E.R.M. reports receiving grants from the National Institute on Drug Abuse, the Agency for Healthcare Research and Quality, and the American Medical Association, and stock options for advising Iolite Health, Inc., unrelated to this work. The content is solely the responsibility of the authors and does not necessarily represent the official views of the National Institutes of Health, the Agency for Healthcare Research and Quality, the American Medical Association, and Iolite. L.E.S. reported participating in a volunteer AI research advisory committee with Abridge, Inc., and consulting to Genentech and Medtronic unrelated to this work. S.Y.O. reported equity in Avo, Inc. The contents of this manuscript represent the view of the authors and do not necessarily reflect the position or policy of the U.S. Department of Veterans Affairs or the United States Government.
Protection of Human and Animal Subjects
The study was determined to be exempt from review by the Yale University Institutional Review Board (approval no.: HIC 2000038118) before participant recruitment.
Supplementary Material
References
- 1.Seth P, Carretas R, Rudzicz F. The utility and implications of ambient scribes in primary care. JMIR AI. 2024;3:e57673. doi: 10.2196/57673. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 2.Shah S J, Devon-Sand A, Ma S P et al. Ambient artificial intelligence scribes: physician burnout and perspectives on usability and documentation burden. J Am Med Inform Assoc. 2025;32(02):375–380. doi: 10.1093/jamia/ocae295. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 3.Ma S P, Liang A S, Shah S J et al. Ambient artificial intelligence scribes: utilization and impact on documentation time. J Am Med Inform Assoc. 2025;32(02):381–385. doi: 10.1093/jamia/ocae304. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 4.Haberle T, Cleveland C, Snow G L et al. The impact of nuance DAX ambient listening AI documentation: a cohort study. J Am Med Inform Assoc. 2024;31(04):975–979. doi: 10.1093/jamia/ocae022. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 5.Duggan M J, Gervase J, Schoenbaum A et al. Clinician experiences with ambient scribe technology to assist with documentation burden and efficiency. JAMA Netw Open. 2025;8(02):e2460637. doi: 10.1001/jamanetworkopen.2024.60637. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 6.Hassan H, Zipursky A R, Rabbani N et al. Special topic on burnout: clinical implementation of artificial intelligence scribes in healthcare: a systematic review. Appl Clin Inform. 2025 doi: 10.1055/a-2597-2017. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 7.Sloss E A, Abdul S, Aboagyewah M A et al. Toward alleviating clinician documentation burden: a scoping review of burden reduction efforts. Appl Clin Inform. 2024;15(03):446–455. doi: 10.1055/s-0044-1787007. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 8.Hobensack M, Levy D R, Cato K et al. 25 × 5 symposium to reduce documentation burden: report-out and call for action. Appl Clin Inform. 2022;13(02):439–446. doi: 10.1055/s-0042-1746169. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 9.Levy D R, Rossetti S C, Brandt C A et al. Interventions to mitigate EHR and documentation burden in health professions trainees: a scoping review. Appl Clin Inform. 2025;16(01):111–127. doi: 10.1055/a-2434-5177. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 10.Melnick E R, Harry E, Sinsky C A et al. Perceived electronic health record usability as a predictor of task load and burnout among US physicians: mediation analysis. J Med Internet Res. 2020;22(12):e23382. doi: 10.2196/23382. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 11.REDCap Consortium . Harris P A, Taylor R, Minor B L et al. The REDCap consortium: building an international community of software platform partners. J Biomed Inform. 2019;95:103208. doi: 10.1016/j.jbi.2019.103208. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 12.Harris P A, Taylor R, Thielke R, Payne J, Gonzalez N, Conde J G. Research electronic data capture (REDCap)–a metadata-driven methodology and workflow process for providing translational research informatics support. J Biomed Inform. 2009;42(02):377–381. doi: 10.1016/j.jbi.2008.08.010. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 13.Hart S G. Nasa-task load index (NASA-TLX); 20 years later. Proc Hum Factors Ergon Soc Annu Meet. 2006;50(09):904–908. [Google Scholar]
- 14.Bangor A, Kortum P T, Miller J T. An empirical evaluation of the system usability scale. Int J Hum Comput Interact. 2008;24(06):574–594. [Google Scholar]
- 15.Adams C, Walpola R, Schembri A M, Harrison R. The ultimate question? Evaluating the use of net promoter score in healthcare: a systematic review. Health Expect. 2022;25(05):2328–2339. doi: 10.1111/hex.13577. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 16.TrendBurden: Pulse Survey on Excessive Documentation Burden for Health Professionals | AMIA - American Medical Informatics AssociationAccessed February 13, 2025 at:https://amia.org/about-amia/amia-25x5/trendburden-pulse-survey
- 17.Levy D R, Withall J B, Mishuris R G et al. Defining documentation burden (DocBurden) and excessive DocBurden for all health professionals: a scoping review. Appl Clin Inform. 2024;15(05):898–913. doi: 10.1055/a-2385-1654. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 18.R Core Team R: A Language and Environment for Statistical ComputingPublished online in 2024. Accessed July 4, 2025 at:https://www.R-project.org/
- 19.Wickham H, Chang W, Henry Let al. ggplot2: Create Elegant Data Visualisations Using the Grammar of GraphicsPublished online April 23, 2024. Accessed February 13, 2025 at:https://cran.r-project.org/web/packages/ggplot2/index.html
- 20.Bangor A, Kortum P, Miller J. Determining what individual SUS scores mean: adding an adjective rating scale. J Usability Stud. 2009;4(03):114–234. [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.


