Skip to main content
Applied Clinical Informatics logoLink to Applied Clinical Informatics
. 2025 Sep 11;16(4):1041–1052. doi: 10.1055/a-2657-8087

The Influence of Artificial Intelligence Scribes on Clinician Experience and Efficiency among Pediatric Subspecialists: A Rapid, Randomized Quality Improvement Trial

H Stella Shin 1,2,3, Herb Williams 2, Nikolay Braykov 2, Afrin Jahan 2, Jeremy Meller 2, Evan W Orenstein 1,2,4,
PMCID: PMC12425613  PMID: 40675605

Abstract

Background

Artificial intelligence (AI) scribes may reduce the documentation burden and improve clinician experience through generative AI automatically producing provider note sections from recordings of patient–provider encounters.

Objective

We aimed to examine the impact of AI scribes on clinician experience, clinician efficiency, and business efficiency measures among pediatric subspecialty physicians.

Methods

We randomized pediatric subspecialty providers with ≥0.5 clinical full-time equivalent and stable electronic health record (EHR) log metrics to use Microsoft/Nuance Digital Ambient eXperience (DAX) Copilot from May 1, 2024, to July 31, 2024 (intervention group) or controls. Using difference-in-differences, we compared quantitative measures of subjective clinician experience using the KLAS Net EHR Experience survey, objective measures of clinician efficiency from EHR logs (e.g., pajama time), and business efficiency measures. At-the-elbow support checked in with intervention providers approximately weekly, and we assessed the sentiment of qualitative comments.

Results

Twelve providers were randomized to the intervention and 11 to the control group. One intervention provider stopped using DAX due to ineffectiveness. In the intervention group, DAX was used to populate one or more characters in 53% of visit notes (range across providers: 10.6–98.2%). Nine intervention and eight control providers completed pre- and postsurveys. KLAS Net EHR Experience improved among intervention providers from 52.6 (70th percentile) to 75.2 (99th percentile) but dropped from 37.3 (38th percentile) to 30 (14th percentile) among control providers. Experiencing burnout dropped from 8 (89%) to 5 (56%) among intervention providers but remained stable at 3 (38%) in the control group. There was no significant change to pajama time (−9.4 minutes per scheduled day, 95% CI: −41.2 to +22.4), time in notes per encounter (+0.2 minutes per note, 95% CI: −6.6 to +6.9), or work Relative Value Units (wRVUs) per encounter (−0.03, 95% CI: −0.5 to +0.44). Of 48 qualitative comments, 69% had a positive sentiment, 15% neutral, and 17% negative.

Conclusion

Among pediatric subspecialists, AI scribes improved clinician experience and burnout without changing charting time or EHR work outside work hours.

Keywords: electronic health records, burnout, artificial intelligence, quality improvement, randomized controlled trial

Background and Significance

Physician burnout is a long-term stress reaction characterized by emotional exhaustion, depersonalization, and decreased personal achievement. 1 In 2023, 48% of physicians reported experiencing at least one symptom of burnout, and the estimated cost of burnout to the U.S. health care system ranges from $2 to 6 billion per year. 2 3 Time spent documenting in electronic health records (EHRs) outside of work hours has been associated with burnout, 4 5 suggesting that solutions to reduce documentation burden 6 could alleviate burnout. For example, in-person and virtual medical scribes have demonstrated reductions in both documentation burden and burnout. 7 8 9 10 However, medical scribes remain expensive with limited adoption, especially in lower-income specialties. 11

Artificial intelligence (AI) scribes, also known as ambient documentation, use recordings of patient visits and AI algorithms to automatically generate transcripts and physician note elements, such as history of present illness (HPI), physical exam, and assessment and plan (A&P) sections. AI scribes provide several potential advantages over other solutions, including near-instantaneous availability after the visit is recorded, sectioning of note elements that allow providers to use their existing note templates, transcripts for verification of AI-generated note elements to improve provider trust, and scalability. Many vendors now offer AI scribe products with different levels of EHR integration, customization, and performance in different settings.

As this new technology has rolled out, many institutions have implemented pilot programs and shared evaluations showing improvement in burnout scores, decreased documentation time, and some suggesting business efficiency improvements. 12 13 14 15 However, most of these studies were performed as postimplementation surveys, 14 pre–post examinations without a control group, 12 13 14 15 16 17 18 19 20 or as non-randomized quasi-experimental studies with volunteer physicians. 21 22 23 To our knowledge, no randomized trials of AI scribes have been published to date. Additionally, many prior reports have been shared primarily as press releases or conference presentations, 12 13 14 15 16 17 lacking the strength of peer review in published studies. Finally, few studies have included pediatric physicians or pediatric subspecialists. 14 18

Rapid, randomized quality improvement (QI) trials aim to bring together the causal inference power of randomized controlled trials with the practical, operational, and speed requirements of QI. 24 25 While generally not used for higher-risk studies (e.g., new biologic therapies or procedures), they are a powerful approach to study how health care services are delivered, allowing health systems to have greater confidence in the utility of interventions being proposed for QI when compared with pre/post or other evaluation methods. Examples include texting patients for vaccine reminders, encouraging providers to engage in tobacco cessation counseling, improving uptake of preventive care, and many others. 26 27 28 29 30 31 32 In this single-center rapid, randomized QI trial, we examined the impact of AI scribes on clinician experience, clinician efficiency, and business efficiency measures among pediatric subspecialty physicians to determine whether to adopt, adapt, or abandon the strategy for physician burnout reduction and patient access.

Methods

Setting

This work was done at Children's Healthcare of Atlanta (CHOA), a tertiary care academic pediatric health system in the Southeastern United States with a single enterprise implementation of Epic Systems as its EHR. The intervention period was from May 1, 2024, to July 31, 2024. EHR logs were compared with the 12 months prior to the intervention period, from May 1, 2023, through April 30, 2024. The CHOA Institutional Review Board deemed this work non-human subjects research (approval no.: STUDY00001962) as a QI project aimed at reducing physician burnout, rather than a project whose primary goal was to create generalizable knowledge.

Provider Cohort

We recruited pediatrics subspecialty physicians who met the following inclusion criteria.

  1. At least 0.5 clinical full-time equivalent (FTE).

  2. Stable Epic Signal© (a vendor-provided tool for provider efficiency analytics) measures of pajama time (see definition in “Data Collection and Outcome Metrics” section) 33 and time in notes per encounter during the baseline period as judged by visual inspection by two physician informaticists (H.S.S. and E.W.O.) with disagreements resolved by consensus. “Stability” was defined as relatively constant values during the baseline period.

  3. Opportunity for improvement in pajama time and time in notes per encounter relative to specialty-specific and peer benchmarks, as judged by two physician informaticists (H.S.S. and E.W.O.) with disagreements resolved by consensus. “Opportunity for improvement” was based on comparison to peers in the same specialty.

  4. Had Epic Haiku (a provider-facing mobile application containing a subset of EHR features) on their phone (or willing to install it).

  5. Division Chief approval for the individual provider to participate in this pilot study.

  6. Verbal consent of the provider.

Providers were recruited across seven pediatric specialties for which the AI scribe vendor, Microsoft/Nuance Digital Ambient eXperience (DAX) Copilot had a product at the time. For each specialty, two to four providers meeting the criteria above were identified and then recruited for randomization. Those that completed baseline surveys were then randomized within specialties to intervention or control ( Fig. 1 ). Of note, in some specialties, an odd number (1 or 3) of providers completed baseline surveys—in this case, we randomized a larger fraction (1 or 2) to the intervention group to use up all budgeted licenses. However, this approach led to some imbalances between the intervention and control groups by specialty ( Table 1 ).

Fig. 1.

Fig. 1

Study design for the rapid, randomized quality improvement trial. AI, artificial intelligence.

Table 1. Characteristics of control and intervention participants.

Control Intervention
Female (%) 6/11 (54.5%) 9/12 (75%)
Specialty ( n )
 Endocrinology
 Gastroenterology
 Nephrology
 Neurology
 Rheumatology
 Orthopedics
 Pulmonology
 Psychiatry a
2
2
0
2
2
0
2
1
2
2
1
2
1
1
2
1
Years posttraining (mean) 15.7 10.3
Clinical FTE (mean) 0.90 0.78

Abbreviation: FTE, full-time equivalent.

a

One Behavioral Health provider was enrolled in Microsoft/Nuance Digital Ambient eXperience (DAX) Copilot, but not used beyond the training sessions.

Intervention

DAX Copilot integrated into Epic Haiku was launched for all intervention providers on May 1, 2024, with the generally available model at the time. DAX Copilot was chosen as the initial vendor given the health system's prior relationship with Nuance through their Dragon Medical One system. Of note, on June 11, 2024, the model was updated per DAX Copilot's early access program. The planned intervention period was set for 3 months (ending July 31, 2024), although intervention providers were allowed to continue using DAX Copilot if they wanted. Written consent was required for each patient/family prior to beginning recording. Training was focused on DAX Copilot use (as opposed to general EHR proficiency) and included the following:

  • A meeting between a physician informaticist (H.S.S.) and the user to walk through the mechanics of how to use DAX Copilot prior to going live.

  • Sharing of training videos from the AI scribe vendor (DAX Copilot).

  • At least one visit by CHOA's support-on-site team (at-the-elbow trainers) during the first 2 weeks of use.

  • Shadowing of a convenience sample of four providers by the Microsoft/Nuance DAX training team.

In addition, CHOA's support-on-site team continued to provide ongoing visits every 2 to 4 weeks for each provider and asked if the app worked, how accurate the provider perceived the transcription to be, challenges around consent, and perceived efficiency gains from using the Copilot.

Data Collection and Outcome Metrics

We compared outcomes in five domains, including use and engagement, scribe performance, clinician experience, clinician efficiency, and business efficiency ( Table 2 ). Our primary outcomes were (1) Net EHR Experience and Experiencing Burnout according to the KLAS EHR Experience survey, 34 (2) pajama time (average minutes a provider spent in charting activities on weekdays outside working hours (7 a.m. to 5:30 p.m.) or outside scheduled hours on weekends or non-scheduled holidays) and Time in Notes per Appointment according to Epic Systems© Signal data, 33 and (3) outpatient adjusted work Relative Value Units (wRVUs) per encounter according to an existing local revenue cycle dashboard.

Table 2. Measurement framework for artificial intelligence scribe evaluation.

Measurement concept Metric Source
Use and engagement Percentage of eligible encounters in which an AI scribe was used and contributed at least one character to the note. Numerator: DAX CopilotDenominator: Epic Signal
Scribe performance Percentage of manual note characters Epic Signal
Clinician experience Net EHR Experience Score a KLAS EHR Experience Survey
Clinical practice enhanced through EHR use KLAS EHR Experience Survey
50% or less ambulatory chart closed same day (provider perception) KLAS EHR Experience Survey
6+ hours of at-home charting per week (provider perception) KLAS EHR Experience Survey
Experiencing burnout a KLAS EHR Experience Survey
Likely to leave organization within 2 years KLAS EHR Experience Survey
Clinician efficiency Pajama time a Epic Signal
Time in notes per appointment a Epic Signal
Business Efficiency Outpatient adjusted wRVUs per encounter a Local Revenue Cycle Dashboard
Percentage of appointments closed same day Epic Signal

Abbreviations: AI, artificial intelligence; EHR, electronic health record; DAX, Microsoft/Nuance Digital Ambient eXperience.

a

Primary outcomes.

Survey data were collected among control and intervention participants prior to the start of the treatment period (May 1, 2024) and again after 3 months of DAX Copilot use (July 31, 2024). Epic Signal data and wRVU data were collected over a 12-month baseline period from May 2023 to April 2024 and through the treatment period from May 2024 to July 2024.

Analysis

We performed a Difference-in-Differences (D-i-D) analysis to compare outcomes between intervention and control groups. Briefly, the D-i-D design accounts for secular trends in outcomes of interest by comparing how an outcome changes in the intervention group during the treatment period with that same change in the control group. Thus, the difference pre/post in the intervention group is subtracted from the difference pre/post in the control group to get the crude D-i-D estimation, which can be adjusted through regression-based methods. 35 For survey-based metrics, we compared only the crude D-i-D for outcomes. For Epic Signal data-based metrics and wRVUs per encounter, we compared summary metrics during the baseline and treatment periods for both control and intervention groups and visually inspected the difference via run charts. For our primary outcomes, we calculated 95% confidence intervals using linear regression with variables including the intervention group, treatment period, and interaction between intervention group and treatment period, where the coefficient for the interaction term was used as the D-i-D estimate. For non-primary outcomes, we calculated p -values comparing the two groups using two-sample t -tests, considering p  < 0.05 as statistically significant. Analyses were performed in R Statistical Programming software. 36

In addition to our quantitative analyses, we analyzed comments elicited from our support-on-site (at-the-elbow) training team from visits to intervention providers. These visits occurred approximately once per month for each provider but were oriented around operational constraints. At these visits, the support-on-site team would ask about any technical issues using the software, any experiences to share, and whether the provider felt the platform made them faster or more efficient. The support-on-site team would email a summary after each visit to the two physician informaticists (H.S.S. and E.W.O.) who classified comment sentiment as positive, negative, or neutral in consensus sessions, and example comments were also extracted.

Results

The AI scribe was implemented with 12 intervention providers, including 10 pediatric medical subspecialists, 1 pediatric behavioral health provider, and 1 pediatric surgical subspecialist (orthopaedics). Results were compared with 11 control providers in similar specialties ( Table 1 ).

Use and Engagement

Of the 12 intervention providers, 1 stopped using DAX Copilot because the notes were consistently not helpful despite training and observation from both CHOA's support-on-site team and review with the vendor. This provider cared for medically complex patients who also had substantial behavioral health challenges, with many appointments lasting over an hour and containing a myriad of medical and psychological discussions. In discussion with the provider, vendor, and support-on-site team, it was felt that the model struggled to organize as a medical subspecialty note or as a behavioral health note, leading to variable note styles inconsistent with the approach of that provider. Of the 11 providers who continued using DAX Copilot through the treatment period, adoption ranged from 10.6 to 98.2% use for eligible encounters ( Table 3 ). In qualitative comments, providers noted the AI scribe was generally more useful to them for new patients than established patients (where providers often would copy-forward old notes).

Table 3. Adoption of artificial intelligence scribes by participant ( n  = 11 a ) .

 User  Total appointments during treatment period  Appointments with AI scribe use  Usage rate (%)
 A  216  23  10.6
 B  274  49  17.9
 C  83  15  18.1
 D  162  58  35.8
 E  68  30  44.1
 F  241  146  60.6
 G  121  80  66.1
 H  222  149  67.1
 I  370  268  72.4
 J  136  112  82.4
 K  171  168  98.2

Abbreviation: AI, artificial intelligence.

a

While 12 providers were initially in the intervention group, 1 Behavioral Health provider was enrolled in Microsoft/Nuance Digital Ambient eXperience (DAX) Copilot, but did not use it beyond the training sessions. That provider is excluded from this table.

Scribe Performance

While we did not have measures of scribe “correctness” or specific edits to the note sections added by the AI scribe, we estimated scribe performance using the proportion of the note comprising manually entered characters (i.e., typed by the provider as opposed to templated text, links from other portions of the chart, dot phrases, copy-forwarded sections, or text added by the AI scribe). This proxy measure should increase if the AI scribe content was incorrect (requiring manual editing by the provider) or if it was insufficient (e.g., the provider had to manually type their own A&P). The proportion of manually entered characters decreased in the intervention group from 11.8% during the baseline period to 8.9% in the treatment period ( p  = 0.04). In the control group, the manual note character percentage was not significantly changed (8.9% baseline to 9.4% during the treatment period, p  = 0.42). The drop in manual note characters in the intervention group was temporally associated with the time of the intervention ( Fig. 2 ).

Fig. 2.

Fig. 2

Proportion of ambulatory progress notes comprising manually entered characters over time for intervention and control participants.

Clinician Experience

Among intervention participants, 9 of 12 completed the KLAS EHR Experience Survey before and after the treatment period, while 8 of 11 control participants completed both surveys ( Table 4 ). The Net EHR Experience score improved in the intervention group from 52.6 (70th percentile for physicians across 131 other health systems using Epic Systems) to 75.2 (99th percentile), while the score decreased from 45.1 (58th percentile) to 37.6 (38th percentile) in the control group. The proportion of providers stating they were experiencing burnout symptoms dropped from 89% (8/9) to 56% (5/9) in the intervention group, while it remained constant at 38% (3/8) in the control group. Perceived difficulty closing ambulatory charts the same day increased slightly in the intervention group (63–86%) while decreasing slightly in the control group (88–63%). Perceived time spent at home charting was unchanged in both groups. At baseline, the intervention group had two providers who answered they were likely to leave the organization within 2 years, which decreased to only one after the treatment period. By contrast, the control group had no providers at baseline or after the treatment period who answered they were likely to leave within 2 years.

Table 4. KLAS EHR Experience survey scores pre- and posttreatment period for control and intervention groups.

Control( n  = 8 of 11 submitted pre- and postsurveys) Intervention( n  = 9 of 12 submitted pre- and postsurveys)
Pre Percentile compared with other Epic Orgs ( n  = 131) Post Percentile compared with other Epic Orgs ( n  = 131) Pre Percentile compared with other Epic Orgs ( n  = 131) Post Percentile compared with other Epic Orgs ( n  = 131)
Net EHR experience score 45.1 58 37.6 38 52.6 70 75.2 99
Experiencing burnout 3 (38%) 33 3 (38%) 33 8 (89%) 0 5 (56%) 1
Clinical practice enhanced through EHR use 4 (50%) 11 3 (38%) 3 6 (67%) 72 7 (78%) 93
50% or less ambulatory charts closed same day 7 (88%) 0 5 (63%) 6 5 (63%) 6 6 (86%) 0
6+ hours of at-home charting 7 (88%) 1 7 (88%) 1 7 (78%) 2 7 (78%) 2
Likely to leave org within 2 years 0 99 0 99 2 (22%) 20 1 (11%) 92

Abbreviation: EHR, electronic health record.

Clinician Efficiency

Pajama time, a measure of work outside of work hours from Epic Signal, was unchanged in the intervention group (47.2 minutes at baseline, 47.4 minutes during the treatment period, p  = 0.97) and the control group (75.5 minutes at baseline, 92 minutes during the treatment period, p  = 0.32). In a D-i-D analysis, the estimated treatment effect of the AI scribe was a reduction in pajama time of 9.4 minutes per day (95% CI: −41.2 to +22.4, p  = 0.56). On visual inspection of a run chart of pajama time, while the intervention group is markedly lower at baseline and during the treatment period, there is no temporal association of improvement with the AI scribe intervention ( Fig. 3 ). Time in notes per appointment ( Fig. 4 ) was grossly unchanged in both the intervention group (12.1 minutes at baseline, 11.6 minutes in the treatment period, p  = 0.68) and the control group (20.7 minutes at baseline, 20.5 minutes in the treatment period, p  = 0.89) with the D-i-D treatment estimate of a reduction of 0.16 minutes (95% CI: −6.61 to +6.92, p  = 0.96).

Fig. 3.

Fig. 3

Pajama time (a measure of work outside of work hours) over time for intervention and control participants. DAX, Microsoft/Nuance Digital Ambient eXperience.

Fig. 4.

Fig. 4

Time in notes per appointment over time for intervention and control participants. DAX, Microsoft/Nuance Digital Ambient eXperience.

Business Efficiency

We compared monthly averages of outpatient adjusted wRVUs for intervention and control providers ( Fig. 5 ). Among intervention providers, mean wRVUs per encounter increased from 1.81 during the baseline period to 1.84 in the treatment period, and the control group from 1.56 in the baseline period to 1.62 during the treatment period. The D-i-D estimate of the treatment effect of −0.03 wRVUs per encounter was not significant (95% CI: −0.52 to +0.44, p  = 0.9).

Fig. 5.

Fig. 5

Outpatient adjusted wRVUs per appointment over time for intervention and control participants. DAX, Microsoft/Nuance Digital Ambient eXperience; wRVU, work Relative Value Unit.

Same-day chart closures ( Fig. 5 ) were unchanged in the intervention group (55.2% at baseline, 54.1% during the treatment period, p  = 0.54) and the control group (56.0% at baseline, 61.3% during the treatment period, p  = 0.16).

Qualitative Comments

We reviewed 48 comments elicited from our support-on-site training team visiting intervention providers. Of these, 33 (69%) had a positive sentiment, for example, one comment from the training team after visiting a provider stated: “No technical issues. With newest algorithm, [I] noticed more accurate and logical in where it put things than the previous algorithm. No consent issues. Feels it makes her faster.” Seven (15%) were neutral, for example, “The assessment & plan is not ideal as I cannot use medical lingo with patients…and needs a lot of modification after. The HPI is good and captures everything and more.” Eight (17%) were negative, for example,

“I've been trying this week to use it for pre-charting as well, but I've found that when I use it for the patient interview, it seems to disregard our conversation in that circumstance…it seems to skip it altogether and go straight to the A/P, so probably won't be using it to pre-chart any more.”

Of note, all eight negative comments were from the first 3 weeks of implementation and prior to the model update on June 11, 2024.

In addition, we received unsolicited feedback e-mails from several enthusiastic participants with high adoption rates, including

“I definitely will advocate to continue to use this for our practice, it has definitely been helpful to me—to have conversation with family without looking at charts, engage more, have more precise and complete information from family in their own perception, and reduction in time for me working on documentation after hours.”

“The AI scribe—in just the few times I have used it, albeit not perfect, I can see will be an order of magnitude better once I get better at it. It allows the interaction with the patient to be more personal as there is not an extra person in the room. I don't have to worry about turnover of staff, and I don't have to worry about inconsistency of scribe help.”

Discussion

In this rapid, randomized QI trial of AI scribes among pediatric subspecialty physicians, subjective measures of clinician experience substantially improved while objective measures of clinician efficiency and business measures were grossly unchanged. We found wide variability in usage and engagement with the AI scribe, including one provider who stopped using it altogether, while another used it for 98% of encounters. This variation highlights the importance of understanding the contexts in which AI scribes work most effectively, for example, based on provider characteristics (e.g., specialty, training level, comfort with technology, communication style, etc.), care settings (e.g., outpatient vs. inpatient or emergency department), patient characteristics (e.g., new vs. established, disease processes, language, demographic or cultural factors, etc.), and technical implementation (e.g., EHR integrated vs. standalone, user customization options, verbal vs. written consent, AI scribe functions beyond documentation such as preparing orders or patient instructions, etc.). Although our study sample size was too small to make inferences about these factors, it provides one of the earliest evaluations of AI scribes for pediatric subspecialists.

Prior studies of AI scribes have more consistently demonstrated benefits to clinician efficiency, either based on EHR log data 22 23 37 38 39 or clinician reports of increased efficiency. 20 40 Despite subjective improvements in clinician experience, we did not see any change in time working outside of work hours or time in notes per encounter. This difference in findings may be due to random variation, inadequate statistical power due to our small sample size, differences in provider or patient population (i.e., pediatric subspecialists in our study as opposed to adult primary care and medical subspecialties in most prior studies), or our recruitment and comparison strategy. To our knowledge, no other published study has used a randomized design to evaluate the use of AI scribes, relying on either pre–post comparisons within the intervention group or comparison to a control group of providers who did not express interest in using the AI scribe. Early technology adopters may be different from their peers, leading to unmeasured confounders and results that may be due to selection bias instead of the effect of the intervention.

While the effects of AI scribes on clinician wellness and efficiency have been promising from other studies, health systems also require a business case to establish and maintain AI scribe programs. Han et al. 3 estimated that the cost of physician burnout to the health system is $7,600 per employed physician per year. Additional business cases may come from improved capture of diagnoses leading to increased wRVUs, which has been reported in press releases, although, to our knowledge, not in published studies. 13 In our study, there was no substantial or statistically significant difference in wRVUs per encounter, though the small sample size limits conclusions. Future randomized studies with larger sample sizes could help quantify this potential benefit with higher precision.

AI scribe models are being frequently updated and improved—it is likely that their impact on clinical efficiency may change over time and across clinical contexts. Ongoing evaluation is, therefore, critical to determine where and when AI scribes are effective to guide implementation focus and to identify model improvement opportunities. Rapid, randomized QI trials could be a very useful mechanism for this challenge. Many QI studies suffer from methods at high risk of bias and inappropriate causal inference. 41 However, large, lengthy randomized trials are generally infeasible for health system QI teams due to the staff effort required, the regulatory burden, and inability to make a conclusion at the speed health systems need to solve clinical and operational problems. The rapid, randomized QI trial can solve both of these problems, increasing confidence that the QI team has identified an intervention that works in a timely fashion to facilitate improvement. 24 25 26 Because (1) their primary goal is local improvement as opposed to creation of generalizable knowledge, (2) projects involve minimal risk, and (3) positive results are quickly incorporated into practice, many IRBs, including ours for this study, treat these trials as non-human subjects research, limiting the regulatory burden. 27 31 42 These criteria apply well to AI scribe implementations and could accelerate knowledge of how to leverage these systems most effectively if applied across many health systems, vendors, and clinical contexts.

Limitations

Our study has several limitations, most notably its small sample size and short observation period. In addition to limiting statistical power, the influence of particular circumstances for a small number of individuals may have an outsize influence on the results. While randomized studies aim to balance known and unknown confounders, with our small sample size, there were noticeable differences between the intervention and control groups at baseline, including higher KLAS Net EHR experience scores, lower pajama time and time in notes per encounter, and a higher percentage of manual note characters—all in the intervention group. While the D-i-D method aims to account for these baseline differences, it is also possible, for example, that the control group was already at the floor for the percentage of manual note characters and that the intervention group only improved to the floor because of their higher baseline. We also did not prespecify power calculations or aim to balance baseline factors between the two groups after randomization.

This study was also conducted at a single institution with specific workflows, requirements for written consent, and likely other local factors that may limit generalizability, even to other pediatric subspecialty settings. Our study inclusion criteria were also unique compared with many other studies—as opposed to identifying volunteer providers interested in the AI scribe technology, we included subjective assessments of EHR log data to identify opportunities for increased efficiency. This approach may yield greater insights into the utility of AI scribes for this population, but may not be as applicable to implementations that focus on other factors (e.g., interest in AI scribes) for recruitment. Additionally, this evaluation only used one AI scribe vendor, and even within the study, the model itself was updated and led to noticeable differences in AI scribe output. The intervention period was also short (3 months) compared with our baseline period (12 months) for Epic Signal-based outcomes, which could lead to differences from seasonality rather than the intervention, although we did not see clear seasonal differences in run charts. While we aimed to apply training methods equally to all participants, scheduling concerns may have led to some differences in the dose and type of training that may explain some differences in adoption. Training was focused on DAX Copilot use, but it is also possible that the increased presence of trainers may have provided other EHR efficiency tips associated with the improvements in EHR experience. Providers were also unblinded to their intervention, which could have led to Hawthorne effects, particularly if providers wanted to keep access to the AI scribe technology and therefore wanted it to look good in evaluation. Many of our objective measures of clinician efficiency (e.g., pajama time) were vendor-delivered and did not undergo additional validation to ensure their accuracy. Similarly, the KLAS Net EHR Experience Survey is a proprietary survey that provides benchmarks but has not been validated against burnout or other relevant endpoints. Our analyses of qualitative comments did not have an explicit codebook or interannotator reliability and were done in tandem by consensus of two physician informaticists, which could lead to bias. Finally, we only evaluated a subset of concepts that may be affected by AI scribes. In particular, we did not adequately evaluate patient or family experience or sentiment associated with recording visits or use of AI, we did not use a validated burnout/wellness instrument, and we did not assess differences in scribe performance for patients with limited English proficiency.

Conclusion

In our study, AI scribes improved subjective clinician experience with the EHR among pediatric subspecialists, even without observable changes in EHR log-based measures of efficiency. AI scribes have the potential to substantially reduce clinician burnout through the presumed mechanism of greater efficiency, though other mechanisms, such as reduced cognitive burden or better patient interactions, may also play a role. The landscape is rapidly changing with many vendors in the AI scribe marketplace and frequent updates to the underlying large language models powering these technologies. In that environment, ongoing evaluation comparing performance across vendors, clinical settings, and implementation contexts will be critical to obtain the benefits of these technologies without adding operational burden and cost for settings where they work less well. Rapid, randomized QI trials combined with more robust qualitative assessment to understand why AI scribes work or do not work in certain contexts would likely accelerate the ability to leverage these technologies for maximum benefit.

Clinical Relevance Statement

AI scribes have the potential to substantially reduce physician burnout by reducing documentation burden during and after clinical visits. However, the degree of benefit in different clinical settings and workflows remains unknown. In this single-center, quality randomized trial, there was high variability in AI scribe use by providers. Nonetheless, AI scribe users reported improved EHR experience and reduced symptoms of burnout compared with controls. However, there were no significant changes to objective measures of physician documentation burden, including pajama time, time in notes per encounter, or business efficiency measures.

Multiple-Choice Questions

  1. In assessing the impact of AI scribes, which of the following outcomes would be considered a subjective measure of clinician experience?

    1. Time spent on documentation per patient encounter as defined by EHR logs

    2. Number of clicks required to complete a form

    3. Physician satisfaction survey results

    4. Number of patient visits per half-day of clinic

    5. Time spent outside of work hours on documentation per half-day of clinic, based on EHR logs

Correct Answer : The correct answer is option c. Surveys asking physicians about their satisfaction would provide a subjective measure of their experience. By contrast, time spent on documentation per encounter or after hours is an objective measure of provider documentation time using automatically collected EHR logs, but these may or may not be related to providers' subjective experience. Similarly, the number of clicks required is a useful usability measure, and the number of patient visits per half-day is helpful for operational decision-making, but these do not reflect the providers' subjective experience.

  • 2. An investigator aims to understand if AI scribes reduce physician burnout. Which of the following study designs has the highest rigor for answering the question of interest?

    1. Pre/Post study comparing burnout among providers before and after using AI scribes.

    2. Trial comparing burnout among providers randomized to use AI scribes versus control providers.

    3. Cross-sectional survey comparing burnout among users and non-users of AI scribes.

    4. Case–control study comparing use of AI scribes among providers with or without high burnout scores.

Correct Answer : The correct answer is option b. A trial in which a set of providers is recruited and then randomized to use or not use ambient documentation is the most rigorous and most likely to determine the influence of ambient documentation on burnout in the population studies. Through randomization, known and unknown confounders are likely to be distributed evenly between the two groups, better isolating the impact of ambient documentation alone. A pre/post study can be helpful, but may be subject to secular trends in burnout over time or other factors postimplementation. A cross-sectional survey is also useful, but the association of ambient documentation and burnout may be confounded by unrelated factors, such as technology literacy of adopters versus non-adopters, organizational decision-making for who gets access to ambient documentation at different times, etc. Similarly, a case–control study comparing burned-out and non-burned-out providers may not demonstrate ambient documentation if it is not available widely to the sample of interest, and associations may again be confounded by similar factors like technology literacy and organizational decision-making.

Funding Statement

Funding None.

Conflict of Interest E.W.O. is the co-founder and has equity in Phrase Health, a clinical decision support analytics company. He also served as principal investigator on R41 and R42 grants with Phrase Health from the National Library of Medicine (NLM) and National Center for Advancing Translational Science (NCATS). He has received salary support from the NLM and NCATS, but no direct revenue from Phrase Health. Other authors have nothing to disclose.

Protection of Human and Animal Subjects

The CHOA Institutional Review Board deemed this work non-human subjects research (approval no.: STUDY00001962) as a QI project aimed at reducing physician burnout, rather than a project whose primary goal was to create generalizable knowledge.

References

  • 1.American Medical Association . Advocacy in action: Reducing physician burnout.2024. Accessed July 18, 2025 at:https://www.ama-assn.org/practice-management/physician-health/advocacy-action-reducing-physician-burnout
  • 2.American Medical Association . Physician burnout rate drops below 50% for first time in 4 years.2024. Accessed July 18, 2025 at:https://www.ama-assn.org/practice-management/physician-health/physician-burnout-rate-drops-below-50-first-time-4-years
  • 3.Han S, Shanafelt T D, Sinsky C A et al. Estimating the attributable cost of physician burnout in the United States. Ann Intern Med. 2019;170(11):784–790. doi: 10.7326/M18-1422. [DOI] [PubMed] [Google Scholar]
  • 4.Wu Y, Wu M, Wang C, Lin J, Liu J, Liu S. Evaluating the prevalence of burnout among health care professionals related to electronic health record use: Systematic review and meta-analysis. JMIR Med Inform. 2024;12:e54811. doi: 10.2196/54811. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 5.Adler-Milstein J, Zhao W, Willard-Grace R, Knox M, Grumbach K. Electronic health records and burnout: Time spent on the electronic health record after hours and message volume associated with exhaustion but not with cynicism among primary care clinicians. J Am Med Inform Assoc. 2020;27(04):531–538. doi: 10.1093/jamia/ocz220. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6.Levy D R, Withall J B, Mishuris R G et al. Defining documentation burden (DocBurden) and excessive DocBurden for all health professionals: A scoping review. Appl Clin Inform. 2024;15(05):898–913. doi: 10.1055/a-2385-1654. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 7.Sloss E A, Abdul S, Aboagyewah M A et al. Toward alleviating clinician documentation burden: A scoping review of burden reduction efforts. Appl Clin Inform. 2024;15(03):446–455. doi: 10.1055/s-0044-1787007. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 8.Corby S, Ash J S, Mohan V et al. A qualitative study of provider burnout: do medical scribes hinder or help? JAMIA Open. 2021;4(03):ooab047. doi: 10.1093/jamiaopen/ooab047. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 9.Micek M A, Arndt B, Baltus J J et al. The effect of remote scribes on primary care physicians' wellness, EHR satisfaction, and EHR use. Healthcare (Amst) 2022;10(04):100663. doi: 10.1016/j.hjdsi.2022.100663. [DOI] [PubMed] [Google Scholar]
  • 10.Mishra P, Kiang J C, Grant R W. Association of medical scribes in primary care with physician workflow and patient experience. JAMA Intern Med. 2018;178(11):1467–1472. doi: 10.1001/jamainternmed.2018.3956. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 11.Ziemann M, Erikson C, Krips M. The use of medical scribes in primary care settings: A literature synthesis. Med Care. 2021;59 05:S449–S456. doi: 10.1097/MLR.0000000000001605. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12.Abridge. The University of Vermont Health Network Signs Enterprise Deal with AI Leader Abridge To Improve Clinician Wellbeing. Accessed July 18, 2025 at:https://www.abridge.com/press-release/uvm-health-network-announcement
  • 13.Nuance . Northwestern Medicine deploys DAX Copilot embedded in Epic within its enterprise to improve patient and physician experiences. August 15,2024. Accessed July 18, 2025 at:https://news.nuance.com/2024-08-15-Northwestern-Medicine-deploys-DAX-Copilot-embedded-in-Epic-within-its-enterprise-to-improve-patient-and-physician-experiences
  • 14.Dbouk R, Shanks D, Mishuris R G, Yu F P. PAC02 Pajama Time(less) Stories - Early Experiences with Ambient AI Documentation. Paper presented at: UGM Conference; August 19–22,2024; Verona, WI
  • 15.Patel P, Lu H, Curren M, Ahrensfield K. UGM126 Listen Up! Two Success Stories with Ambient AI Tools. Paper presented at: UGM Conference; August 19–22,2024; Verona, WI
  • 16.Garcia P, Shah S, Ma S, Delahaie C. PAC18 Ambient Note Generative AI Success Stories_Stanford Health Care. Paper presented at: XGM Conference; May2024; Verona, WI
  • 17.Dupre B, Daigrepont T. PAC18 Ambient Note Generative AI Success Stories_Franciscan Missionaries of Our Lady Health System. Paper presented at: XGM Conference; May2024; Verona, WI
  • 18.Misurac J, Knake L A, Blum J M.Impact of ambient artificial intelligence notes on provider burnoutmedRXiv 2024.07.18.24310656. Preprint athttps://doi.org/10.1101/2024.07.18.24310656 [DOI] [PMC free article] [PubMed]
  • 19.Shanks Det al. Enhancing clinical documentation workflow with ambient artificial intelligence: Clinician perspectives on work burden, burnout, and job satisfactionmedRXiv 2024.08.12.24311883. Preprint athttps://doi.org/10.1101/2024.08.12.24311883 [DOI] [PMC free article] [PubMed]
  • 20.Galloway J L, Munroe D, Vohra-Khullar P D et al. Impact of an artificial intelligence-based solution on clinicians' clinical documentation experience: Initial findings using ambient listening technology. J Gen Intern Med. 2024;39(13):2625–2627. doi: 10.1007/s11606-024-08924-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 21.Owens L M, Wilda J J, Hahn P Y, Koehler T, Fletcher J J. The association between use of ambient voice technology documentation during primary care patient encounters, documentation burden, and provider burnout. Fam Pract. 2024;41(02):86–91. doi: 10.1093/fampra/cmad092. [DOI] [PubMed] [Google Scholar]
  • 22.Tierney A A, Gayre G, Hoberman B et al. Ambient artificial intelligence scribes to alleviate the burden of clinical documentation. NEJM Catal. 2024;5(03):x. [Google Scholar]
  • 23.Owens L M, Wilda J J, Grifka R, Westendorp J, Fletcher J J. Effect of ambient voice technology, natural language processing, and artificial intelligence on the patient-physician relationship. Appl Clin Inform. 2024;15(04):660–667. doi: 10.1055/a-2337-4739. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 24.Horwitz L I, Krelle H A. Using rapid randomized trials to improve health care systems. Annu Rev Public Health. 2023;44:445–457. doi: 10.1146/annurev-publhealth-071521-025758. [DOI] [PubMed] [Google Scholar]
  • 25.El-Kareh R, Brenner D A, Longhurst C A. Developing a highly-reliable learning health system. Learn Health Syst. 2022;7(03):e10351. doi: 10.1002/lrh2.10351. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 26.Horwitz L I, Kuznetsova M, Jones S A. Creating a learning health system through rapid-cycle, randomized testing. N Engl J Med. 2019;381(12):1175–1179. doi: 10.1056/NEJMsb1900856. [DOI] [PubMed] [Google Scholar]
  • 27.Austrian J, Mendoza F, Szerencsy A et al. Applying A/B testing to clinical decision support: Rapid randomized controlled trials. J Med Internet Res. 2021;23(04):e16651. doi: 10.2196/16651. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 28.Major V J, Jones S A, Razavian N et al. Evaluating the effect of a COVID-19 predictive model to facilitate discharge: A randomized controlled trial. Appl Clin Inform. 2022;13(03):632–640. doi: 10.1055/s-0042-1750416. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 29.Rosen K, Krelle H, King W C et al. Effect of text message reminders to improve paediatric immunisation rates: A randomised controlled quality improvement project. BMJ Qual Saf. 2025;34(05):339–348. doi: 10.1136/bmjqs-2024-017893. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 30.Tai-Seale M, May N, Sitapati A, Longhurst C A. A learning health system approach to COVID-19 exposure notification system rollout. Learn Health Syst. 2021;6(02):e10290. doi: 10.1002/lrh2.10290. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 31.Wardi G, Owens R, Josef C, Malhotra A, Longhurst C, Nemati S. Bringing the promise of artificial intelligence to critical care: What the experience with sepsis analytics can teach us. Crit Care Med. 2023;51(08):985–991. doi: 10.1097/CCM.0000000000005894. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 32.Brenner A T, Rhode J, Yang J Y et al. Comparative effectiveness of mailed reminders with and without fecal immunochemical tests for Medicaid beneficiaries at a large county health department: A randomized controlled trial. Cancer. 2018;124(16):3346–3354. doi: 10.1002/cncr.31566. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 33.Epic Systems Corporation . Signal - Efficiency Portal. Accessed July 18, 2025 at:https://signal.epic.com/Documentation/MetricReference
  • 34.KLAS Research . KLAS Arch Collaborative. Accessed July 18, 2025 at:https://engage.klasresearch.com/klas-arch-collaborative/
  • 35.Wing C, Dreyer M.Making sense of the difference-in-difference design JAMA Intern Med 20241841250–1251.. Accessed July 18, 2025 at:https://jamanetwork.com/journals/jamainternalmedicine/article-abstract/2822389 [DOI] [PubMed] [Google Scholar]
  • 36.R Core Team. R: A language and environment for statistical computing. Accessed July 18, 2025 at:https://cran.r-project.org/doc/manuals/r-release/fullrefman.pdf
  • 37.Cao D Y, Silkey J R, Decker M C, Wanat K A. Artificial intelligence-driven digital scribes in clinical documentation: Pilot study assessing the impact on dermatologist workflow and patient encounters. JAAD Int. 2024;15:149–151. doi: 10.1016/j.jdin.2024.02.009. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 38.Liu T-L, Hetherington T C, Stephens C et al. AI-powered clinical documentation and clinicians' electronic health record experience: A nonrandomized clinical trial. JAMA Netw Open. 2024;7(09):e2432460. doi: 10.1001/jamanetworkopen.2024.32460. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 39.Haberle T, Cleveland C, Snow G L et al. The impact of nuance DAX ambient listening AI documentation: A cohort study. J Am Med Inform Assoc. 2024;31(04):975–979. doi: 10.1093/jamia/ocae022. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 40.Shah S J, Devon-Sand A, Ma S P et al. Ambient artificial intelligence scribes: physician burnout and perspectives on usability and documentation burden. J Am Med Inform Assoc. 2025;32(02):375–380. doi: 10.1093/jamia/ocae295. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 41.Ralston S L, Brady P W, Kemper A R. Do we really need scholarly quality improvement? JAMA Pediatr. 2019;173(05):413–414. doi: 10.1001/jamapediatrics.2019.0067. [DOI] [PubMed] [Google Scholar]
  • 42.Finkelstein J A, Brickman A L, Capron A et al. Oversight on the borderline: Quality improvement and pragmatic research. Clin Trials. 2015;12(05):457–466. doi: 10.1177/1740774515597682. [DOI] [PMC free article] [PubMed] [Google Scholar]

Articles from Applied Clinical Informatics are provided here courtesy of Thieme Medical Publishers

RESOURCES