Abstract
Ambient artificial intelligence (AI) scribes offer promise for reducing documentation burden, yet their effects on surgical practice have not been well defined. We conducted a pilot study of an ambient AI scribe across 79 ambulatory providers (three surgeons) in a multihospital system from December 2024 to February 2025. Surgeon adoption of the AI scribe ranged from 58% to 90% of visits. We evaluated workload via pre- and post-intervention surveys (NASA-TLX mental demand and perceived rush), burnout rates, scheduling capacity, and Epic Signal metrics (note length, time per note) and compared billing data for high utilizers between August and October 2024 and the pilot period. Average “pajama time” did not change significantly (p=0.55). NASA-TLX mental demand decreased from 14 to 5 (p=0.08) and perceived rush from 15 to 5 (p=0.06). Burnout declined from 67% to 33%. Two surgeons reported the capacity to add three patients per clinic. The billing metric showed no significant changes. Undivided attention scores improved from 3.5 to 4.1 (p<0.0001). This preliminary data shows promise that ambient AI scribes in surgical clinics may reduce documentation burden and burnout, with potential gains in efficiency and throughput. Larger studies are warranted to further confirm these findings.
Keywords: Technology, Health Services, Information Technology
Introduction
Ambient artificial intelligence (AI) emergence has been rapid over the past 5 years. Ambient AI scribes are primarily marketed as a method for reducing physician burden. Indeed, there are studies that show when implemented in single systems, there is a reduction in time spent in the electronic medical record (EMR), but there is not yet any definitive benefit for burnout.1 Some data have found that reductions in reported workload are associated with the introduction of ambient AI.2 Additionally, the literature lacks specialty-specific outcomes for this technology, especially among surgeons.
Surgical documentation is critical in our dynamic patient population as these patients are followed from the clinic or inpatient setting to the operating room, through the postoperative course. Surgeons are not known to write extensive notes, attempting to convey critical information in a concise manner. However, surgeons often fall short of this, with one single center study finding surgeons failed to document symptomatic change in 28% of postoperative notes, and 67% of notes failed to document functional change.3 Another medical record review study identified hand surgeons in a single center were at risk of bias in medical documentation.4 Medical record documentation from surgeons is poorly studied in the literature, but surgical documentation can certainly be improved.
The current literature lacks a detailed analysis of outcomes for surgeons and AI scribes. Further, subjective and objective data for how this emerging technology impacts the practice of surgery are currently limited. The goal of this paper is to evaluate how our single medical group experience with ambient AI may influence surgical practice.
Methods
Setting, selection of participants, and pilot design
The Hawaii Pacific Health Medical Group (HPHMG) elected to pilot an ambient AI device. The group includes four hospital systems in the state of Hawaii. A goal of 75 providers was set and agreed on between HPHMG and the AI purveyor. Surgeons were included to provide specialty representation. While department leadership identified participants, final inclusion required agreement from providers to trial the tool with a goal of ≥60% utilization.
After providers were recruited, they were given a familiarization period before the pilot began. The pilot was run for 3 months, from 1 December 2024 through 28 February 2025. Prior to starting the pilot, providers completed an initial survey and attended a required training session. A final survey was distributed around the midpoint of the 3-month pilot to compare task load, burnout and work hours with the prepilot survey. Epic Signal metrics were collected on pajama time, progress note length, and average time in notes per appointment. Pajama time, as defined by Epic, is ‘charting activities outside 7 AM to 5:30 PM on weekdays, time outside scheduled hours on weekends, and time on unscheduled holidays.’ Utilization of the ambient AI scribe was determined by the proportion of AI notes generated over the total number of clinic visits. High utilizers used the AI scribe in ≥60% of visits, medium utilization 25%–59% of visits, and low utilization ≤25% of visits. Patient satisfaction data for ‘doctor always listens carefully’ was collected and analyzed with a Fisher’s exact test.
Data collection and statistical analyses
Pre-pilot and final pilot surveys were analyzed using the paired student’s t-test for all Likert scale components. Burnout results were compared with a χ2 test and the relationship further defined with an OR. For billing reporting, we focused on high utilizers comparing billing data from August through October 2024 to December 2024 through February 2025. For these providers, professional billing transactions were run through Slicer Dicer and then filtered by department before being grouped by CPT code. Low-volume departments (<100 visits over 3 months) were excluded. Surgeons were included as those specialties listed as general surgery, orthopedic surgery, neurosurgery, urology, and vascular surgery. Surgeons who dropped out for poor uptake were excluded from statistical analysis but included in commentary. Patient satisfaction data for the proportion of patients responding ‘Yes, definitely’ on patient satisfaction surveys to the prompt ‘Did this provider listen to you carefully’ were collected and analyzed with a Fisher’s exact test for high utilizer physicians.
Results
In our pilot of 79 providers, a total of eight providers from surgical specialties participated. Of these, only three utilized the AI tool enough for statistical analysis representing neurosurgery, general surgery, and orthopedic surgery. Two of the three providers met high utilizer status by using the AI tool in at least 60% of notes. The highest usage was the neurosurgeon, using the AI in 90% of clinic visits. Next was the orthopedic surgeon using it in 89% of clinic visits, while the general surgeon used the AI tool in 58% of visits. Since the pilot, the AI scribe has been expanded to 100 providers, including the original 79. In the expanded group, a pediatric surgeon and another general surgeon were added to the cohort, with the pediatric surgeon using the AI in 68% of visits and the general surgeon using it in only 20% of visits in less than 2 months of data collection.
For the three surgeons who participated in the pilot, surveys on perceived workload were collected before and after the pilot (figure 1). On the 20-point scale derived from the NASA Task Load Index (NASA-TLX), average response for mental demand decreased from 14 to 5 (p=0.08). The average response for how rushed the surgeons felt completing clinic notes decreased from 15 to 5 (p=0.06) and how mentally demanding note writing was decreased from 16 to 7 (p=0.24). For self-reported burnout, at the start of the pilot, two of the three surgeons reported burnout symptoms and only one reported symptoms by the end of the pilot.
Figure 1. Surgeon responses for task load surveys. Each question is presented as an estimation plot comparing survey response for each surgeon before AI to their own response with AI. Mean difference is plotted with 95% CI error bars. Group differences were calculated with a paired student’s t-test. ns denotes no significance between groups. AI, artificial intelligence. Net change represents the score change from with AI to before AI.
Epic Signal data were collected for the participating surgeons and are displayed in table 1. Pajama time averaged across the three surgeons did not change significantly (p=0.55). It is important to note that pajama time in Epic includes all activities (documentation, order entry, chart review), whereas time in notes is a more direct measure of documentation. In our data, the neurosurgeon demonstrated a decrease in pajama time (20%, p=0.26) despite an increase in time in notes per appointment (30%, p=0.08) and a 49% (p<0.05) increase in time in notes per day, suggesting a shift of documentation into clinic hours rather than after-hours. By contrast, the orthopedic surgeon experienced both decreased pajama time (29%) and decreased time in notes per day (26%, p<0.01) as well as per appointment (19%, p<0.05), reflecting greater documentation efficiency. The general surgeon had a 26% increase in pajama time (p=0.33), a 7% (p=0.62) decrease in time per note, and a 5% (p=0.74) increase in time in notes per day. Average length of progress notes increased by 22% (p=0.15) for the general surgeon and by 15% (p<0.05) for the neurosurgeon while decreasing by 6% (p=0.16) for the orthopedic surgeon.
Table 1. Epic signal data for general, orthopedic, and neurosurgeons.
| Variables | Categories | General surgeon | Orthopedic surgeon | Neurosurgeon |
|---|---|---|---|---|
| Pajama time, minutes | Before AI | 16.9 | 20.1 | 19.2 |
| With AI | 21.4 | 26.1 | 15.3 | |
| Change | 26% | 29% | −20% | |
| Time in notes per appointment, minutes | Before AI | 1.6 | 4.2 | 4.4 |
| With AI | 1.5 | 3.4 | 5.7 | |
| Change | −7% | −19% | 30% | |
| Time in notes per day, minutes | Before AI | 4.5 | 24.6 | 9.2 |
| With AI | 4.8 | 18 | 13.8 | |
| Change | 5% | −26% | 49% | |
| Note length, characters | Before AI | 2571 | 12,701 | 4616 |
| With AI | 3139 | 11,850 | 5304 | |
|
Change |
22% |
−6% |
15% |
AI, artificial intelligence.
We sought to determine avenues by which adding this AI tool to the provider toolkit may improve revenue. In our prepilot and postpilot surveys, all 79 providers were asked if they thought patients could be added to their clinic schedule. On this 1 to 5-point Likert scale, the average response for the ability to add patients to their schedule did not change (3.4 to 3.4, p>0.9999). For our surgeon cohort, two of the three providers felt they could add three patients each to their clinic schedule.
When looking at billing coding level, we did not find any significant changes for any of our clinics or visit types, but our ability to analyze this was limited by low power when looking at specific clinics as, out of surgical offices, the orthopedic surgeon was the only provider to meet the minimum number of visits for analysis. In the orthopedic surgery clinic, established patient visit billing level did not change after adding the AI tool (p=0.70) (online supplemental table 1). For new office visits and office consult visits, patient volumes did not meet minimum requirements for analysis. Established patient office visits did meet volume thresholds in the orthopedic surgery clinic and are reported in online supplemental table 1.
On the pre- and postpilot surveys, providers were asked to assess several subjective metrics on a scale of 1-. Provider assessed ability to give their undivided attention to their patients increased significantly from 3.5±1 to 4.1±0.8 (p<0.0001) (figure 2). When asked how they perceived their patients’ ability to understand their clinic notes, there was no significant change in score (3.9±0.9 pre-pilot vs. 3.7±0.9 after the pilot, p=0.16) (figure 2). On a scale of 1–5 in the post-pilot survey, providers were also asked to assess if AI improved note quality and work satisfaction. Providers rated improvement of note quality at 3.1 and work satisfaction with AI at 3.7.
Figure 2. Provider survey responses for perceived impact on patient interactions. Of all 79 pilot providers, answers were given on a 1-5 scale. (A) provider responses for perceived ability to give undivided attention to patients during office visits. (B) Can patients understand clinic notes as reported by physicians? Error bars represent one standard deviation. Before AI and with AI groups were compared with a paired student’s t-test. **** denotes p<0.0001; ns denotes no significance. AI, artificial intelligence.
Finally, patient satisfaction scores were also collected for all providers as part of the pilot and were compared with each provider’s prepilot patient satisfaction surveys. In table 2, we show the proportion of patients reporting “Doctor always listens carefully” for our high utilizer surgeons. The general surgeon and neurosurgeon saw non-significant increases in patient satisfaction (4.4% and 6.2% respectively), while the orthopedic surgeon saw a small non-significant decrease (−1.4%).
Table 2. 'Provider always listens to me carefully’ scores before and after AI.
| Specialty | Patient Satisfaction Score of Before AI (%) | Respondents of Before AI (N) | Patient Satisfaction Score of With AI (%) | Respondents of With AI (N) | Net Change | OR (95%CI) | P value |
|---|---|---|---|---|---|---|---|
| General | 90.0 | 30 | 94.4 | 71 | +4.4% |
0.53 (0.14 to 2.26) |
0.420 |
| Orthopedic | 91.4 | 116 | 90.0 | 120 | −1.4% |
1.18 (0.51 to 2.85) |
0.820 |
| Neurosurgery | 89.5 | 105 | 95.7 | 92 | +6.2% |
0.39 (0.13 to 1.15) |
0.110 |
AI, artificial intelligence; CI, confidence interval; OR, odds ratio.
Discussion
This study was completed to better understand how an ambient artificial intelligence scribe may impact physicians and patient interactions. We looked to further sub-analyze how this AI tool may impact surgical practice. Here, we describe outcomes from the entire 79-physician cohort on practice implications including adding patients to the clinic schedule, billing changes for office visits, and ability to give undivided attention to patients. We further analyzed outcomes for three surgeons participating in the pilot specializing in general surgery, orthopedic surgery, and neurosurgery. We found that although surgical participation is limited, there was favorable reception of the AI tool among the three surgeons. Five surgeons, however, had low utilization and were excluded from analysis. Since the pilot ended, there is a new cohort of surgeons using the AI scribe with excellent early uptake. Several outcomes, including reductions in NASA-TLX mental demand and perceived urgency, did not reach statistical significance, reflecting the small sample size and limited power of this pilot. While effect sizes suggest potentially meaningful reductions (e.g., mental demand decreased from 14 to 5; urgency from 15 to 5), these results should be interpreted as exploratory rather than definitive. Confidence intervals for these measures were wide, underscoring the need for larger, adequately powered studies to validate these preliminary findings. However, despite finding decreased perceived demand, we found much variability in the time providers were actually spending in the EMR. It does appear that the trend for pajama time and note writing time is decreasing over the 3-month pilot in our month-to-month analysis, but until there are long-term data, we are unable to make conclusions about impact for these providers. Further, as we are increasing the number of surgeons using the AI tool, we will be better able to understand how time in the EMR and notes change.
We asked the five surgeons who had poor utilization what factors contributed to this outcome. Among the eight surgeons who participated (three high utilizers, five low utilizers), no specific workflow changes were required or reported. Surgeons who derived the most benefit tended to generate new notes at each encounter, whereas those who relied heavily on precharting or copy-forward documentation did not perceive significant time savings. For structured physical exams such as musculoskeletal evaluations, several surgeons opted to continue using their existing Smart Tools or templates in Epic rather than the AI-generated exam sections. One liked the technology and stated it worked well during appointments, but found most of their time was chart prepping for the next day of clinic, so the AI tool did not save time. Another provider would need to make significant edits as they felt notes were written for patients rather than providers, therefore lacking ability to convey complicated information. They also felt that the AI-generated note was too broad for a specialty surgical visit and fell short on complex visits, informed consent, and shared decision-making. This was backed by another provider who commented on poor transference of risks and benefits discussion for procedures into the note, which created short summaries of the interaction. Some of the provider concerns with uptake may have improved with continued use, but documentation of surgical risk remains an important area for improvement. A possible solution to this is a dictation mode that will allow direct verbatim dictation to parts of the note deemed important, such as consenting for surgery. Interestingly, while one orthopedic surgeon struggled, another stated the AI scribe ‘…has been amazing…Life-changing! Really makes practicing in the clinic a lot more efficient.’ Barriers to full adoption among surgeons include dissatisfaction with note output (several reported that notes did not ‘sound like them’ and required extensive editing), device-related challenges (e.g., Android user relying on loaned iPhone), and lack of perceived time saving for providers who relied heavily on previsit chart preparation and copy-forward documentation. These factors may reflect unmet expectations that must be addressed to optimize ambient AI for surgical documentation.
These challenges highlight the need for specialty-specific AI development. For surgeons, accurate documentation of complex encounters, particularly informed consent, risk–benefit discussions, and perioperative decision-making, is essential. Ambient AI scribes will likely require further customization to surgical workflows, including integration with existing specialty templates and Smart Tools in the electronic health record. Another avenue for improvement is the incorporation of surgical-specific training datasets to enhance recognition and documentation of nuanced conversations. Such refinements may improve note fidelity and reduce the need for manual edits, thereby increasing surgeon adoption and trust in AI documentation systems
It is well known that long hours and poor posture contribute to surgeon fatigue and mistakes, so offloading medical record tasks may help to improve surgeons’ outcomes and work satisfaction.5 6 Not only could mental strain be improved, but there may be an opportunity to study physical symptoms of EMR fatigue. An unanticipated benefit of ambient AI scribes is that there is potentially a role for improved ergonomics and physical strain, which may be especially beneficial for surgeon longevity. Physicians left comments during our pilot such as “…not to mention giving my wrists a break,” “less fatigue or pain related to typing,” “less typing has been great for my tendonitis,” and “…I noticed a significant improvement in my neck discomfort and carpal tunnel syndrome.” This subjective feedback creates new avenues for understanding ways in which technology may improve surgical practice.
Improving cost is a key implication of emerging technology that is not well understood with ambient AI scribes on a clinic or hospital system-wide basis. Accurate CPT coding is critical to reimbursement and avoidance of procedure denial, whereas understanding of these topics by graduating residents has been found to be poor.7 This is important as a large study on common reasons for claim denial among a pediatric surgery group found that inaccuracy in coding led to increased likelihood of denial.8 As AI-mediated visit coding improves through algorithm training, there is a role for improving appropriate CPT coding to limit claim denials.
AI scribe companies often market improvements in integration with billing systems, largely aimed at supporting documentation for complexity-based billing. In principle, medical complexity should not change with AI use, although better-structured documentation may allow for more accurate coding of complexity that was previously under-documented. By contrast, time-based billing is unlikely to be affected, as AI does not alter the duration of counseling or decision-making, even if documentation time decreases. One study found criticism in AI scribes for improving billing level coding because as of now, there is too much variation in AI coding that makes it potentially unreliable.9 Given these factors, billing level may not be a reliable primary metric of ambient AI success, though ongoing algorithm refinements could still improve fidelity in complexity-based coding. We are continuing to run analyses for these high utilizer providers outside of the initial pilot period to determine if coding trends or shifts may exist.
In addition to changes in billing practice, increasing patient census in the clinic is a possible source for increased revenue leading to AI scribe cost offset. Two of the three surgeons in our pilot study indicated they could add up to three more patients to their clinic. Adding at least one patient per clinic day is anticipated to cover the cost of the AI scribe, and more patients may improve practice revenue. Including more patients may also help enhance access to care for patients which improves patient satisfaction in addition to monetary incentives for the clinic.10 In addition to this potential measure of improving patient satisfaction, utilization of ambient AI tools may enhance face-to-face communication and thereby make patients feel more involved in the process of their care. We found no significant difference in patient experience scores for the metric “doctor always listens to me carefully” across our three surgeons. Future studies involving more direct patient feedback will aid in better understanding how integration of AI tools affects the doctor-patient relationship.
This study is limited by the small number of surgeons who reached adequate utilization for analysis (three of eight enrolled). As a result, statistical power is low and variability in adoption rates complicates interpretation. We used providers as their own control, comparing pre- and post-intervention outcomes, rather than comparing against non-users within the medical group. This approach was chosen because of the high variability in specialty, documentation styles, and ambulatory workloads across providers, which we felt would confound between-group comparisons. While this design improves internal consistency, it does not eliminate the possibility that seasonal variation, concurrent workflow initiatives, or organizational changes may potentially influence results. Longer-term follow-up and, ideally, multicenter studies including comparator groups will be necessary to more definitively assess causality. These findings should therefore be viewed as preliminary and hypothesis-generating. Expansion of the dataset, inclusion of additional cohorts of providers, and longer follow-up will be necessary to better define the role of ambient AI in surgical practice. Another important limitation is that our most compelling findings were derived from self-reported outcomes, including NASA-TLX workload, perceived urgency, and burnout. While these subjective data highlight the potential of ambient AI to reduce cognitive and emotional burden, they should be complemented in future studies by robust objective endpoints such as independently verified documentation time, patient throughput, or clinic efficiency. In our pilot, Epic Signal data served as an objective measure, but findings varied by provider and the clinical relevance of relatively small numerical differences remains uncertain. Larger cohort studies will be needed to validate whether these subjective improvements translate into measurable workflow and efficiency gains. Despite these limitations, we believe reporting our early experience and incorporating subjective feedback provides important insights for this emerging technology.
Looking forward, the expansion of ambient AI into broader surgical practice, particularly within pediatric surgery where specialist scarcity amplifies individual provider workload and offers the potential to reallocate valuable time toward direct patient care and operative duties. As AI integrates into surgical clinics, there remains much uncertainty in unseen risk outside of this possible benefit described here. For example, it is unclear how ambient AI and the existence of recorded visits will impact medicolegal proceedings, among other common practice management gaps for early career surgeons.11 Continued longitudinal data collection, encompassing both objective workflow metrics and user-driven feedback, will be critical to identify system refinements that drive toward an “ideal scribe” capable of producing finalized notes with minimal clinician edits. Such efforts will inform both the optimization of AI integration and the assessment of long-term impacts on surgeon efficiency, ergonomics, and patient throughput.
In conclusion, ambient AI scribes have shown an unclear impact on surgical practice within this 3-month pilot. While there is improved perceived workload pertaining to note-writing, possible cost benefits such as billing level and adding patients to the clinic schedule are not yet understood. Complex surgical visits and accurate conveyance of the consenting process are still critical areas for improvement based on surgeon commentary. With further algorithm development and integration into our hospital system, we anticipate that surgeons and their practice will benefit from this emerging technology.
Supplementary material
Acknowledgements
Hawaii Pacific Health Medical Group AI working group: Emi Thieme, William Huynh, James Lin, Eric Chang, and Jerome Lee for weekly improvement meetings and ambient AI implementation progression. Abridge, Inc. for providing provider usage data and aid in training providers on the platform.
Footnotes
Funding: The authors have not declared a specific grant for this research from any funding agency in the public, commercial, or not-for-profit sectors.
Provenance and peer review: Part of a Topic Collection. Not commissioned; externally peer-reviewed.
Patient consent for publication: Not applicable.
Ethics approval: This study involved human participants, but the Hawaii Pacific Health Institutional Review Board (Study number: 2024-080) was exempted by the IRB under 45 CFR 46.102(l) and classified the project as quality improvement. Participants gave informed consent to participate in the study before taking part.
Data availability free text: All data are available upon reasonable request to the corresponding author.
Data availability statement
Data are available upon reasonable request.
References
- 1.Tierney AA, Gayre G, Hoberman B, et al. Ambient Artificial Intelligence Scribes to Alleviate the Burden of Clinical Documentation. NEJM Catalyst. 2024;5 doi: 10.1056/CAT.23.0404. [DOI] [Google Scholar]
- 2.Shah SJ, Devon-Sand A, Ma SP, et al. Ambient artificial intelligence scribes: physician burnout and perspectives on usability and documentation burden. J Am Med Inform Assoc. 2025;32:375–80. doi: 10.1093/jamia/ocae295. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 3.Miller JM, Velanovich V. The natural language of the surgeon’s clinical note in outcomes assessment: a qualitative analysis of the medical record. Am J Surg. 2010;199:817–22. doi: 10.1016/j.amjsurg.2009.06.037. [DOI] [PubMed] [Google Scholar]
- 4.Calfee R, Fynn-Thompson E, Stern P. Surgeon Bias in the Medical Record. Orthopedics. 2009;32:732–6. doi: 10.3928/01477447-20090818-07. [DOI] [Google Scholar]
- 5.Schlussel AT, Maykel JA. Ergonomics and Musculoskeletal Health of the Surgeon. Clin Colon Rectal Surg. 2019;32:424–34. doi: 10.1055/s-0039-1693026. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 6.McCormick F, Kadzielski J, Landrigan CP, et al. Surgeon fatigue: a prospective analysis of the incidence, risk, and intervals of predicted fatigue-related impairment in residents. Arch Surg. 2012;147:430–5. doi: 10.1001/archsurg.2012.84. [DOI] [PubMed] [Google Scholar]
- 7.Andreae MC, Dunham K, Freed GL. Inadequate Training in Billing and Coding as Perceived by Recent Pediatric Graduates. Clin Pediatr (Phila) 2009;48:939–44. doi: 10.1177/0009922809337622. [DOI] [PubMed] [Google Scholar]
- 8.Ryan ML, Mutore KT, DeLeon J, et al. Improving Billing and Collections in a High-Volume Pediatric Surgery Practice: Denials-Based Approach. J Am Coll Surg. 2023;236:630–5. doi: 10.1097/XCS.0000000000000559. [DOI] [PubMed] [Google Scholar]
- 9.Bracken A, Reilly C, Feeley A, et al. Artificial Intelligence (AI) – Powered Documentation Systems in Healthcare: A Systematic Review. J Med Syst. 2025;49:28. doi: 10.1007/s10916-025-02157-4. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 10.Volk AS, Davis MJ, Abu-Ghname A, et al. Ambulatory Access: Improving Scheduling Increases Patient Satisfaction and Revenue. Plast Reconstr Surg. 2020;146:913–9. doi: 10.1097/PRS.0000000000007195. [DOI] [PubMed] [Google Scholar]
- 11.Sinyard RD, Veeramani A, Rouanet E, et al. Gaps in Practice Management Skills After Training: A Qualitative Needs Assessment of Early Career Surgeons. J Surg Educ. 2022;79:e151–60. doi: 10.1016/j.jsurg.2022.06.009. [DOI] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Data Availability Statement
Data are available upon reasonable request.


