1. Lukac and colleagues compared two ambient AI scribes with usual documentation approaches among outpatient clinical visits.
2. Both AI scribes showed some improvements in burnout-related measures but continued to produce clinically significant inaccuracies.
Evidence Rating Level: 1 (Excellent)
Study Rundown: Human scribes have been deployed with some success to mitigate physician burnout. However, they pose high costs and accessibility challenges. Digital scribes that incorporate AI technology may reduce this challenge by generating clinical notes from recorded encounters. Lukac and colleagues evaluated two AI scribes, Dragon Ambient eXperience (DAX) Copilot and Nabla, in outpatient practice. Physicians were assigned to DAX, Nabla, or usual care for two months. The primary outcome was change in electronic health record (EHR) time spent writing each note, and secondary outcomes included burnout, cognitive load, and work exhaustion. Nabla users experienced a 9.5% decrease in time-in-note compared to the control group, whereas DAX produced no significant reduction. Both scribe groups showed favorable changes in composite burnout, physician cognitive load, and work exhaustion measures. However, users reported several clinically significant inaccuracies, including one adverse patient safety event. This study showed that AI scribes can improve documentation efficiency, but further development is needed to ensure safety and reliability.
Click here to read the study in NEJM AI
Relevant Reading: Beyond human ears: navigating the uncharted risks of AI scribes in clinical practice
In-Depth [randomized controlled trial]: The trial enrolled 238 outpatient physicians from 14 specialties and randomized them to DAX Copilot (n=79), Nabla (n=79), or usual care (n=80). Physicians were instructed to use the scribes or document per usual methods for two months. The primary outcome was change in EHR time-in-note, calculated as weekly minutes spent writing notes divided by notes completed and averaged monthly. Secondary outcomes included surveys to assess physician burnout, cognitive load, and work exhaustion, as well as usability, inaccuracies, bias, and safety measures. The study found that, compared to control, Nabla reduced time-in-note by 9.5% (95% confidence interval [CI], -17.2 to -1.8%; p=0.02), whereas DAX did not demonstrate a significant reduction (-1.7%; 95% CI, -9.4 to 5.9%; p=0.66). Both tools showed potentially favorable changes in physician burnout, cognitive load, and work-exhaustion scores, but these secondary outcomes were not hypothesis-tested and required further confirmation. Poststudy surveys were completed by 76% of controls and 82% of each intervention group. Occasional inaccuracies were reported, with Nabla having 2.8 mean inaccuracies (standard deviation [SD] 1.0) and DAX having 2.7 (SD 1.1) inaccuracies. One adverse patient safety event was reported, with the event deemed mild by five physician co-authors. This study was limited by incomplete AI scribe uptake and possible survey nonresponse bias. Overall, this study showed that AI scribes have the potential to reduce physician documentation burden, but more long-term studies are needed to validate their capabilities.
Image: PD
©2026 2 Minute Medicine, Inc. All rights reserved. No works may be reproduced without expressed written consent from 2 Minute Medicine, Inc. Inquire about licensing here. No article should be construed as medical advice and is not intended as such by the authors or by 2 Minute Medicine, Inc.




