A primary care physician walks into an exam room, sets a small device on the desk, and says, “I’m going to let this listen to our conversation so I can focus on you instead of my screen.” Minutes after the visit ends, a draft clinical note is waiting in the electronic health record — organized into a history, an assessment, and a plan — generated not by a transcriptionist in another room but by a large language model that processed the audio of the visit itself. A year ago, that scenario barely existed outside of a handful of pilot programs. In 2023, it has become one of the most closely watched applications of generative AI in medicine.

The technology is called ambient AI documentation, or more casually an “AI scribe,” and it has arrived at a moment when clinician burnout tied to documentation is already a well-established crisis. Health systems, EHR vendors, and startups are moving quickly — arguably faster than the evidence base is moving — to bring these tools into exam rooms. That combination of urgent need and early-stage technology makes this a topic worth understanding carefully, cautions included.

Why Is Documentation Burden Fueling Burnout in the First Place?

The link between EHR documentation and clinician burnout did not start with generative AI, and it will not be solved solely by it. A widely cited 2016 time-and-motion study published in the Annals of Internal Medicine found that physicians in ambulatory practice spent nearly two hours on EHR and desk work for every hour of direct face time with patients, plus additional hours of charting after the workday ends — the pattern clinicians have nicknamed “pajama time.” National physician surveys have consistently ranked excessive documentation and bureaucratic burden among the top drivers of burnout, alongside workload and loss of autonomy.

That backdrop matters for understanding why ambient AI scribes have generated so much attention in such a short period. Documentation burden is one of the few burnout drivers that a technology can plausibly address directly, by intervening in the moment the note gets written rather than by asking clinicians to manage stress around an unchanged workload. Whether ambient AI actually delivers on that promise at scale is still an open question in 2023, but the underlying problem it targets is not in dispute.

What Changed to Make Ambient AI Scribes Possible Now?

Voice-recognition dictation tools have existed in clinical settings for years, and human medical scribes — whether in-room or remote — have long been used to offload documentation from physicians. What is new in 2023 is the combination of ambient audio capture with generative AI models capable of turning a natural, unstructured conversation into a structured clinical note without a human listening in real time or transcribing afterward.

The clearest public marker of this shift came in March 2023, when Microsoft’s Nuance division announced Dragon Ambient eXperience (DAX) Express, describing it as the first ambient clinical documentation application to combine conversational and ambient AI with GPT-4, OpenAI’s large language model, delivered through Microsoft’s Azure OpenAI service. Nuance and Epic subsequently announced deeper integration of the tool into Epic’s EHR workflow in June 2023. Other ambient documentation products — some predating the generative AI wave and since upgraded with it, others newly built around it — have followed a similar pattern of rapid rollout across health systems and specialty practices during 2023.

How Does an Ambient AI Scribe Actually Work?

Although products differ in detail, the general workflow described by vendors and early adopters follows a consistent pattern:

Capture. With the patient’s awareness and typically their explicit consent, a microphone-equipped device or an app on a phone, tablet, or computer records the audio of the visit as the clinician and patient talk normally, without either party dictating in a structured or artificial way.

Transcription. Automatic speech recognition converts the audio into text, distinguishing speakers where the technology supports it.

Drafting. A large language model processes that transcript and generates a draft clinical note, typically organized into a standard format such as a SOAP note (Subjective, Objective, Assessment, Plan), pulling out the clinically relevant content from the surrounding conversational language.

Review and edit. The draft is delivered into the clinician’s EHR workflow, usually within minutes, where the clinician is expected to review it, correct or add detail, and formally sign and finalize it as the official medical record — the same way they would review a note drafted by a human scribe or trainee.

That last step is the one vendors and early clinical adopters emphasize most consistently in 2023: the AI-generated text is described as a draft, not a finished note, and the clinician’s review is positioned as a required part of the workflow rather than an optional check.

What Does the Early Evidence Say?

Because ambient AI scribes built on generative AI are new in 2023, the evidence base is thin, and what exists mostly takes the form of early pilots, vendor-reported figures, and small single-institution studies rather than large randomized trials. Several health systems have begun piloting these tools this year and are tracking early results:

  • Mass General Brigham piloted an ambient AI documentation tool starting in July 2023 with a small initial group of physicians, with plans to expand the pilot and formally evaluate its effect on documentation time and clinician experience.
  • Other academic medical centers have launched similar pilots in 2023, generally structured as quality-improvement initiatives that survey clinicians on task load, documentation time, and satisfaction before and after adoption, rather than as blinded controlled trials.
  • A separate multi-year pilot of ambient documentation technology in a dermatology practice — begun before the current generative AI wave and continuing through early 2023 — found daily documentation time fell from about 90 minutes to about 70 minutes among clinicians who stayed with the tool, with the large majority reporting they would be “very disappointed” to lose access to it. That study relied on clinician self-report for satisfaction and did not formally measure note accuracy, which is a limitation worth keeping in mind when generalizing from it.

Taken together, this early evidence points in a promising direction — reduced self-reported documentation time and generally positive clinician sentiment — but it should be read as preliminary. Sample sizes are small, most studies are unblinded and rely heavily on self-report, follow-up periods are short, and none of the generative-AI-specific pilots launched in 2023 have had time to produce peer-reviewed outcomes on burnout at the scale of a large health system. Anyone citing early results, including this article, should describe them as encouraging pilot data rather than settled proof of effect.

What Are the Accuracy and Hallucination Concerns?

The same generative AI capability that lets these tools turn a free-flowing conversation into a coherent note is also the source of their central risk. Large language models are known to sometimes produce “hallucinations” — text that is fluent and plausible-sounding but factually incorrect or entirely fabricated — and researchers studying LLMs broadly, not scribe tools specifically, have documented that this failure mode can appear even in carefully constrained tasks. In a clinical note, a hallucination might mean a symptom the patient did not report, a medication dose that was misheard or invented, or an assessment that does not match what was actually discussed.

Several factors compound this risk in the ambient scribe context specifically:

  • Audio quality and speaker confusion. Real clinical conversations include crosstalk, background noise, accents, medical jargon, and multiple speakers (patient, clinician, family members), any of which can degrade transcription accuracy before the language model ever sees the text.
  • Omission as well as fabrication. A note can be inaccurate not only by adding something that was not said, but by leaving out something clinically important that was.
  • Automation bias. A fluent, well-formatted AI-generated note can look more authoritative than it is, and a busy clinician reviewing dozens of notes a day may be more likely to sign off on subtly wrong text than they would catch an error in their own dictation.
  • Immature evaluation methods. As of 2023, there is no widely agreed-upon, standardized way to measure the accuracy of an AI-generated clinical note at scale, which makes it difficult for health systems to compare tools or set a clear accuracy bar before deployment.

None of this means the technology is unsafe by design — human-authored notes are not error-free either, and human medical scribes make mistakes too. But it does mean the accuracy question is unresolved, and vendors, health systems, and regulators are actively working through what oversight should look like while adoption is still accelerating.

Why Does Clinician-in-the-Loop Review Matter So Much?

Given the accuracy concerns above, the design choice that current ambient AI scribe deployments lean on most heavily is keeping a clinician firmly in the loop. In practice, that means:

  • The AI output is explicitly framed as a draft requiring review, not an auto-signed note.
  • The clinician who conducted the visit — not an administrator or the AI vendor — remains the one who reviews, edits, and formally attests to the note’s accuracy before it becomes part of the permanent medical record.
  • Health system pilots in 2023 have generally paired tool rollout with training on what to check in an AI-drafted note and guidance on when to discard a draft entirely rather than edit it.

This human-in-the-loop model is also where some of the promised time savings can erode if oversight is done carefully: a clinician who reads an AI-generated note as closely as they would review a trainee’s note may not save as much time as one who skims and signs quickly. Health IT leaders piloting these tools in 2023 have generally been candid that the right amount of review time is still being worked out, and that cutting corners on review to maximize time savings would defeat the point of keeping a clinician in the loop at all.

What Should Health IT Leaders Watch For Before Adopting?

Organizations evaluating ambient AI scribes in 2023 are generally weighing a similar set of open questions:

  • Patient consent and disclosure. Because these tools record clinical conversations, clear patient notification and consent practices are essential, and expectations here are still being established as adoption spreads.
  • Data privacy and security. Audio and transcript data typically passes through cloud infrastructure and third-party AI models, raising the same HIPAA and data-governance questions that apply to any cloud-based health IT tool, with the added wrinkle of understanding exactly how a vendor’s underlying language model is hosted and whether patient data is used to further train that model.
  • Specialty and population fit. Early pilots have concentrated in primary care, internal medicine, and a handful of specialties such as dermatology; how well these tools perform in encounters with pediatric patients, non-English speakers, or highly technical subspecialty visits is much less established.
  • Workflow and EHR integration. A tool that produces a good draft note but does not integrate cleanly into existing EHR documentation workflows may add friction rather than remove it.
  • Realistic expectations. Framing these tools as a way to reduce documentation burden — not eliminate it, and not replace clinical judgment — appears to be the more durable expectation-setting approach emerging from 2023 pilots.

Is This Medical Advice or a Product Endorsement?

No. This article is a general overview of ambient AI documentation technology and the early evidence and concerns surrounding it as of 2023. It does not recommend any specific product, evaluate specific vendors, and is not medical, legal, or compliance advice. Health systems and clinicians evaluating these tools should conduct their own diligence, including review of vendor data-handling practices and consultation with their organization’s privacy, compliance, and clinical leadership. Readers seeking authoritative background on clinical documentation burden can consult the Office of the National Coordinator for Health Information Technology, which tracks EHR usability and burden-reduction efforts as part of its federal health IT mandate.

Frequently Asked Questions

What is an ambient AI scribe?

An ambient AI scribe is a documentation tool that uses microphones to capture the audio of a clinical conversation and generative AI to turn that conversation into a draft clinical note. Unlike traditional dictation, the clinician speaks naturally with the patient rather than dictating a structured summary afterward.

Is an AI scribe the same as a human medical scribe?

No. A human scribe is a person, often in the room or listening remotely, who documents in real time. An AI scribe uses speech recognition and a language model to generate a draft automatically, with a clinician reviewing and finalizing the note rather than a person writing it live.

Can AI-generated clinical notes be trusted without review?

Not as of 2023. Large language models can produce fluent but inaccurate text, sometimes called hallucination, so every AI-drafted note is intended to be reviewed, corrected if needed, and formally signed by the treating clinician before it becomes part of the medical record.

Do ambient AI scribes actually reduce burnout?

Early 2023 pilot data suggests reduced documentation time and generally positive clinician sentiment, but studies so far are small, short-term, and largely rely on self-reported outcomes. Larger, longer, and more rigorous research is still needed before drawing firm conclusions about burnout impact.

Clinical practices using these tools generally disclose their use to patients and obtain consent before recording a visit, similar to policies for recording with a human scribe present. Specific consent requirements can vary by state law and organizational policy.