Two years ago, “AI in the EHR” mostly meant one narrow use case: an ambient tool listening to a visit and drafting a note. In 2025, that framing looks dated. Large language models are now woven into far more of the daily workflow — drafting replies to patient portal messages, condensing sprawling charts into readable summaries, suggesting billing codes, and still writing notes from ambient audio. The common thread is a large language model (LLM) sitting between raw clinical data and a human who has to act on it, generating a first draft a person is expected to check.

That expectation — that a clinician always reviews before signing — does a lot of quiet work in the industry’s safety case. It’s also the piece most worth scrutinizing, because the evidence so far shows review happens unevenly, and the tools still make mistakes that are easy to miss under time pressure. This piece walks through where generative AI has actually landed in EHR workflows as of 2025, what the early data says about accuracy and adoption, who’s partnering with whom, and the governance questions — including a federal transparency rule — shaping how these tools get scrutinized.

What Does “Generative AI in the EHR” Actually Mean in 2025?

The phrase covers a handful of distinct workflows that share an underlying technology (large language models, often the same foundation models used elsewhere) but differ a lot in stakes and maturity:

  • In-basket message drafting — an LLM reads an incoming patient portal message and a snippet of the chart, then proposes a reply for a physician or care team member to edit and send.
  • Ambient documentation (“AI scribes”) — an LLM turns a recorded or transcribed patient encounter into a structured clinical note.
  • Chart and encounter summarization — an LLM condenses a long history, a hospital stay, or a stack of prior notes into a shorter summary for a clinician about to see the patient.
  • Coding and documentation-improvement assistance — an LLM reads a note and suggests ICD-10 or CPT codes, or flags where documentation doesn’t support the level of service billed.

Each inserts a generative model at a different point in the chart, with different consequences if it gets something wrong. A hallucinated detail in a discharge summary a specialist skims for two minutes carries different risk than a hallucinated detail in a coded diagnosis that reaches a payer. That distinction matters more than the shared “generative AI” label suggests, and it’s a big reason vendors and health systems apply different levels of human review to each use case.

How Are In-Basket Drafting Tools Actually Performing?

In-basket drafting has the most publicly available outcome data so far, largely because patient portal message volume ballooned after 2020 and health systems were eager for anything that might blunt inbox burden. The picture that’s emerged is more mixed than early vendor demos suggested.

A retrospective study of EHR audit logs at a large New York City health system, examining how 75 clinicians used AI-drafted patient message replies over roughly a year, found utilization of the AI draft was modest — around 19% of eligible messages — though it climbed as the health system refined its prompts, and was highest among physicians rather than other care team roles (npj Digital Medicine). Drafts shaved a few percentage points off message turnaround, even with a small increase in the clicks needed to finalize a reply. Physicians and support staff also wanted different things from the draft: physicians preferred terse, information-dense replies, while other roles favored warmer language — a reminder that “good AI output” isn’t one target.

A broader systematic review pooling 23 studies on generative AI for patient message drafting concluded that AI-drafted replies were often comparable in quality and empathy to human-written ones, but flagged inconsistent performance and real, unresolved concerns about patient safety and oversight when a draft goes out with insufficient review (npj Health Systems).

The takeaway for 2025: in-basket drafting reduces some friction, but adoption is uneven, output quality depends heavily on prompt design and role, and it hasn’t yet delivered the dramatic burden relief some early pilots implied.

What Is Chart Summarization Good and Bad At?

Chart and encounter summarization is arguably the highest-volume generative AI use case in the EHR, because almost every clinical workflow involves reading through prior documentation before deciding what to do next. The research on how well LLMs do this job, however, cautions against treating summaries as a substitute for reading the underlying chart.

A scoping review of large language models in clinical documentation found accuracy varies sharply by task complexity: LLMs do reasonably well summarizing standardized, template-driven documents like routine discharge instructions, but accuracy drops for complex, high-acuity cases such as neurosurgical operative reports, where factual errors were more common than in human-authored text. One evaluation of LLM-generated emergency department encounter summaries found only about a third of GPT-4 summaries were completely free of any error across every category assessed. Another framework for quantifying clinical safety in medical text summarization measured a roughly 1.5% hallucination rate and a 3.5% omission rate across nearly 13,000 clinician-annotated sentences — numbers that sound low until multiplied across the summaries a large health system generates in a single day.

The recurring failure modes are fabricated details (a lab value, medication, or diagnosis never in the source record), omissions (a real and sometimes important detail left out), and distortions (a detail present but subtly mischaracterized — a “possible” finding rendered as definite). None are exotic; they’re the ordinary failure modes of language models applied to a domain where a missed detail can matter far more than in casual use.

What About Coding Assistance and Ambient Notes?

Coding and Documentation-Improvement Assistance

Generative AI has also moved into revenue-cycle and coding workflows, where LLMs and natural language processing read unstructured notes, operative reports, and discharge summaries to suggest ICD-10, CPT, and HCPCS codes, or to flag clinical documentation improvement (CDI) opportunities where documentation doesn’t fully support the billed level of care. The computer-assisted coding segment captured the largest share of the AI-in-medical-coding market in 2025.

Two coding-specific caveats are worth flagging. First, most deployed tools are explicitly positioned as computer-assisted coding — a suggestion layer a human coder reviews and finalizes — rather than fully autonomous coding, and vendors that claim high automation still build in human audit sampling. Second, coding errors carry a different risk profile than a clinical documentation error: an incorrect or unsupported code can trigger a payer audit, a compliance finding, or false claims exposure, which is why CDI and compliance teams, not just clinicians, are typically part of the review loop.

Ambient Notes: Still the Anchor Use Case

Ambient AI scribes — tools that listen to (or transcribe) a patient encounter and generate a structured note — remain the most mature and widely deployed generative AI application in the EHR, and by 2025 had moved past pilot status at many large health systems. Industry estimates put the number of vendors in the ambient scribing space at 60 or more as of early 2025, with projections that a majority of practices could be using some form of ambient documentation by year’s end. Reported outcomes at specific institutions include per-encounter time savings and, in some cases, improvements in self-reported clinician well-being, though the evidence still leans heavily on health-system-reported figures and observational studies rather than large randomized trials. A narrative review flagged a caution that applies broadly: implementations report real efficiency gains alongside a persistent rate of documentation omissions and occasional clinically significant hallucinations, particularly in complex encounters.

Because ambient documentation was the first generative AI use case to reach scale in the EHR, it also established the review-and-sign workflow pattern — clinician reviews and edits before finalizing — that other use cases inherited.

Which EHR-AI Partnerships Are Shaping the Market?

A handful of partnerships illustrate how EHR vendors are bringing generative AI into their platforms — largely by partnering rather than building every model from scratch, then layering their own workflow integration on top.

Epic’s most visible partnership has been with Microsoft, which brought Azure OpenAI Service models into Epic’s platform for uses including drafting patient message replies and natural-language chart queries, alongside Nuance’s Dragon Ambient eXperience (DAX) Copilot, which reached general availability embedded in Epic’s clinical workflow. Epic has separately deepened integration with independent ambient-documentation vendors such as Abridge through its Epic Workshop developer program, and by 2025 had begun rolling out its own native ambient scribe capability. Other major EHR vendors have pursued comparable strategies, pairing in-house AI development with specialty-vendor partnerships for ambient documentation, in-basket drafting, and coding assistance.

These examples illustrate the broader market pattern — vendor partnerships accelerating feature delivery — rather than endorsing any specific product; the underlying models, integration depth, and safety review vary by deployment, and health systems should treat vendor claims about accuracy and time savings as a starting point for their own validation, not a substitute for it.

What Are the Governance and Safety Concerns?

The central governance tension is straightforward: these are probabilistic text-generation systems inserted into workflows that assume a human will catch their mistakes, and the evidence above shows that assumption doesn’t always hold. A few concerns keep surfacing across the research:

  • Hallucination and omission risk. Summarization and drafting tools fabricate, omit, or distort clinical details at rates that are individually small but meaningful at scale, and errors cluster in complex cases — precisely where clinicians most need a reliable summary.
  • Uneven review in practice. Studies of in-basket drafting found AI-suggested replies aren’t always reviewed with the same scrutiny as a blank-page draft, partly because a plausible-looking, well-formatted suggestion can create a false sense of confidence — a dynamic sometimes called automation bias.
  • Liability and accountability. When an AI-assisted note, reply, or code contains an error that reaches a patient or payer, questions about who is accountable — the clinician who signed it, the health system that deployed the tool, or the vendor that built it — remain unsettled in most jurisdictions.
  • Training and trust gaps. National physician surveys have found a majority of doctors now use AI tools in some form, most commonly for documentation and research summarization, but also report wanting more structured training on safe use, and identify clear liability frameworks and independent safety validation as prerequisites for greater trust (American Medical Association).

None of this means generative AI shouldn’t be used in the EHR — the efficiency gains in several of these studies are real. It does mean the tools are early enough in their evidence base that “the AI drafted it” should prompt more scrutiny, not less, particularly for anything touching a diagnosis, a medication, or a billing code.

How Does the HTI-1 Rule’s Transparency Requirement Fit In?

One concrete regulatory response comes from the Assistant Secretary for Technology Policy (ASTP, formerly the Office of the National Coordinator for Health IT) and its HTI-1 final rule, which took effect in 2024 and updated certification requirements for EHR technology. HTI-1 requires certified EHR technology to expose a defined set of “source attributes” for decision support interventions (DSIs) — plain-language information about how a predictive or AI-driven feature works, what data it was developed and validated on, and its known performance limitations — so a health system can judge whether it is fair, appropriate, valid, effective, and safe, a standard the rule shorthands as FAVES (HealthIT.gov).

HTI-1’s DSI transparency requirements were written with predictive algorithms — risk scores, deterioration alerts, and similar tools — most explicitly in mind, and it remains an open question how cleanly that framework extends to generative, LLM-based features like in-basket drafting or summarization, which produce open-ended text rather than a single risk score. This is also a moving regulatory target: subsequent ASTP rulemaking has proposed narrowing some AI-transparency requirements introduced under HTI-1, so health systems and vendors should verify current requirements against ASTP’s latest guidance rather than treating the original 2024 rule as the final word. The larger point holds regardless of the exact rule text: transparency about how an AI feature was built and validated — not just how well it performs in a vendor’s demo — is the direction regulators, health systems, and physician organizations are all pushing toward.

The Bottom Line for 2025

Generative AI has moved from a single flashy use case to a set of tools embedded at several points in the EHR workflow, and the honest 2025 assessment is “genuinely useful, still imperfect, unevenly reviewed.” In-basket drafting saves some time but isn’t used as often as hoped. Chart summarization works well on routine documents and worse on complex ones. Coding assistance accelerates revenue-cycle work but is deployed as a human-reviewed suggestion layer, not an autonomous system. Ambient scribes remain the most mature use case but still carry a measurable hallucination and omission rate. And the governance framework meant to keep all this transparent and accountable is itself still being written and rewritten. None of this is medical advice or a product recommendation; it’s a snapshot of where the evidence and the regulatory conversation stood as of 2025, and both are likely to look different a year from now.

Frequently Asked Questions

Is generative AI in the EHR the same thing as an ambient AI scribe?

No. Ambient scribes are one application — turning a recorded encounter into a note — but generative AI in the EHR also includes in-basket message drafting, chart and encounter summarization, and coding assistance. All four rely on similar large language model technology but serve different workflows with different risk profiles.

How accurate are AI-generated chart summaries?

Accuracy varies by complexity. Research shows LLMs perform reasonably well on standardized documents like routine discharge summaries but less reliably on complex cases; one emergency department study found only about a third of AI-generated summaries were completely free of any error across all categories assessed.

Do clinicians actually use the AI-drafted patient message replies?

Usage has been lower than early enthusiasm suggested. One large health system study found roughly 19% of eligible in-basket messages actually used the AI-generated draft, improving somewhat as the health system refined its prompts, with physicians adopting it more than other care team roles.

What is the HTI-1 rule and why does it matter for AI in the EHR?

HTI-1 is a federal EHR certification rule from ASTP (formerly ONC) that requires certified health IT to disclose plain-language “source attributes” about AI-driven decision support features, so health systems can assess whether a tool is fair, appropriate, valid, effective, and safe before relying on it.

Can AI fully automate medical coding without a human reviewing it?

Most deployed generative AI coding tools function as computer-assisted coding — suggesting codes for a human coder to review and finalize — rather than fully autonomous systems. Vendors advertising high automation levels still typically build in human audit sampling given the compliance risk of coding errors.