A radiology triage algorithm, an ambient documentation assistant, a sepsis-prediction score, and a large language model summarizing discharge notes can all be running inside the same health system at the same time — approved by different departments, validated (if at all) against different data, and monitored by nobody in particular. That fragmentation is exactly the problem clinical AI governance in 2025 is trying to solve. It is no longer a theoretical exercise: with intake volume for AI tools climbing sharply year over year at many systems, and a first-of-its-kind accreditor framework now in circulation, health systems are converging on a recognizable set of governance building blocks — an intake and inventory process, a review committee, local validation before go-live, bias and equity testing, post-deployment drift monitoring, and transparency practices tied to federal certification requirements. None of this is fully standardized yet, and organizations are assembling it at different speeds and with different resources. But the shape of the emerging consensus is clear enough to describe.
Why Ad Hoc AI Adoption Stopped Being Tenable
Clinical AI arrived in most health systems the way earlier software did: department by department, often through an EHR vendor’s built-in feature or a point solution a service line champion advocated for. That pattern works reasonably well for tools with narrow, well-understood behavior. It works poorly for statistical and generative models, because their behavior is not fully specified in advance — a sepsis model or an ambient scribe can behave differently across patient populations, documentation styles, and even time, in ways that are not obvious from a vendor’s marketing material or even from a single validation study.
The accreditation and standards landscape has started to reflect that gap. In September 2025, the Joint Commission and the Coalition for Health AI (CHAI) jointly released initial guidance — the Responsible Use of Artificial Intelligence in Healthcare (RUAIH) framework — as, by their description, the first framework of its kind from a U.S. accrediting body aimed at helping healthcare organizations adopt AI safely, effectively, and ethically. It is non-binding guidance rather than a certification requirement, and the two organizations have said follow-on, more implementation-focused governance playbooks are still being developed from health-system workshops. Several published case reports from large academic medical centers and cancer centers describe formal AI governance committees now tracking dozens of models — including large language model pilots — under a single oversight structure, a scale of coordination that informal, department-level review cannot realistically provide.
The AI Governance Committee: Who Should Be in the Room
Most published governance models converge on a standing, multidisciplinary committee rather than a single department signing off on AI tools. Representation commonly includes clinical informatics, data science or analytics, compliance and legal, information security and privacy, health equity, and frontline clinical leadership from the specialties most affected by a given tool. Some organizations pair a system-level steering committee with project-specific subcommittees so that deep technical review of, say, an imaging model does not require the same reviewers as an ambient documentation tool.
The committee’s job is broader than a single yes/no approval. Recurring responsibilities described in health-system governance write-ups include:
- Setting the intake and evaluation criteria every proposed AI tool must clear before pilot or production use.
- Deciding the level of scrutiny a tool receives based on its clinical risk — a note-summarization assistant with a human reviewing every output typically warrants a lighter review than a model that can trigger an unreviewed clinical action.
- Assigning ongoing monitoring ownership, so that “who is watching this model in six months” has an answer before go-live, not after a problem surfaces.
- Reviewing and periodically re-approving tools already in production, since a model that performed well at approval can still degrade later.
Intake and Inventory: Knowing What You Already Have
A governance committee cannot govern tools it does not know exist. The starting discipline most frameworks describe is a formal intake process paired with a living inventory — a registry that captures, for every AI tool in use or under consideration, its clinical purpose, the data it was trained and validated on, its risk classification, the vendor or internal team responsible, and its current approval and monitoring status. Several health systems have reported sharp year-over-year increases in AI intake requests, driven partly by embedded AI features arriving inside existing EHR and clinical software rather than through a formal purchasing process — which is precisely the pathway an inventory is meant to catch. Shadow deployment, where a clinical team turns on a vendor’s AI feature without routing it through governance, is a commonly cited failure mode; a documented intake process gives staff an unambiguous door to walk through.
Local Validation: Vendor Performance Data Is a Starting Point, Not a Verdict
Perhaps the most consistent theme across recent clinical AI literature is that a model’s published or vendor-reported performance does not reliably predict how it will perform at a new institution. Patient populations, documentation habits, lab reference ranges, device makes, and clinical workflows all vary across sites, and a model trained or validated elsewhere can carry those differences forward as blind spots. Peer-reviewed work has gone as far as arguing that external validation studies should be treated as a starting hypothesis and that recurring local validation, not a one-time external study, is the more defensible standard for deployed clinical AI.
In practice, local validation typically means running the model against retrospective or “silent” real-time data from the deploying institution — recording what the model would have predicted without yet acting on it — and comparing those outputs to actual clinical outcomes before clinicians ever see a live score or alert. Where that silent-mode testing reveals a performance gap against the local population, some organizations recalibrate thresholds or, for adaptable models, retrain on local data rather than deploying the vendor’s default configuration as-is. This step is what turns “the vendor’s clinical validation study” into “evidence this works for our patients,” and a growing number of governance frameworks treat it as a non-negotiable gate before production release, not an optional refinement.
Bias and Equity Assessment
Because most clinical AI models learn from historical care patterns, they can inherit and even amplify inequities embedded in that history — a model trained predominantly on one demographic group, or on a dataset shaped by historically unequal access to testing or treatment, can perform unevenly once deployed broadly. CHAI’s Assurance Standards Guide names fairness, equity, and bias management as one of its core evaluation domains alongside usefulness, safety and reliability, transparency, and security, and calls for bias evaluation across the AI lifecycle rather than as a one-time pre-launch check.
Concretely, this typically means stratifying validation performance by relevant subgroups — race and ethnicity, age, sex, language, insurance status, or clinical site — rather than reporting only an aggregate accuracy figure, since a model can look strong overall while underperforming for a specific population. Equity review is also where governance committees are increasingly expected to ask harder questions about training data provenance: whose data trained this model, whether those patients resemble the population it will now be used on, and whether known gaps in historical data collection (for example, underrepresentation of certain racial groups in some clinical trial and EHR datasets) are likely to show up as blind spots in the model’s output.
Post-Deployment Monitoring and Drift
Governance does not end at go-live. A model’s performance can degrade over time even without any change to its code, a phenomenon generally called model or data drift — driven by shifts in patient population, changes in clinical practice, new lab assays or documentation templates, or even changes in an unrelated upstream system that feeds the model’s inputs. The NIST AI Risk Management Framework, organized around four core functions — Govern, Map, Measure, and Manage — frames this as a continuous lifecycle activity rather than a launch milestone: “Measure” calls for ongoing testing, evaluation, and validation of deployed models, while “Manage” addresses monitoring for incidents and drift and responding when performance shifts.
Health-system governance frameworks generally translate that into recurring, scheduled performance reviews — not just a mechanism for clinicians to report problems anecdotally — with a named owner accountable for each production model, defined performance thresholds that trigger a re-review, and a documented process for retraining, recalibrating, or retiring a model whose real-world performance has fallen out of tolerance. Several published committee case studies describe formal annual or more frequent policy and model reviews as the norm, rather than an indefinite “set and forget” approval.
Transparency and the HTI-1 Connection
Governance and regulation increasingly intersect at the question of what information clinicians and health systems are entitled to know about the AI embedded in their certified health IT. The Office of the National Coordinator’s (now the Assistant Secretary for Technology Policy, ASTP) HTI-1 final rule established what ONC has described as first-of-its-kind transparency requirements for AI and other predictive algorithms embedded in certified electronic health record technology — introducing the concept of Decision Support Interventions (DSIs) that must expose a defined set of “source attributes” describing how the tool was developed, validated, and intended to be used. ASTP’s implementation materials describe an expanded set of attributes for evidence-based DSIs and a considerably larger set specific to predictive DSIs, covering areas such as the intervention’s purpose, development details, and information about the data used to train and validate it, all intended to give health systems and clinicians a consistent baseline for judging whether a given predictive tool is fair and fit for their population.
Governance committees are a natural consumer of this information: HTI-1 source attributes are, in effect, the structured version of the same fairness and validity questions a local bias-and-validation review is trying to answer. A well-run intake process should request and retain a tool’s DSI documentation as part of its file, both to satisfy the transparency intent of the rule and because that same documentation is useful evidence when a model’s performance is later questioned.
Emerging Frameworks Health Systems Are Drawing On
No single mandatory standard governs clinical AI oversight as of 2025, so most organizations are assembling their governance model from several complementary sources rather than adopting one framework wholesale:
- CHAI (Coalition for Health AI) — a rapidly growing membership coalition of health systems, vendors, and academic institutions that has published an Assurance Standards Guide and, with the Joint Commission, the RUAIH accreditor-aligned guidance described above; CHAI has also been developing assurance-lab certification concepts and a model “nutrition label” intended to standardize how a tool’s validation and performance data are disclosed.
- NIST AI Risk Management Framework — a voluntary, sector-agnostic framework whose Govern-Map-Measure-Manage structure many health systems use as scaffolding for AI risk policy, even though it was not written specifically for clinical settings.
- ONC/ASTP certification requirements (HTI-1) — a binding regulatory floor, but one that applies specifically to DSIs embedded in certified health IT and to the developers of that certified technology, not to every AI tool a health system might use.
- Peer-reviewed institutional governance case studies — increasingly, single-institution and cancer-center governance models published in journals such as npj Digital Medicine, which document real committee structures, intake volumes, and early lessons rather than prescribing a universal template.
Because these sources overlap rather than align perfectly, health systems building a program in 2025 are largely doing integration work — mapping CHAI’s assurance domains, NIST’s lifecycle functions, and HTI-1’s transparency attributes onto a single intake form and committee charter — rather than adopting a finished, off-the-shelf governance package.
The Practical Starting Point
For an organization without a formal program yet, the recurring advice across these sources points toward a workable sequence rather than a chart, mirroring what is described in the Agency for Healthcare Research and Quality’s broader guidance on health IT safety practices: inventory what AI is already running, however informally; stand up a multidisciplinary committee with real authority to pause or reject a deployment; require local validation, including a silent-mode comparison against actual outcomes, before any tool goes live; build bias and equity review into that same validation step rather than treating it as separate; and commit, in writing, to who monitors each model after launch and what triggers a re-review. None of that requires waiting for a finished national standard — and given how quickly CHAI, ASTP, and the Joint Commission’s guidance are still evolving, health systems that wait for a settled framework before starting are likely to be governing a much larger and harder-to-untangle AI footprint by the time one arrives.
This article is intended for informational purposes for health IT and clinical informatics professionals and does not constitute legal, regulatory, or medical advice. Health systems should consult qualified compliance, legal, and clinical counsel when designing or updating AI governance programs.
Related reading
- Clinical Decision Support Without Alert Fatigue
- AI-Enabled Remote Monitoring: Smarter Alerts or More Noise?
Frequently Asked Questions
What is clinical AI governance?
Clinical AI governance is the set of policies, committees, and processes a health system uses to review, validate, monitor, and retire AI tools used in patient care — covering intake and inventory, local performance validation, bias and equity assessment, post-deployment drift monitoring, and transparency about how each tool works and was tested.
What is a clinical AI governance committee responsible for?
A clinical AI governance committee typically sets intake and evaluation criteria, classifies tools by clinical risk, approves or rejects deployments, assigns ownership for ongoing monitoring, and periodically re-reviews tools already in production to confirm they still perform as expected on the current patient population.
Why does local validation matter if a vendor already validated its AI model?
Vendor or published validation reflects performance on that vendor’s or study’s population, which can differ from a given health system’s patients, documentation habits, and equipment. Peer-reviewed research has found that AI models can perform meaningfully worse at a new site, making site-specific local validation, ideally in silent mode before go-live, a standard governance expectation.
How does HTI-1 relate to clinical AI governance?
HTI-1, from ONC/ASTP, requires certified health IT to expose “source attributes” for embedded predictive and evidence-based Decision Support Interventions, describing how each tool was developed and validated. Governance committees use this transparency data as structured evidence when assessing a tool’s fairness and fitness for local use.
What frameworks are health systems using for AI governance in 2025?
Common reference points include the Coalition for Health AI’s (CHAI) Assurance Standards Guide and its joint guidance with the Joint Commission, the NIST AI Risk Management Framework’s Govern-Map-Measure-Manage structure, ONC/ASTP’s HTI-1 transparency requirements, and published institutional governance case studies from academic medical centers.
