A hospitalist finishes rounds and opens the next chart. A small colored flag sits at the top: elevated sepsis risk, score 7, generated eleven minutes ago from vitals, labs, and nursing documentation nobody explicitly asked the system to analyze. The flag might be the earliest warning of a life-threatening complication. It might also be one of hundreds fired that week on patients who were never septic. The clinician has no way to tell which just from the number, and that gap between what a risk score claims to know and what it actually knows is the central problem in the predictive analytics healthcare has spent the last decade building into its electronic health records.
These models did not appear overnight. Sepsis early-warning scores, 30-day readmission predictors, and inpatient deterioration indices are now standard features in most major EHR platforms, running continuously in the background of hospital operations. Vendors and health systems built them because the underlying promise is real: EHRs already contain the vital signs, labs, and documentation that precede clinical deterioration, often hours before a human notices the pattern. The question that became unavoidable in 2021, after a high-profile external validation of one widely deployed sepsis model found serious performance gaps, is whether the models health systems have already deployed actually deliver on that promise, and what it takes to know the difference.
What Predictive Analytics Actually Do Inside the EHR
Predictive analytics in this context are statistical or machine-learning models that ingest structured EHR data in something close to real time and output a risk score or alert. The three most common categories in production today are:
- Sepsis prediction models, which scan vital signs, white blood cell counts, lactate, and other labs on a rolling basis to flag patients who may be developing sepsis before it is clinically obvious
- 30-day readmission risk models, which run near discharge using diagnosis codes, prior utilization, social factors when available, and lab trends to flag patients likely to bounce back to the hospital
- Inpatient deterioration scores, such as early-warning scores that continuously recalculate a composite risk from vital signs to catch patients trending toward a rapid response or ICU transfer
Each model type shares a basic architecture: a set of predictor variables, a defined outcome window, a scoring threshold that triggers an alert, and a delivery mechanism inside clinical workflow (a dashboard flag, a best-practice advisory, a nursing task). What varies enormously — and determines whether the model helps or harms — is how rigorously that architecture was validated before it influenced care.
How These Models Get Built and Validated
Model development in this space generally follows a derivation-then-validation sequence, though how faithfully organizations follow it varies widely. A model is derived on a historical dataset, where developers select predictors, fit a statistical or machine-learning algorithm, and tune it to distinguish patients who had the outcome from those who did not. Internal validation follows, typically on a held-out slice of the same dataset or the same population in a later time period.
The step far less consistently done well is external validation — testing the model on a population it was never trained on, ideally at a different institution with a different patient mix, documentation style, and care pattern. A model can perform excellently in internal validation and still fail externally, because it may have learned patterns specific to the derivation hospital’s coding habits, order sets, or demographics rather than the underlying clinical signal. This is precisely the failure mode that became a national conversation in 2021.
Two statistical properties matter most when judging any of these models, and both are frequently underreported:
- Discrimination, usually reported as the area under the receiver operating characteristic curve (AUC or AUROC), which measures how well the model ranks patients by risk. An AUC of 0.5 is no better than chance; 1.0 is perfect separation. Most clinically useful models fall somewhere between 0.70 and 0.85, though thresholds vary by task.
- Calibration, which measures whether a predicted probability actually matches observed outcome frequency — if the model says a patient has a 20% risk, do roughly 20% of similarly scored patients actually experience the outcome? A model can have acceptable discrimination and still be badly miscalibrated, which is exactly the combination that undermines clinical trust once alerts start firing on patients who never develop the predicted condition.
The 2015 TRIPOD statement (Transparent Reporting of a multivariable prediction model for Individual Prognosis Or Diagnosis) is the most widely cited reporting framework for this work, and it explicitly calls for both discrimination and calibration to be reported, along with a clear description of the derivation cohort, handling of missing data, and any external validation performed. Proprietary vendor models embedded in commercial EHRs have not always been held to that level of published, peer-reviewed transparency — part of what made the 2021 external validation controversy so consequential.
The Epic Sepsis Model: A Cautionary Case
The Epic Sepsis Model (ESM) is a proprietary sepsis-prediction tool built into the Epic EHR and deployed at hundreds of U.S. hospitals, generating a continuously updated risk score intended to flag developing sepsis before clinical signs become obvious. In June 2021, Wong and colleagues published an independent external validation of the ESM in JAMA Internal Medicine — the first large, peer-reviewed look at how the model performed outside Epic’s own reported figures.
The retrospective cohort study examined 27,697 patients across 38,455 hospitalizations at Michigan Medicine between December 2018 and October 2019, of which sepsis occurred in about 6.6% of hospitalizations. The findings, as reported in the study, were markedly worse than clinicians using the tool may have assumed:
- Hospitalization-level discrimination was reported at an AUC of 0.63 (95% CI, 0.62–0.64) — well short of what would typically be considered strong discrimination for a high-stakes clinical alert, though time-horizon-level AUCs were somewhat higher, in the 0.72–0.76 range
- At the vendor-recommended alert threshold, the model reportedly identified only about a third of sepsis cases (sensitivity of roughly 33%), meaning it missed a majority of the patients it was meant to catch
- Positive predictive value at that same threshold was reported around 12%, meaning most alerts fired on patients who did not go on to develop sepsis
- The model reportedly generated alerts on a substantial share of all hospitalized patients — cited at roughly 18% — which the authors linked to meaningful alert-fatigue risk
- Calibration was described as poor across the time horizons evaluated, meaning the numeric risk scores did not reliably track actual sepsis probability
An accompanying editorial in the same issue, titled “The Epic Sepsis Model Falls Short—The Importance of External Validation,” argued a broader point: proprietary clinical algorithms deployed at scale, particularly ones embedded automatically in widely used EHR platforms, warrant the same independent, peer-reviewed external scrutiny expected of a new drug or device, rather than relying primarily on vendor-reported performance. Worth noting: this was one external validation at one health system, using that system’s own definition of sepsis onset and its own patient population; the authors cautioned against over-generalizing a single-site result to every ESM deployment nationally, even as the study prompted numerous hospitals to reassess their own sepsis alerting.
What the Controversy Changed
The practical effect of the JAMA Internal Medicine findings was to accelerate a conversation health IT leaders were already having: that a model’s reported performance at the vendor level, or even at its original derivation site, does not guarantee comparable performance against a different hospital’s patient mix, documentation patterns, and outcome definitions. Several health systems that had adopted the ESM reported locally re-validating or re-tuning their thresholds in response, and the episode became a widely cited reference point in subsequent discussions of algorithmic transparency requirements for commercial clinical software.
Calibration, Local Validation, and Ongoing Monitoring
The ESM case reinforced a lesson that model-validation researchers had been emphasizing for years: a predictive model is not a one-time purchase, it is an ongoing clinical intervention that requires the same lifecycle discipline as a new drug protocol. Health systems considering or currently running an embedded EHR risk model generally need, at minimum:
- Local validation before go-live, testing the model’s discrimination and calibration against the institution’s own historical data rather than relying solely on vendor-reported or another site’s published metrics
- A defined alert threshold chosen deliberately for the local population’s outcome prevalence and clinical tolerance for false positives, rather than accepting a vendor default that may reflect a very different case mix
- Periodic re-validation, since case mix, documentation habits, and even the EHR’s own underlying data structure can drift over time — a phenomenon sometimes called “model decay” — such that a model that performed acceptably at go-live may no longer be well calibrated a year or two later
- Outcome tracking after deployment, comparing alert-triggered interventions against actual outcomes to confirm the model is producing clinical benefit and not just alert volume
Bias and Fairness Concerns in Risk Models
Calibration and discrimination describe how well a model performs on average across a population. They do not, on their own, reveal whether a model performs equitably across different patient subgroups — and a widely cited 2019 study in Science by Obermeyer and colleagues showed why that distinction matters. The study examined a different type of commercial healthcare algorithm, one used to identify patients for extra care-management resources, and found it exhibited significant racial bias: at a given risk score, Black patients were considerably sicker than white patients, because the model had been trained to predict healthcare costs as a proxy for health need. Because less money was historically spent on care for Black patients in the underlying data, the cost-based proxy systematically underestimated their actual clinical need. Retraining the algorithm on a more direct measure of health status meaningfully reduced the disparity.
That study was not about a sepsis or readmission model specifically, but the underlying mechanism generalizes directly to EHR-embedded predictive analytics: any model trained on historical data inherits whatever inequities are baked into that data, whether through proxy outcome variables, unequal documentation practices, or differential access to the diagnostic tests that feed the model’s inputs. A sepsis model trained largely on data from one demographic mix, or a readmission model that uses utilization history as a predictor in a population with unequal access to outpatient follow-up, can encode those disparities into its scores even when race is never used as an explicit input variable.
This is why bias assessment is increasingly treated as a distinct validation step, not a byproduct of checking overall AUC. A model can show strong aggregate discrimination and calibration while still performing meaningfully worse — missing more true cases, or over-alerting — for a specific subgroup, and that gap is invisible unless someone explicitly tests for it by stratifying performance metrics across race, ethnicity, sex, age, insurance status, and other relevant subgroups before and after deployment.
Governance: Who Should Own These Models
The technical validation questions above only matter if a health system has a governance structure that requires them to be asked before a model goes live and revisited after. AHRQ’s Patient Safety Network has highlighted algorithmic and AI-driven clinical tools as an area requiring dedicated ethics and governance attention, distinct from traditional clinical decision support oversight, precisely because these models can scale a flawed assumption across an entire hospital system in a way a single clinician’s error cannot.
A functioning governance process for predictive models embedded in the EHR generally includes:
- A multidisciplinary review body — clinical informatics, biostatistics or data science, frontline clinicians who will act on the alerts, and increasingly a health-equity or bias-review function — with authority to approve, modify, or decline to deploy a model, rather than treating vendor-bundled analytics as pre-approved
- Transparency about model logic and training data sufficient for reviewers to assess plausible sources of bias, even when the underlying algorithm is proprietary
- A defined threshold-setting and re-validation cadence, so a model’s alert cutoff is a local clinical decision, not a fixed vendor default left unexamined for years
- A retirement or suspension pathway for models that external validation, local monitoring, or subgroup analysis shows are underperforming, mirroring the way a hospital would respond to a medication or device with a safety signal
None of this is a reason to abandon predictive analytics in the EHR. Early-warning and risk-stratification tools have a genuine, evidence-supported role in surfacing signal that busy clinicians might otherwise miss. But the ESM external validation made clear that “the EHR flagged it” is not, by itself, evidence that a model works — and that the health systems getting real value from these tools treat validation, calibration checks, bias auditing, and governance as permanent operational functions rather than a one-time step before go-live.
This article is for general informational purposes only and does not constitute medical, clinical, or health IT implementation advice. Health systems evaluating or governing predictive analytics tools should consult their own clinical, biostatistical, compliance, and legal teams.
Related reading
- Clinical Workflow Optimization: Mapping and Redesigning Care Processes
- Hospital at Home: Delivering Acute Care Outside the Hospital
Frequently Asked Questions
What is predictive analytics in an EHR?
Predictive analytics in an EHR refers to statistical or machine-learning models that continuously analyze existing patient data — vitals, labs, diagnoses, documentation — to generate a risk score for an outcome such as sepsis, hospital readmission, or clinical deterioration, typically surfaced to clinicians as an in-workflow alert or dashboard flag.
What did the JAMA Internal Medicine study find about the Epic Sepsis Model?
The 2021 external validation, based on roughly 38,000 hospitalizations at one health system, reported an AUC around 0.63, sensitivity near 33% at the recommended threshold, positive predictive value around 12%, and poor calibration — findings that suggested the model performed considerably worse than many users may have assumed and generated substantial false-positive alert volume.
Why does external validation matter for clinical risk models?
A model can perform well on the data it was built and tested on yet fail elsewhere because it has implicitly learned patterns specific to that site’s patient mix, documentation habits, or outcome definitions. External validation at a different institution is the main way to confirm a model generalizes rather than simply fitting its original environment.
How can EHR predictive models introduce bias?
Models trained on historical data can inherit real-world inequities, especially when they rely on proxy variables like healthcare cost or utilization instead of direct measures of clinical need. A 2019 study found exactly this pattern in a commercial care-management algorithm, where using cost as a stand-in for health need disadvantaged Black patients, illustrating a bias risk relevant to any EHR-based prediction model.
Who should be responsible for validating and governing these models?
Effective governance typically involves a multidisciplinary committee spanning clinical informatics, biostatistics, frontline clinicians, and health-equity review, with the authority to require local validation, set alert thresholds deliberately, monitor performance and subgroup fairness over time, and suspend or retire underperforming models rather than treating vendor defaults as permanent.
