Why Do 93% of Health Systems Deploy AI but Only 4% Have a Strategy?
Years ago I watched a Stripe dashboard cross a million dollars in a single day.
I should have been elated. I remember feeling something closer to nausea, because I knew the product was bad. Not catastrophically bad. Just bad enough that the returns were coming, the support tickets were coming, and none of it was visible yet because volume was drowning it out. Scale doesn't fix problems. It hides them, then hands them back to you with interest.
I thought about that dashboard when I read the new Center for Connected Medicine and KLAS research on health system AI. Ninety-three percent have deployed third-party AI. Sixty-three percent describe their AI strategy as developing or ad hoc. Four percent call it advanced.
Four.
What does it mean that 93% have deployed AI but only 4% have a strategy?
It means procurement is outrunning governance. Health systems are buying faster than they can validate, and the gap between those two speeds is where the risk lives. The same research found that while 92% of organizations test AI before deployment, fewer than half have a dedicated data platform or sandbox to do that testing in. So the testing is happening. It's just happening in slide decks and vendor demos rather than against your own de-identified patient population.
That's not a technology problem. That's a million-dollar-day-with-a-bad-product problem.
Why is the bottleneck infrastructure rather than tools?
Because a model is only as good as the data it can reach and the pipeline that carries the result into a workflow. The CCM and KLAS respondents named their real obstacles plainly: manual spreadsheet workarounds, metrics that don't agree across departments, unstructured clinical data. None of that is fixed by buying a better algorithm. UPMC's response was to build a real-world data platform called Ahavi to evaluate third-party models in silico before they touch live care, which is roughly the most sensible thing I've read from a health system this year.
Most organizations can't do that yet. Which means most organizations are running vendor claims straight into production and calling the first quarter a pilot.
What should a health system validate before it deploys a patient-facing AI?
Four things, and they're all unglamorous. What the model does when it's wrong. Who sees the escalation and how fast. Whether the output lands in a workflow a human already uses or in a new queue nobody opens. And whether you can audit any single interaction six months later when someone asks.
We publish our outcome data and safety methodology because those are the questions that actually get asked in procurement, and vague answers should end the conversation. Over a million patient interactions, zero critical adverse events, 85% weekly engagement against a 15-20% industry baseline. Those numbers mean nothing without the escalation architecture underneath them, and that architecture is the product.
Does the ROI measurement gap have a fix?
Yes, and it's narrower than people want it to be. Pick one workflow with a countable outcome and instrument it before you launch. The Qventus CIO research on operationalizing AI found that around 80% of health system leaders struggle to quantify AI returns at all. That's not because AI has no return. It's because nobody wrote down the baseline before switching the thing on.
No-show rate on a specific service line. Thirty-day readmission on a specific DRG. Time from discharge to first successful patient contact. One number, measured for eight weeks before, eight weeks after. Do that and the ROI question answers itself in a way a board can read. Skip it and you'll be arguing about attribution for two years. Our pricing model is built around that logic on purpose, because a per-workflow structure is the only one you can honestly attribute anything to.
Why does self-hosted matter for this?
Governance is easier when the data doesn't leave. That's the whole argument. If you're a system with a mature data platform and a real security posture, the last thing you want is patient conversations flowing through a vendor's cloud on a black-box dependency you can't inspect or replace. HANA is fully open-source and self-hostable, with no OpenAI dependency, and you can read exactly how it deploys inside your own environment in the technical documentation. Some of that decision was philosophical. Most of it was that health system CISOs kept asking the same question and I got tired of a bad answer.
Key takeaways
The industry's AI problem has shifted from access to discipline. Everyone can buy a model now, which means the differentiator is validation infrastructure, escalation design, and the willingness to write down a baseline before you launch. Four percent of health systems describe their strategy as advanced, and that number will not improve through more procurement. It improves by narrowing scope, measuring one thing properly, and treating vendor claims as hypotheses rather than findings.
My daughter is ten. She told me once, in the middle of an argument about screen time, that she works for herself and not for me. She's right, and I say the same thing to my team and to every health system I work with. Nobody works for the technology. If the tool isn't visibly serving the people using it inside a quarter, kill it.
FAQ
How do we test a voice AI vendor without a sandbox environment? Run a shadow deployment on a low-risk workflow with full transcript review for the first two weeks. Appointment reminders or post-visit check-ins are the standard starting points because the harm surface is small and the volume is high enough to read signal quickly.
Should AI governance sit with IT or with clinical leadership? Both, with a named executive owner who is accountable for outcomes rather than uptime. The systems doing this well have a single senior sponsor, a review committee with clinical representation, and a registry of every AI feature in production with its current value metric attached.
What does a realistic first year of patient-engagement AI look like? One workflow, one service line, one measured outcome for the first quarter. Expansion to adjacent workflows in quarters two and three, once escalation patterns are understood and staff trust the handoffs. You can read how existing deployments sequenced it, or if you'd rather talk it through directly, book time with me.
