Why Health Systems Keep Buying AI Tools and Getting No Outcomes
I watched a Stripe dashboard cross a million dollars in a single day once. DTC company, my second one, and the number was real. Money moved. Cards cleared. I screenshotted it.
The product was bad.
Not fraudulent, just bad, and the thing nobody warns you about scale is that it doesn't fix anything, it just makes whatever you already had louder. We had a returns problem and a retention problem and a supply chain running on somebody's spreadsheet, and every one of those got worse as revenue went up, because volume finds every crack you were politely ignoring.
I think about that dashboard constantly when I read health system AI reports.
Why do most health system AI pilots never scale?
Because a pilot proves the model works and says nothing about whether the organisation does. Qventus surveyed more than 60 health system CIOs and chief AI officers for its 2026 Beyond the Pilot report and the gap comes out stark: most systems have governance in place and pilots running, and only 4% have scaled AI with measurable outcomes. Eighty percent struggle to quantify returns at all.
Ninety four percent of those same leaders say delay would put them at a competitive disadvantage.
So, near total agreement on urgency, near total absence of measured results. That isn't a technology gap. That's an operating gap, and no vendor sells a fix for it.
Is the problem the models or the workflow?
The workflow, and it isn't close. BCG's 2026 guidance on becoming an AI-first provider puts a number on the split: 20% of effort belongs on data and technology infrastructure, 10% on algorithms and use case models, and 70% on workflow redesign, role redefinition and adoption.
Most organisations invert that completely.
They spend nine months evaluating vendors and models, which is the 10%, then bolt the winner onto a process designed in 1998 around a fax machine. BCG's line is blunt and correct: AI efforts struggle to scale not because of technology limitations but because providers layer AI onto workflows never built for real time intelligence.
Ask a system to change nothing and you'll get precisely what you asked for.
Where does the friction actually live?
At the seams. A recent Becker's piece on patient access reaching an architectural inflection point describes this better than I can, so I'll tell you what it says and add my own scar tissue.
Voice AI resolves cancellations and confirmations. Self scheduling removes routine bookings before they hit a queue. Engagement platforms cut reactive inbound through reminders. Each category does its job.
Then a reminder goes out, the patient calls back with a question, and the voice agent has no record of the reminder. A cancellation lands at 9pm and the tool that could fill the slot doesn't hear about it until the morning sync. A new patient needs insurance verification, provider logic and visit type selection before anything can be booked, and no single tool holds all three.
Those aren't edge cases. Those are the exact moments that decide whether the strategy was worth anything, and each one crosses a boundary nobody designed.
You didn't buy a bad tool. You bought seven good ones that don't talk to each other.
What does vendor sprawl actually cost?
More than the licences. In the same Qventus data, 66% of leaders name limited IT resources from managing multiple AI vendors as a top three execution obstacle, and 51% spend between 11% and 25% of total IT bandwidth on vendor management, integrations and implementations alone. Seventeen percent spend up to half of it.
Only 4% say they have adequate resources and processes in place for that work.
Meanwhile 87% say tech fatigue, meaning staff exhaustion from learning new systems before seeing any value, is a barrier to adoption. Every additional tool taxes the same nurses you bought the tools to protect.
Should you wait for your EHR vendor to ship it?
Depends entirely on what waiting costs you, and the industry's answer to that has moved fast. Seventy four percent of surveyed leaders name reliance on their EHR vendor's AI roadmap as a top execution barrier. Asked whether they'd wait 18 months for an EHR feature or deploy a proven third party solution in three, only 22% would wait. The previous year, 52% said wait.
That's a big swing in twelve months.
The filter those CIOs describe is a good one. Does this touch the bottom line directly. Does it touch a mission critical patient experience. Can I survive not having it for eighteen months. If that's yes, yes and no, stop waiting. It's also why HANA is fully open source and self hosted, because a health system that can't run the thing inside its own boundary hasn't reduced dependency, it's just changed who it depends on.
Where should a health system start in 2026?
Start with the loop where the data is already accessible and one person clearly owns the outcome. For most systems that's patient access and follow up, which is why 72% of surveyed CIOs named scheduling and care access as their biggest near term opportunity.
Pick one service line. Set a baseline nobody can argue with in six months. Redesign the workflow before you deploy anything into it, and give that redesign the 70% of attention it deserves.
Then measure things that survive scrutiny. Reach rate, escalation volume, time to resolution, slots recovered. Our deployment data is published in those terms, with the underlying outcome research alongside it, because a 31:1 ROI figure means nothing unless somebody can see what went into the numerator.
One more thing, and it's the part everyone skips. Decide what a human still owns before you launch, not after the first bad call. Across our clinical use cases the escalation rules are the least sophisticated part of the entire system and the single reason we've run 1M+ patient interactions without a critical adverse event.
Key Takeaways
The scaling problem in health system AI isn't a model problem. Only 4% of systems have scaled with measurable outcomes while 94% believe delay is dangerous, and the space between those two numbers is filled with workflow debt, vendor sprawl and pilots nobody measured properly.
The systems getting results treat this as infrastructure rather than a project portfolio. They own the sequencing instead of outsourcing it to a platform vendor, they put the bulk of their effort into redesigning how work actually happens rather than into model selection, and they consolidate deliberately instead of approving every tool that demos well.
Volume doesn't fix a broken process. It never has. That Stripe dashboard was proof, and a million dollars in a day bought me nothing except a faster route to finding out what was wrong. If you want to see the cost side before committing to anything, our pricing is public.
Happy to walk through your access baseline and find where the seams are. Book a slot here.
FAQ
Why do health system AI pilots fail to scale?
The most common causes are workflow debt and measurement gaps, not model performance. Only 4% of surveyed health systems have scaled AI with measurable outcomes, and 80% report difficulty quantifying returns at all, which makes it nearly impossible to justify expansion to a finance committee.
How much effort should go into workflow redesign versus model selection?
BCG recommends roughly 70% of effort on workflow redesign, role redefinition and adoption, 20% on data and technology infrastructure, and only 10% on algorithms and use case models. Most organisations invert that ratio and then wonder why adoption stalls.
Is it better to wait for an EHR vendor's AI roadmap or buy third party?
Willingness to wait has dropped sharply. Only 22% of surveyed CIOs would wait 18 months for an EHR feature versus deploying a proven third party solution in three months, down from 52% the previous year. The practical test is whether the workflow touches revenue or a mission critical patient experience.
