Why Do 89% of Healthcare AI Agents Never Reach Production?
I spent a season crossing Australia with a circus. Actual circus, actual desert, feels like another lifetime and I still think about it most weeks.
The aerialists were extraordinary. Years of training, silks, drops that pulled your stomach into your throat. And in town after town, the biggest crowd gathered around the guy spinning fire chains on the ground, doing something you could learn in a month.
Not because the crowd was stupid. Because they could see what he was doing, instantly, without anyone explaining it to them.
I think about that every single time a health system walks me through its AI strategy.
Why do so few healthcare AI agents ever reach production?
Because most of them are aerial silk. Corti reported in February that only 11% of AI agents reach production in healthcare, against a projected shortage of 10 million health workers by 2030. Eighty-nine percent die somewhere between the pilot deck and the go-live date.
They don't die because the models are bad. They die because the workflow being automated got picked for how impressive it sounds in a board meeting, not for whether anyone can watch it working by Friday.
Isn't the answer better infrastructure?
Partly, and I'd rather people took the argument seriously than waved it off. The clinical intelligence infrastructure case is right that a sepsis model or a prior auth tool isn't a platform, and that identity resolution, terminology mapping, audit logging and workflow delivery are where the money actually goes. IKS Health puts a number on it: roughly 70% of AI success comes from workflow and orchestration rather than from the model itself.
But I've watched plenty of organisations use "we need the foundation first" as permission to spend eighteen months building something nobody has yet asked to use. Cloud data lakes run $100,000 to $1M. Middleware for legacy integration, $50,000 to $500,000. That's real money burned before a single patient notices a single thing.
Infrastructure is necessary. It's just a terrible first deliverable.
What does the fire chain version look like in a health system?
Something a clinician can watch happen and understand without a slide. Post-discharge follow-up calls. No-show recovery. Pre-op instruction confirmation. Chronic care check-ins that used to sit on a coordinator's list until Thursday and now don't.
Boring, legible, countable. You can put a number on it in three weeks. How many patients did we reach, how many escalations did we catch, how many hours came back to the care team.
I learned this lesson expensively once before, in a completely different industry. I watched a Stripe dashboard tick past a million dollars in a single day and felt like a genius, and the product underneath it was not good. Scale hides problems. It hides them beautifully, right up until the moment it stops. A pilot that looks impressive and can't be measured is doing the same trick to you, just slower and with a smaller number attached.
We've run more than a million patient interactions with zero critical adverse events, across five countries and three languages, and almost none of it looks clever from the outside. That's the point, honestly. The deployments we publish are all one workflow, done properly, then widened.
How do you tell real infrastructure from expensive sprawl?
One test. When you add the second use case, does it cost less than the first one did?
If yes, you built a platform. If every new tool means a fresh EHR integration, a fresh compliance review, a fresh vendor contract and a fresh data silo, then you bought four point solutions with a shared invoice. IKS Health describes exactly this failure mode: four integrations, four reviews, four silos, zero shared intelligence between them.
The second test is uglier. Can you turn it off? If the answer involves a committee and a renewal date, you don't have governance, you have a hostage situation. Open-source and self-hosted deployment exists partly for that reason, and it's why we built HANA that way instead of renting our own runtime from a model provider.
What should you actually fund this year?
One workflow, one owner, one number, ninety days.
Pick the workflow where the data is already accessible and the clinical owner already cares. Instrument the baseline before you launch, because a program with no pre-launch number can never defend its own budget when the budget gets defended. Then make the second deployment cheaper than the first, and the third cheaper than the second, and let the platform emerge from the sequence rather than preceding it by a year and a half.
Costing exercises tend to break where people expect them to hold, so run the arithmetic on what this actually costs against the coordinator hours you're currently burning. The technical setup takes days rather than quarters, which is itself part of the argument.
Key Takeaways
Eighty-nine percent of healthcare AI agents never reach production, and the failure is rarely technical. It's a selection problem. Organisations pick the most sophisticated available use case instead of the most legible one, then discover that legibility was the exact quality that would have carried it through governance, adoption and the budget cycle.
Infrastructure matters enormously, and it should be built as a consequence of shipping rather than a prerequisite for it. Fire chains first. The silks can wait until people trust you enough to look up.
FAQ
Why do most healthcare AI pilots fail to scale? The common failure is choosing use cases by impressiveness rather than by measurability and workflow fit. Programs without a pre-launch baseline, a named clinical owner and a result visible inside one quarter tend to stall at pilot no matter how good the model is.
Should a health system build AI infrastructure before deploying use cases? Build them together, weighted heavily toward shipping. Start with one workflow where data access and clinical ownership already exist, then let each subsequent deployment reuse the integration, governance and monitoring you built for the first one.
How do you measure whether an AI deployment is working? Pick one operational number and one clinical number before launch. For patient outreach that usually means reach rate and escalations caught, alongside downstream measures like readmission or no-show rate. If the metric can't be read off inside ninety days, it's the wrong metric.
If you want to argue about which workflow is your fire chain, grab a slot and bring your volumes.
