All posts
AI Infrastructure
Hana Health
September 1, 2026

Why Do Most Healthcare AI Pilots Never Reach Production?

Matteo

In a previous life I watched a Stripe dashboard cross a million dollars in a single day. DTC company, mine, and I remember sitting very still because it felt like the number might notice me and go away.

The product was bad.

Not fraudulent, not broken, just bad. Thin margins, returns we couldn't explain, a support queue growing faster than the revenue line. Scale didn't fix any of that. Scale hid it, which is worse, because hidden problems keep compounding while everyone in the room is looking at the good chart. It took us about eighteen months to find out how expensive that was.

I think about that a lot when health systems show me their AI pilot count. Fourteen pilots. Twenty-two pilots. Great chart. How many are in production?

Why do AI pilots stall before production?

Because most organisations buy AI as a point solution and then discover it needed to be infrastructure. A 2026 blueprint on clinical intelligence infrastructure frames the failure modes precisely: models with weak data inputs, pilots that never reach production, governance reviews that start too late, and clinician-facing tools that add clicks instead of removing them.

That last one is the killer. If your AI adds a click, nurses will route around it within a fortnight, and no steering committee on earth will save it.

The pilot succeeds because a pilot is a demo with a champion. Production has no champion. Production has an on-call rota, a Sunday night, and someone who wasn't in the procurement meeting.

What does treating AI as infrastructure actually mean?

It means the boring parts get budget before the model does. Patient identity resolution across fragmented records. Terminology mapping. Audit logs that are complete rather than mostly complete. Decision rights for when a model gets limited or switched off, written down before anyone needs them.

Nobody gets promoted for terminology mapping. Everybody gets fired when it's missing.

The systems getting real returns aren't the ones with the best model access. They're the ones who decided ownership early, wired the thing into an existing workflow instead of beside it, and instrumented the baseline before launch so they could prove impact later. We publish how our platform integrates for exactly this reason, because the integration question is the real question and everyone pretends it's the last one.

Does the evidence actually support scaling these programmes?

Sometimes yes, sometimes emphatically no, and telling the difference is the entire job. Take remote monitoring. A randomized trial across 19 hospitals published in JAMA Network Open in June 2026 tested four remote monitoring strategies after sepsis and respiratory infection hospitalisation. Days at home at 90 days were identical to usual care in every arm. Among patients 65 and older, the monitoring arms did worse.

CMS reimburses this. It's a funded category. And the best available trial says it doesn't do the thing.

Now compare that to a nine-hospital study in npj Digital Medicine where virtual-nursing-supported discharge produced 30-day emergency readmission rates of 3.7% against 13.3% for traditional discharge, matched on baseline risk. Both are technology-enabled care at a distance. One works. The difference isn't the technology, it's whether a human interaction actually happened at the right moment.

How should a health system sequence the investment?

Start where the intervention is proven and the workflow already exists. Post-discharge follow-up, care gap closure, pre-procedure prep. Not because they're exciting, but because you can prove them inside two quarters and proof is the currency that buys the next thing.

Set the baseline before you launch. Reach rate, response latency, readmission rate, no-show rate. If you don't instrument the before, you will spend the following year arguing about attribution with a finance team who is, annoyingly, correct to ask.

Then pick the deployment model deliberately. We built HANA fully open-source and self-hosted with no dependency on a single model vendor, which sounded ideological until the first health system asked where the PHI goes and we could answer in one sentence. The story of why we made that choice is less noble than it sounds. It was mostly about being able to sleep.

What does good look like once it's running?

Boring, mostly. That's the tell.

Across five countries and three languages we're past a million patient interactions with zero critical adverse events, and the reason isn't clever prompting. It's scope. The agent gathers, documents and escalates. It doesn't diagnose. Engagement sits around 85% weekly against a 15 to 20% industry baseline, and clinics land near 31:1 on ROI, which you can pick apart in the deployment case studies rather than accept from a founder on the internet.

The circus taught me the shape of this. I crossed the Australian desert with one, and the acts that drew crowds were never the technically hardest. Fire chains beat aerial silk every single night. Silk is harder, more beautiful, more impressive to other performers. Fire chains are legible from thirty metres away while you're holding a beer.

Most healthcare AI is aerial silk being sold to people holding beers.

Key Takeaways

The gap between pilot and production is not a technology gap, it's an ownership and instrumentation gap, and it shows up as governance arriving late and clinicians routing around tools that cost them clicks. Fix the boring layer first: identity, terminology, audit, decision rights, and a named owner who can turn the thing off.

Be genuinely evidence-led about which use cases you scale, including when the evidence is inconvenient. The 2026 remote monitoring trial is a gift to anyone building a serious programme, because it draws a hard line between passive data collection and actual contact. Passive monitoring of older patients after sepsis reduced days at home. A discharge conversation inside seven days reduces readmission odds by around a fifth. Same budget line, opposite results.

Start narrow, measure the baseline, and treat the cost model as an operations question rather than a software one. The clinical use cases worth automating first are always the repetitive ones your staff already resent.

FAQ

Should we build or buy clinical voice infrastructure? Buy the conversation layer, own the data and the deployment. Building voice orchestration from scratch takes eighteen months you don't have, but renting a black box that holds your PHI creates a governance problem you'll be explaining to your board for years.

How long before a post-discharge programme shows measurable impact? Roughly one to two quarters if you instrumented the baseline. Reach rate moves in week one, follow-up appointment attendance in month one, readmission rate needs a full 90-day cohort before anyone should believe the number.

What's the single most common reason these programmes fail? Nobody owns it after go-live. The pilot had a champion with a budget and a deadline, and production has a rota. Name the owner and the escalation path before the first call goes out.

If you're sequencing an AI programme for 2027 and want a second opinion from someone who has broken this several times, book a slot with me.