All posts
Insights
Hana Health
August 26, 2026

The Healthcare AI Vendor Claims Nobody Can Verify (And What We Do Instead)

I remember watching a Stripe dashboard cross a million dollars in a single day. Direct-to-consumer company, my second one. Champagne, screenshots, the whole thing.

The product underneath that number was actually kind of bad.

Scale doesn't fix a bad product. It just hides it faster, and for longer, until the day it can't anymore.

I think about that dashboard every time I read a healthcare AI vendor's landing page. Forty-seven percent increase in bookings. Twelve-x ROI. Forty percent fewer readmissions. Big numbers, no methodology, a six-week measurement window, no control group in sight. Doesn't mean the tool is bad. Just means nobody's checked yet, and that's a gap that gets a lot more dangerous in healthcare than it ever was in ecommerce.

A missed appointment is annoying. A missed red flag on a post-surgical call is a different category of problem entirely, and it's the one most vendor decks quietly skip past on the way to the ROI slide.

Why is there almost no peer-reviewed evidence for AI voice agents in healthcare?

Because the field moved faster than the research could keep up with it.

A recent evidence review scanned six major research databases, PubMed, JMIR, npj Digital Medicine, BMC Health Services Research, JAMIA, and arXiv, and found zero peer-reviewed studies measuring the impact of voice AI on inbound appointment booking. A separate Duke metanarrative review screened 3,415 citations on AI scheduling and included eleven studies worth citing. Every single one was algorithmic, none of them voice-based. The vendor numbers you're reading on homepages come almost entirely from uncontrolled case studies.

What should you actually ask a vendor before you believe their numbers?

Ask what the comparison group was, and over what time period.

A booking lift measured against last month, with no control arm and a two-week window, tells you almost nothing about steady-state performance. Ask whether the claim is about outbound reminders, which have real backing, a meta-analysis of ten randomized trials found reminded patients were more likely to show up, or inbound scheduling, which is essentially unstudied. Ask what happens when the agent gets it wrong. If a vendor can't answer that last one specifically, with numbers, walk.

Why does zero critical adverse events matter more than an ROI multiple?

Because an ROI claim describes what the vendor wants to be true, and an adverse event count describes what actually happened to a patient.

My daughter is ten. She told me once, out of nowhere, about something completely unrelated to any of this, I work for you, not the other way around. It's stuck with me since as basically a description of how I think about every team and every tool we build. The agent works for the patient outcome. Not for the demo. Not for the board deck. We've run more than a million patient interactions and tracked every one of them, and the number that matters to us isn't the engagement percentage. It's the zero next to critical adverse events.

How do you build something that can survive this kind of scrutiny?

You build it so someone else can check your work.

HANA runs fully open source and self-hosted, documented at docs.hana.health, with no dependency on a single closed model vendor. That means a health system's own team can audit exactly what the agent said, when, and why it escalated or didn't. That's not a compliance checkbox tucked into an appendix. That's the actual answer to how do I know this is safe, which is the only question that matters once you're past the sales deck and looking at real deployment history.

What does the evidence actually support right now?

Less than the marketing suggests, but more than nothing.

Outbound reminders reliably improve attendance, that part's solid. Structured post-discharge contact within a defined window reliably lowers readmission risk, a Kaiser Permanente study on heart failure patients found a 19% drop in 30-day readmission odds with contact inside 7 days of discharge. What's unproven is the inbound scheduling lift every vendor advertises above the fold. CMS backs this distinction with real money too: the Hospital Readmissions Reduction Program can cut a hospital's Medicare payments by up to 3% for excess readmissions, which is exactly why the proven interventions deserve more attention than the speculative ones. Build your evaluation around what's actually been shown to work, and treat everything else as an unverified pilot, because that is exactly what it is.

Key Takeaways

Vendor claims about AI voice agents in healthcare mostly come from uncontrolled case studies, not peer-reviewed research, and that gap is wider for inbound scheduling than it is for outbound outreach. The right response isn't to avoid the technology, it's to demand the same rigor you'd demand from a drug trial: control groups, defined time windows, and a clear account of what happens when the system gets it wrong. Adverse event tracking is the metric that should decide a purchase, not the ROI multiple on the homepage, and before you compare pricing across vendors, compare their evidence. A vendor that can show you real deployment history alongside a clean safety record is worth far more than a booking percentage with no denominator attached.

FAQ

Are AI voice agent ROI claims in healthcare accurate?

Some are directionally true, but almost none are independently verified. Most published figures come from vendor case studies with short measurement windows and no control groups, so treat any specific percentage as a starting hypothesis, not a guarantee.

What's the difference between outbound and inbound voice AI evidence?

Outbound reminder calls have real peer-reviewed backing, a meta-analysis of ten randomized trials found reminded patients were more likely to keep their appointments. Inbound scheduling claims, the type behind most homepage headlines, currently have no peer-reviewed support at all.

What should a health system require before signing an AI voice agent contract?

Ask for adverse event data, not just ROI projections, along with a clear escalation protocol and an explanation of what happens when the system encounters something it wasn't trained for. If a vendor can't produce that, the contract is premature.

If you want the actual data behind our numbers instead of a slide with a big percentage on it, book a discovery call and I'll show you the real thing.