All posts
Voice AI
Hana Health
August 12, 2026

Every Voice AI Vendor Claims 40% More Bookings. Where's the Evidence?

I built a mental health app once. Beautiful thing, honestly. Clean onboarding, mood tracking, a breathing exercise with a little circle that expanded and contracted, all of it designed for bipolar patients who needed to catch an episode before it caught them.

Fifteen percent engagement.

I stared at that number for about three weeks. Told myself it was an onboarding problem, then a notification problem, then a segmentation problem, then I ran out of problems to blame and had to sit with the actual answer, which is that people going through the worst days of their life do not want to open an app to tell you about it.

So we threw it out. Called them instead, with AI. Eighty-five percent weekly engagement, and that number has held across more than a million patient interactions since.

I bring this up because I read a piece from Deepgram on the evidence behind voice AI patient engagement claims and it made me uncomfortable in a way I think is useful. Their argument, roughly: everybody in this category is quoting 30% to 50% booking lifts, and almost nobody can show you where the number came from.

They're right. And I sell voice AI.

Is there peer-reviewed evidence for AI voice agents in patient scheduling?

For inbound scheduling, no. Not yet. A search across PubMed, JMIR, npj Digital Medicine, BMC Health Services Research, JAMIA and arXiv turns up zero randomized trials, zero systematic reviews, and zero primary outcome studies testing whether an AI voice agent answering the phone books more appointments than a human scheduler.

That's not a gap in the marketing. That's a gap in the science.

There is real published evidence for adjacent things. A prospective evaluation in npj Digital Medicine ran 1,606 AI voice calls to 1,431 cardiac catheterization patients and reported an 87.9% call completion rate in its second phase, with system errors dropping to 2.6%. That's a real study with real numbers. It's also about pre-procedural preparation, not inbound booking. Different workflow, different claim.

If you're a clinic owner being sold on inbound automation, know which one you're buying.

Why do vendor case studies overstate voice AI results?

Because the measurement window is short, the control group doesn't exist, and the denominator is whatever makes the number look best.

Look at the pattern. A vendor reports a 47% increase in booked appointments, and it turns out to be web assistant conversions only, not total appointment volume across channels. Another reports 21%, measured two weeks after go-live. Another reports 880% ROI without disclosing the methodology. Another cites 489,000 touchpoints, which is a measure of how much the system did, not how much it changed.

None of that is fraud. It's just what happens when a company measures itself.

I've been on the other side of this. In my DTC days I watched a Stripe dashboard cross a million dollars in a single day while the product underneath it was quietly bad. Scale hides problems. Volume metrics hide problems even better, because they always go up.

What should a clinic actually measure before buying voice AI?

Three things, and none of them are on the vendor's slide.

Call completion rate: did the patient reach the end of the conversation with the thing resolved? First-call resolution: did they have to call back? Patient callback rate: how much rework did the system create for your front desk?

Then measure your own baseline first. If 30% of your calls currently go unanswered, automation will look miraculous. If your answer rate is already 95%, the honest expected gain is small, and any vendor telling you otherwise is selling you your own inefficiency back at a markup. We publish our own outcome data and methodology for exactly this reason, because the alternative is asking people to trust a number with no shape behind it.

Run the math on your actual volume before you look at anyone's pricing. Not after.

Does outbound patient outreach evidence apply to inbound scheduling?

No, and conflating the two is the most common sleight of hand in this category.

The outbound evidence is genuinely solid. A meta-analysis of ten randomized trials found reminded patients were meaningfully more likely to attend. A 2026 systematic review of automated discharge instructions covering 34,386 patients across 13 studies found automated phone calls delivered the most consistent patient interaction of any modality, with completion rates between 44% and 56%.

That evidence supports outbound calling. It says nothing about whether an AI should answer your inbound line.

There's also a finding in that literature people skip past: a 2023 randomized trial found that adding live staff calls on top of automated reminders dropped no-show rates further than automation alone. Automation helps. It doesn't replace the value of somebody actually reaching a person.

What does honest voice AI performance look like?

Boring, mostly. Measured over months, not weeks. Reported with the denominator attached.

When I was crossing the Australian desert with a circus, the performers pulling the biggest crowds weren't doing the hardest acts. Aerial silk is objectively more difficult than spinning fire chains. Nobody cared. Fire chains won every night, because people could see immediately what was happening and why it was hard.

Healthcare AI has an aerial silk problem. Everyone is demoing the complicated thing when the market needs the legible thing: did you call the patient, did they answer, did the problem get solved, can you prove it.

Our own number is 85% weekly engagement against a 15% to 20% industry baseline, across a million-plus interactions with zero critical adverse events. I'll say the boring part too: that's outbound follow-up in specific clinical use cases, measured over sustained deployments across five countries and three languages, and it is not a claim about inbound scheduling. Our deployment write-ups carry the same caveats.

Key Takeaways

The category is real and the evidence base is thin, and both things being true at once is the actual situation. Peer-reviewed support exists for automated outbound follow-up and pre-procedural calling. It does not yet exist for inbound AI scheduling, so treat any inbound booking lift you're quoted as a vendor estimate rather than a finding.

Measure your own baseline before you evaluate anyone. The gain a voice AI system delivers is mostly a function of how many calls you're currently dropping, which means the honest answer to "how much will this help" is different for every clinic and no vendor can know it from a demo.

And ask for the denominator. Every time. A vendor who can tell you what they measured, over how long, against what, is a vendor who has actually looked.

FAQ

Are there peer-reviewed studies on AI voice agents in healthcare? Yes for some workflows, no for others. Pre-procedural calling and automated post-discharge outreach both have published evidence, including a 2026 npj Digital Medicine study of 1,606 AI-conducted preparation calls. Inbound appointment scheduling has no peer-reviewed outcome studies as of 2026.

How do I evaluate a voice AI vendor's performance claims? Ask three questions: what's the denominator, how long was the measurement window, and was there a control or baseline period. Claims measured within two weeks of go-live, or expressed as raw touchpoint counts, tell you about system activity rather than clinical or financial impact.

Is voice AI worth it if my clinic already answers most of its calls? For inbound, probably less than the marketing suggests, because your upside is capped by the calls you're currently missing. Outbound follow-up is a different question entirely, since most practices do almost none of it and the published evidence for automated outbound contact is considerably stronger.

If you want to pressure test any of this against your own numbers instead of ours, grab a slot and let's look at them together.