All posts
Voice AI
Hana Health
September 22, 2026

Is There Any Real Evidence That AI Voice Agents Improve Patient Engagement?

I built a mental health app once. Beautiful thing. Mood tracking for bipolar patients, clinician dashboard, the lot.

Fifteen percent of them used it weekly.

Fifteen. After a year of work, after the user interviews, after everything I thought I knew from actually training as a clinical psychologist before I started building software for a living. I sat with that number for a long time, and then I threw the whole thing out and had an AI call the patients on the phone instead, and eighty five percent of them picked up and talked every week. That was the start of HANA. Not a strategy. A failure that turned into one.

So when I read Artera's roundup of the research on generative AI voice agents last week, something itched. They're right. The evidence is thin. And I'm part of the industry stacking claims on top of it.

Is there any peer-reviewed evidence that AI voice agents book more appointments?

No. Not yet. A 2026 review of the evidence base searched six major databases including PubMed, JMIR, npj Digital Medicine and JAMIA and found zero peer-reviewed studies measuring whether an AI voice agent books more appointments than a human scheduler. A Duke metanarrative review screened 3,415 citations, included eleven studies on AI scheduling, and not one of them was voice-based.

Every booking-lift number you've seen on a landing page is a vendor case study.

Mine included. We publish ours in HANA's research section and I'd honestly rather you poke holes in them than take them on faith.

So why does every vendor quote the same numbers?

Because we're all citing each other. Somebody reports a 21% appointment lift measured two weeks after go-live, somebody else rounds it into a range, a third company puts "30 to 50%" in a deck, and within a year that range is industry consensus without a single control group behind it. Two-week measurement windows. No baseline. No comparison arm.

It's the Stripe dashboard problem. Back in my DTC days I watched revenue cross a million dollars in a single day and felt like a genius. The product was bad. Scale hides problems, it doesn't solve them, and a big number on a slide hides them even better.

What should a clinic actually measure instead?

Three things, and none of them are call volume. Call completion rate: did the patient finish the thing they called about. First-call resolution: did they finish it without calling back. Patient callback rate: how many come back because the first interaction failed. Those three tell you whether the automation worked. "Calls handled" tells you the phone rang.

A call routed confidently to the wrong clinic closes as a success on every dashboard that counts calls handled.

Measure the outcome, not the activity. The use cases where this actually holds up are the dull ones. Reminders. Post-visit follow-up. Intake. Repetitive work with a clear finish line.

Does voice alone fix patient engagement?

It doesn't, and I think that's the most useful thing in the Artera piece. A booking is rarely one clean call. It spans intake, prep instructions, a reminder, a reschedule, and a follow-up, moving across voice and text and portal over weeks. Automate the phone call on its own and you've fixed one link in a chain that still snaps everywhere else.

Worse, a standalone bot creates a new silo with its own version of the patient's status, which somebody then has to reconcile against the record the staff actually work from. Reconciliation is where the errors live.

What does honest measurement look like?

It looks boring. You set a baseline before you launch, you pick a comparison group, you measure for ninety days instead of fourteen, and you publish the number even when it's worse than you hoped.

I spent years in an industry where the Australian circus taught me more about this than any board meeting. Crossing the desert with those performers, the acts drawing the biggest crowds weren't the most technically complex ones. Fire chains beat aerial silk every night. Not because fire chains are harder. Because people could actually see what was happening.

Same with clinical outcomes. A number people can verify beats an impressive one they can't. Across more than a million patient interactions in five countries we've had zero critical adverse events, and that's the figure I care about most, because it's the one that would be obvious if it were false.

Key Takeaways

The peer-reviewed literature backs voice agents for structured, high-volume, non-diagnostic work, and it backs nothing beyond that yet. Every booking-lift claim in the market today, ours included, comes from vendor-reported data with short measurement windows and no control arms. That doesn't mean the tools don't work. It means you should trust your own pilot data over anybody's marketing, measure completion and resolution rather than volume, and treat voice as one channel in a coordinated workflow rather than a product that fixes engagement by itself. Set your baseline before go-live, or you'll never know what changed. And if a vendor won't show you their methodology, that's the methodology.

FAQ

Are AI voice agents HIPAA compliant out of the box?

No. Compliance requires a signed Business Associate Agreement covering every subprocessor that touches PHI, not just the primary vendor. If the speech recognition, the language model and the storage layer are separate companies, each one needs its own BAA. Ask for the paperwork before any patient data moves. Our integration and deployment documentation covers how a self-hosted setup changes that calculus.

How long should a pilot run before I trust the numbers?

Ninety days minimum, with a baseline captured before launch. Two-week windows catch the novelty period and miss the regression. Most claims in the market come from exactly those two-week windows.

Is it cheaper to build this in-house or buy it?

It depends on whether you want to own the stack. Building looks cheaper on paper until you count BAAs across vendors, concurrency scaling, QA tooling and ongoing maintenance. HANA is open source and self-hosted, which is a third option that people forget exists, and the cost breakdown is public.

If you want to pressure-test your own numbers against ours, grab a slot on my calendar and bring the ugly data. That's the interesting conversation.