All posts
Voice AI
Hana Health
Voice AIPatient EngagementAugust 31, 2026

Voice AI in Healthcare Has an Evidence Problem, and I'm Glad Someone Finally Said It

Matteo

I built a mental health app once. For bipolar patients. Beautiful thing, honestly, and I was proud of it in the way you're proud of something before anyone uses it. Engagement: 15%.

Fifteen.

Which means 85 out of every 100 people I'd designed it for opened it, shrugged, and went back to their lives. I told myself the number was fine. Industry standard. Everyone sits at 15%. That's the story you tell yourself when the thing you made doesn't work and you're not ready to say so out loud.

Then we stopped building screens and just called people. With AI. Engagement went to 85%, same patients, same clinical goal, the number flipped completely upside down. That's the day HANA actually started, though I didn't know it yet.

I'm thinking about all this because Deepgram published a piece breaking down what the evidence actually supports for AI voice agents in patient engagement, and the answer is uncomfortable for people like me. Zero peer-reviewed studies exist on inbound AI voice scheduling. Every 30% to 50% booking lift you've seen on a vendor landing page traces back to a case study with no control group, no sample size, and a measurement window shorter than a bad flu.

I sell voice AI for a living. And I think they're completely right.

Does AI voice actually improve patient engagement?

For outbound follow-up, yes, and the evidence there is real. For inbound scheduling, nobody has proven it in a peer-reviewed setting, and you should be suspicious of anyone who says otherwise. The Duke metanarrative review screened 3,415 citations on AI scheduling and found eleven relevant studies. Every single intervention was algorithmic. Markov decision processes, XGBoost, that kind of thing. None were voice.

That gap matters because clinics keep buying inbound tools using outbound evidence. Reminder meta-analyses are solid. They do not transfer.

Our own numbers, which we publish rather than hint at, sit on the outbound side: 85% weekly engagement against a 15% to 20% industry baseline, across more than a million patient interactions.

Why does engagement matter more than reach?

Because reach is easy and engagement is the thing that actually changes outcomes. A pilot published in JACC's quality improvement series makes this brutally clear. Cardiac Solutions ran daily digital outreach with 375 heart failure patients after discharge. Overall 30-day readmission dropped to 15.5% against a 24% national historical average.

Then they split the cohort by engagement.

Patients who opened more than half the messages readmitted at 7.7%. Patients who opened fewer than half readmitted at 21.6%, which is basically the national number. Same program. Same clinic. Same discharge. The only variable that moved was whether the patient actually engaged, and it swung readmission by a factor of nearly three.

You can send a million messages into a void. The void doesn't get healthier.

What should a clinic measure instead of vendor claims?

Call completion rate, first-call resolution, and patient callback rate. Those three tell you whether patients finished what they started, got the right outcome, and didn't have to redo it. A headline booking percentage tells you almost nothing without a baseline and a control.

Also test the speech stack against your own patients before you sign anything. Real recordings. Regional accents, a speakerphone in a noisy waiting room, someone code-switching between English and Spanish mid-sentence. Generic speech models trained on podcasts and call center audio fall apart on exactly that, and word error rate is what decides whether the call converts or the patient hangs up. We built multilingual handling across five countries and three languages because our first deployments taught us this the hard way, not because it looked good in a deck.

Is voice AI safe for patient-facing calls?

It's safe when it's scoped and escalated properly, and it's dangerous when it isn't. Our agents don't diagnose and don't prescribe. They ask, they listen, they document, and they hand off to a human the moment something looks clinically off. Across a million-plus interactions we've had zero critical adverse events, and I'd rather report that number honestly than round it into something prettier.

The other half of safety is where the thing runs. HANA is fully open-source and self-hosted, no OpenAI dependency, so patient audio stays inside your perimeter rather than travelling to a vendor's cloud you can't inspect. The integration path is documented publicly for the same reason. If you can't read how it works, you can't govern it.

What does the ROI actually look like?

For the clinics running us, it's roughly 31:1, and the math is boring in a good way. Missed follow-up drives avoidable readmissions and lost quality revenue. Empty slots cost between $150 and $300 each in lost production. You're not buying magic. You're buying the calls that your front desk was never going to make on a Tuesday afternoon with four people in the waiting room.

I watched a Stripe dashboard cross a million dollars in a single day once, at a DTC company where the product was genuinely bad. Scale hides problems beautifully. So when someone shows me a big number with no denominator, I don't get excited anymore. I ask what's underneath it. Our pricing is published for the same reason.

Key Takeaways

The honest position in 2026 is that outbound patient follow-up by voice has real evidence behind it and inbound scheduling doesn't yet. Engagement, not reach, is the variable that moves readmissions, and the heart failure data shows a threefold swing depending on it. Vendor case studies aren't research, and you should measure completion, resolution, and callback rates on your own population rather than trusting a headline. Test your speech stack on real patient audio before you commit. And insist on architecture you can inspect, because governance you can't verify isn't governance.

FAQ

Is there peer-reviewed evidence for AI voice agents in patient scheduling? Not for inbound scheduling. A search across PubMed, JMIR, npj Digital Medicine, JAMIA and other databases turns up no primary outcome studies. Outbound reminder evidence is much stronger, backed by randomized trials, but the two workflows aren't interchangeable.

How is HANA's 85% engagement figure measured? It's weekly engagement across live clinical deployments, meaning the share of enrolled patients who actually complete a follow-up interaction in a given week. The industry baseline for digital patient engagement tools sits at 15% to 20%, which is where my old app lived.

What's the fastest way to know if voice follow-up fits my clinic? Look at where calls are going unmade rather than where they're going unanswered. Post-visit checks, chronic care check-ins, and recall campaigns are where the gap usually sits. If you want to walk through your own numbers, grab a slot on my calendar and we'll model it against your volume.