Do AI Voice Agents Actually Improve Patient Engagement? What the Evidence Really Says
I built a mental health app for people with bipolar disorder. Beautiful thing. Mood tracking, sleep logs, a little chart that showed you your own patterns across ninety days, the kind of screen you put on stage and everyone claps politely.
Engagement was 15%.
Fifteen. Out of a hundred people who genuinely needed the thing, eighty five never opened it twice, and I sat with that number for months because I'm a clinical psychologist by training and I know exactly what happens in the gap between one appointment and the next, and no amount of good interface design was going to close it.
So we threw it out and called them instead. AI on the phone, once a week, asking how they were.
85%.
That inversion is basically why HANA exists. But I read something recently that made me want to be honest about the part of this industry nobody says out loud, so let's be real for a minute.
Is there peer-reviewed evidence that AI voice agents improve patient engagement?
For inbound appointment booking, no. Zero peer-reviewed studies exist. A 2026 review of the evidence base searched PubMed, JMIR, npj Digital Medicine, BMC Health Services Research, JAMIA and arXiv and found nothing. Not a weak study. Nothing. A Duke metanarrative review screened 3,415 citations, included eleven on AI scheduling, and every single intervention was algorithmic rather than voice-based.
That should bother you if you're buying. It bothers me and I'm selling.
Why do vendor numbers look so good then?
Because most of them are measured by the vendor, over a short window, with no control group. One widely quoted 47% booking lift at an academic medical centre turns out to describe web assistant conversions only, not total appointment volume across every channel. Another vendor's 21% lift was measured within two weeks of go-live. Two weeks. That's not a result, that's a honeymoon.
I've been on the other side of a number like that. In my DTC days I watched a Stripe dashboard cross a million dollars in a single day and felt like a genius. The product was bad. Scale hides problems, it doesn't solve them, and a two week pilot hides them beautifully.
What does the evidence actually support?
Outbound contact. That's the part with real research behind it. A meta-analysis of ten randomised trials found reminded patients were meaningfully more likely to attend, and a 2023 trial found something more interesting on top of that: adding live human calls to automated reminders dropped no-show rates further than automated reminders alone.
Read that again. The automation helped. The voice helped more. Which reframes the entire question from can a machine replace the call to how close can a machine get to the call, and that's a much more honest question to build a company around. It's the one we track in our own outcomes research.
Does engagement itself predict clinical outcomes?
Yes, and this is the finding I'd put on a poster. In a digital outreach programme after heart failure hospitalisation, patients who engaged with more than half of their messages had a 30-day readmission rate of 7.7%. Patients who engaged with less than half sat at 21.6%, essentially the national average. Same programme. Same content. The only variable that moved was whether the patient actually showed up to the conversation.
So engagement isn't a vanity metric. It's the mechanism. Everything downstream of it is arithmetic.
What should a clinic measure instead of vendor lift claims?
Four numbers, measured against your own baseline, over at least ninety days. Contact rate, meaning what share of the target list you actually reached. Completion rate, meaning how many of those calls finished the protocol instead of dying at hello. Escalation accuracy, meaning how often a flagged patient genuinely needed a human. And callback rate, because rework is the tell. Map those to the specific workflows you're automating before you look at a single ROI slide, then check the cost per interaction against what those four numbers are worth to you.
Why did the phone work when the app didn't?
I crossed the Australian desert with a circus once. Genuinely, a whole travelling show, tiny towns, dust everywhere. And the thing I noticed night after night was that the performers pulling the biggest crowds weren't doing the most technically difficult acts. Aerial silk is astonishing and it takes years. Fire chains take months and everybody stops walking.
The app was aerial silk. The phone call is fire chains. Your patient already knows how to answer a ringing phone, she learned it when she was six, and she doesn't have to download anything or remember a password or find the thing she agreed to at discharge while she was still on painkillers. Accessible beats impressive. Every time. Our deployment results keep saying the same boring thing.
Key takeaways
There is no peer-reviewed evidence that AI voice agents lift inbound appointment booking, and any vendor telling you otherwise is quoting their own marketing back at you. What is well evidenced is that proactive outbound contact improves attendance, and that live voice improves it further than automation alone. Engagement is the mechanism that drives every downstream number, so measure contact rate, completion rate, escalation accuracy and callback rate against your own baseline over ninety days rather than trusting a two week pilot. And when you're choosing between a sophisticated interface and a boring channel your patients already use, choose boring. We run 85% weekly engagement against a 15 to 20% industry baseline across more than a million patient interactions, and the reason isn't clever. It's that the phone rings and people pick it up.
Frequently asked questions
Are AI voice agents safe for clinical follow-up?
They are when they're scoped to structured check-ins and escalate to a human on any red flag rather than improvising clinical judgement. Across more than a million interactions we've recorded zero critical adverse events, and that's a function of tight scope and fast escalation, not model size.
How long before a clinic sees results?
Contact and completion rates move in the first two weeks because they're direct effects. No-show and readmission effects need a full quarter to separate from seasonal noise. Anyone promising clinical outcomes inside a month is measuring the honeymoon.
Do patients mind talking to an AI?
Less than you'd expect, and considerably less than they mind nobody calling at all. Disclosure helps rather than hurts, because what patients react badly to is feeling tricked, not feeling automated.
If you want to argue with any of this, or you've got a follow-up list nobody has time to work through, book twenty minutes with me and bring your baseline numbers. Those are the only ones that matter.
