All posts
Patient Engagement
Hana Health
August 11, 2026

What a Study on 1,431 Cardiac Patients Actually Proves About AI Phone Calls

I threw out an app once. Spent months on it. Built it for bipolar patients, clean design, smart reminders, the whole thing. Engagement sat at 15%. Then we tried something almost embarrassingly simple: we called the same patients with an AI voice agent instead of pinging their phone. Engagement jumped to 85%. That gap is basically the entire reason HANA exists.

So when I saw a peer reviewed study out of a cardiac catheterization center this year, I read the whole thing twice. Not because it was about us. Because it was the first time I'd seen someone else's data land in almost exactly the same place ours did, with none of the incentive to make it look good.

Can an AI voice agent safely call patients before a procedure?

Yes, according to the study published in npj Digital Medicine. Researchers deployed an AI voice assistant called Sofiya to make pre-procedural calls to patients scheduled for cardiac catheterization. Over roughly six months, Sofiya placed 1,606 calls to 1,431 patients, walking them through instructions, collecting clinical data, and answering questions before handing off to a nurse when needed.

Here's the number that matters. Completion rates, meaning patients who made it through the whole script and answered every clinical question, climbed to an average of 86.4% in the first 90 day phase and held at 87.9% in the second. That's not a demo number. That's what happens after the novelty wears off and real patients with real anxiety about a heart procedure pick up the phone.

Why did the completion rate keep climbing instead of dropping off?

Because the first three months were a tuning period, not a launch. Phase one paired clinicians directly with the developers, catching every place the script confused someone or the AI misread an answer. Phase two handed the system to nurses running it in the real world, and the completion rate held steady with fewer errors, not more.

This matches something I learned running a circus act in the Australian desert of all places (feels like another lifetime, honestly). The acts that pulled the biggest crowds weren't the most impressive ones. Aerial silk looks incredible and draws three people. Fire chains are simple, loud, and repeatable, and they draw a hundred. Complexity doesn't survive contact with a real audience. Neither does an AI script that only works when a developer is standing next to it.

What happens when the AI doesn't know the answer?

It hands off to a human, and that's the part people underestimate. AI system errors dropped from 6.0% in phase one to 2.6% in phase two, and every one of those errors routed to a nurse instead of leaving a patient stuck. That's the actual design question in healthcare voice AI. Not "can it sound human," but "does it know when to stop pretending it knows something."

We built HANA the same way. Over a million patient interactions, zero critical adverse events, because the agent is trained to escalate, not to bluff. You can see how that shows up across different clinical use cases we've deployed, from post-discharge follow-up to chronic condition check-ins.

Is 85% engagement realistic outside a hospital pilot?

It has to be, or none of this matters. A study is one center, one specialty, one very motivated research team. Other vendors are seeing similar signal at commercial scale too. Assort Health reported over 60% conversion on AI outreach for appointment scheduling and up to 90% on rescheduling campaigns, run across hundreds of practices, not one research center. That's the bar we hold ourselves to, and it's why our own research page tracks engagement the same unglamorous way this study did. Not vanity metrics. Completion rates. Adverse events. The stuff that actually tells you if patients are being reached or just being pinged.

What should a clinic actually measure before trusting an AI to call patients?

Three things, in this order. Completion rate, meaning did the patient get through the conversation and not just pick up and hang up. Escalation accuracy, meaning did the system know when to bail to a human. And error rate over time, meaning did the whole thing get safer as it scaled or riskier. Most vendor pitches lead with call volume or cost savings. Those are the least important numbers in the room.

I watched a Stripe dashboard cross a million dollars in a single day once, at a company I'd later leave in pieces. The dashboard looked incredible. The product underneath it was quietly broken, and scale just meant more people found out slower. Volume hides problems. Completion rate and error rate don't.

How do you deploy this without breaking trust with patients?

Slowly, with a human watching closely for the first stretch, exactly like the study did. The 90 day phase one wasn't bureaucracy. It was the only way anyone earned the right to hand the calls to nurses in phase two. Look at how the case studies from our own rollouts read: nobody skipped the supervised period, because that's where you find out what patients actually push back on before it becomes a pattern.

Key Takeaways

An AI voice agent calling 1,431 real cardiac patients completed 86 to 88% of its conversations end to end, with errors dropping by more than half as the system matured under real world use. The lesson isn't that AI can talk to patients. Plenty of things can talk. It's that a system built to escalate instead of guess is what makes the completion rate and the safety record climb together instead of trading one for the other. That's the same pattern we saw going from a 15% engagement app to an 85% engagement voice agent, and it's the same pattern this study just proved in a setting with far higher stakes than ours started in.

FAQ

Can AI voice agents safely handle pre-procedure patient calls? Yes. A peer reviewed study of 1,606 calls to cardiac catheterization patients found completion rates of 86 to 88%, with the AI escalating to a nurse whenever it hit a question outside its scope.

What's the biggest risk with AI patient outreach? The AI not knowing when to stop. Systems that guess instead of escalating create the adverse events that make healthcare AI dangerous. Escalation accuracy matters more than how natural the voice sounds.

Does AI patient engagement actually work outside of pilots? It has to be measured the same way a pilot is measured: completion rate, escalation accuracy, and error rate over time, not call volume. HANA has run this model across 1M+ interactions in 5 countries with zero critical adverse events.

If you're trying to figure out whether this works for your clinic specifically, book a discovery call and I'll walk you through the numbers that actually matter.