How Should a Health System Actually Test an AI Voice Agent Vendor?
You've sat through this demo. I know you have.
The agent books a routine appointment in under a minute, the voice is warm, the dashboard glows green, and everybody in the room does that little nod people do when they've already decided, and then someone at the back asks what happens if the patient says they're not sure they can afford the copay. Silence. The sales engineer laughs and says "great question, let's take that offline."
It never comes back online.
What's the biggest mistake health systems make when evaluating AI voice agents?
They test the easy call. Nearly every vendor can book a simple appointment with a cooperative patient who speaks clearly and wants exactly what the script expects. That proves almost nothing about how the agent behaves on the calls that actually cost you money: the confused patient, the reschedule with three constraints, the post discharge check in where someone mentions chest tightness halfway through a medication question.
A 2026 buyer's guide to AI voice agents from Hello Patient put it bluntly. A simple booking proves the agent can complete a simple booking. That's it. It doesn't prove billing, specialty scheduling, outbound conversion, or EHR write back.
Which workflow should you bring to the vendor demo?
Bring the hardest workflow your staff handles every week, not the one on the brochure. Pick a real, messy, recurring call type (a post discharge follow-up for heart failure, a colonoscopy prep reminder for a patient who's already rescheduled twice, a care gap call in Spanish) and make the vendor run it live, end to end, including the handoff to a human and the note that lands in your EHR.
Then change one variable mid call. Have the "patient" switch languages. Or go quiet. Or mention something clinical that should trigger escalation.
Honestly, that's where you learn everything. Not in the happy path. In the swerve.
Why does outbound matter more than inbound for health systems?
Because the expensive failures in healthcare are the calls nobody makes. Inbound answers the patient who's already motivated enough to pick up the phone and wait on hold. Outbound reaches the patient who isn't, which is the patient most likely to miss a follow-up, skip a refill, and show up in your ED on day nine.
The operational appetite is clearly there. In a Lirio and Sage Growth Partners survey covered by Fierce Healthcare, 60% of health system executives named automating patient outreach as their top priority, yet 35% hadn't invested in outreach AI at all, while 83% had already bought documentation tools. We're buying the scribe before we've fixed the phone.
How do you measure engagement without getting fooled by vanity metrics?
Measure completed conversations and completed actions, not calls placed or messages sent. "We made 40,000 calls" is a vanity number. "We reached 31,000 patients, 26,000 finished the check in, 900 were escalated to a nurse, and 820 of those escalations happened within the hour" is a program you can run a health system on.
I learned to be allergic to vanity metrics the painful way. Years ago, in my DTC life, I stared at a Stripe dashboard that hit a million dollars in a day. The product was bad. The number was real and it meant nothing, and it took me a long time (and, eventually, crying uncontrollably in a meeting with my brain completely fried) to admit that. The body keeps score even when the dashboard doesn't.
So ask every vendor for the denominator. Always.
What safety evidence should you require before going live?
Ask for real world volume and real world adverse event data, not benchmark scores. Benchmarks tell you how a model does on a test. You need to know how the agent behaved across hundreds of thousands of actual patient conversations, how often it escalated, how often it should have escalated and didn't, and who reviewed those calls.
For context on what "good" looks like, HANA has run more than a million patient interactions across five countries and three languages with zero critical adverse events, and you can dig into what we've shared on our research page. You should hold every vendor, including us, to that kind of transparency. If they won't show you the escalation log, that's your answer.
Should a health system own the AI or rent it?
If patient conversations are becoming core infrastructure, you should at least be able to see inside the box. A voice agent that talks to your patients every day is making clinical adjacent decisions in your name, and "trust us, it's proprietary" isn't a governance strategy. Self hosted and open source options let your security team audit the model, keep PHI inside your perimeter, and avoid getting locked into one foundation model provider's pricing and policy changes.
That's the reason we built HANA fully open source, self hosted, with no OpenAI dependency. Your engineers can read the technical documentation before your procurement team reads a single slide. And if you want to understand why a clinical psychologist ended up building phone infrastructure, the short version is on our about page: I once built an app for bipolar patients that got 15% engagement, threw it away, and started calling people with AI instead. Engagement hit 85%. I haven't looked back.
My daughter is ten, and I tell her I work for her, not the other way around. Your vendor should feel the same way about your patients.
Key Takeaways
Most AI voice agent evaluations fail because they test the easy call, and a simple booking tells you nothing about billing, escalation, outbound conversion, or write back into the record. Health system leaders should bring their hardest recurring workflow to every demo, change a variable mid call, and watch what happens at the handoff. Outbound outreach is where the real clinical and financial risk lives, and executives already know it, even though most AI budgets have gone to documentation first. Judge programs on completed conversations and completed actions, demand real world safety and escalation data, and favor architectures your own team can inspect.
FAQ
What's the most important question to ask an AI voice agent vendor?
Ask them to run your hardest real workflow live, including escalation to a human and documentation in your EHR. How the agent handles an unexpected turn tells you more than any slide.
How can a health system compare AI voice vendors fairly?
Use the same scripted scenario for every vendor, include at least one mid call curveball, and score on completed actions, escalation accuracy, and write back. Ask each vendor for the denominator behind every success metric.
Is open source AI safe enough for patient conversations?
Yes, when it's paired with clinical guardrails, human escalation, and real world monitoring. Open source often improves safety because your own team can audit the model and keep patient data inside your infrastructure.
If you'd like to pressure test HANA on your hardest workflow, I'm happy to run it live with you. Grab 20 minutes with me here and bring the call your team dreads most.
