What a $1M Day on Stripe Taught Me About Reading Vendor AI Claims
Years ago I ran a DTC company that did about $12M a year. One day I watched the Stripe dashboard cross $1M in a single day. I remember sitting there feeling like we'd made it, genuinely proud, checking the number every ten minutes like it might disappear.
The product was bad. Not catastrophic, just mediocre in ways that scale was actively hiding from me. Big revenue numbers make you stop asking hard questions. That's the lesson I never unlearned, and it's the first thing I think about every time I read a healthcare AI vendor's landing page promising 12x ROI or a 47% lift in something, with no methodology attached and no independent set of eyes on the number.
Why Should You Be Skeptical of Voice AI ROI Claims in Healthcare?
Because most of them come from the vendor grading their own homework. Scroll any AI voice agent page in this space and you'll see numbers like an average 12x ROI or booking lifts in the 20 to 50% range, almost always self-reported, almost never independently verified, almost never disclosing sample size or measurement window. That doesn't make the technology fake. It makes the marketing unreliable as a way to decide who to trust with your patients' calls.
Is There Any Peer-Reviewed Evidence for Voice AI in Patient Scheduling Yet?
Not really, and it's worth knowing that going in. A thorough review of the published evidence found zero peer-reviewed studies on inbound AI voice scheduling across six major research databases, including PubMed, JMIR, and JAMIA. A separate Duke metanarrative review screened 3,415 citations on AI scheduling broadly and found eleven relevant studies, every single one algorithmic rather than voice-based. If you're procuring for an academic medical center where evidence matters to the committee signing off, that gap is real and it's not going away because a vendor's homepage says "clinically validated."
Why Do Vendor Case Studies Look So Convincing and Prove So Little?
Because they measure the easiest thing to measure, over the shortest window that still looks good. One well-known vendor reports a 47% increase in bookings at a major academic center, but that figure applies to their web assistant only, not phone volume, and it doesn't reflect total scheduling across channels. Another reports hundreds of thousands of touchpoints avoided with no control group and no comparison to what would have happened anyway. None of this is lying, exactly. It's the Stripe dashboard problem, dressed up in a case study PDF. The number is real. What it actually tells you is much smaller than the headline implies.
What Should You Actually Ask a Vendor Before You Sign Anything?
Ask what the denominator is. A vendor claiming 40% call automation needs to tell you 40% of what population, measured over what window, compared to what baseline. Ask whether the number includes total volume or just the subset that made it easy to report. Ask for the raw counts, not just the percentage. And ask what happens on the calls that don't go well, because that failure rate tells you more about the product than any highlight reel.
How Does HANA Handle This Differently?
We publish the boring number alongside the good one. Over 1 million patient interactions logged, with zero critical adverse events, across 5 countries and 3 languages. Not a 90-day pilot window. Not a subset of the easiest channel. The full run. We also run self-hosted, open source, on purpose, so a health system's technical team can actually look inside the thing instead of trusting a sales deck. That's slower to sell. It's also the only version of this I'd trust if it were calling my own family.
What Does "Evidence-Based" Actually Mean Once You Strip the Marketing Language Off It?
It means the number survives someone else checking it. My daughter is ten, and she has a habit of telling me, dead serious, "I work for you, not the other way around," whenever I try to make her do something that's actually for my convenience and not hers. I think about that line constantly with AI in healthcare. The tool should work for the patient's outcome, not for the vendor's next funding round slide. If a claim only survives inside the company that made it, it isn't evidence. It's marketing wearing evidence's clothes. See how we approach this across our use cases.
Key Takeaways
Key takeaways. Most voice AI ROI claims in healthcare are vendor self-reported, with no peer-reviewed backing yet across major research databases. That's not disqualifying, but it means the burden is on you to ask for denominators, raw counts, and failure rates before you trust a headline number. Case studies measure the easiest thing over the shortest flattering window, which is exactly the trap a big revenue number set for me years ago on a much smaller stage. The way through it is transparency you can independently verify, not another slide with a bigger percentage on it.
FAQ
Is there peer-reviewed research on AI voice agents for healthcare scheduling? As of 2026, no systematic reviews or randomized trials exist on inbound AI voice scheduling specifically, according to a review across PubMed, JMIR, JAMIA, and other major databases. Evidence for outbound reminder calls is stronger and better established.
Why do vendor ROI numbers vary so widely, from 12x to 880%? Because they're measuring different things over different windows with no shared methodology. Some count web conversions, some count avoided outbound calls, some count total campaign ROI. Without a disclosed denominator, none of these numbers are comparable to each other.
What should a health system ask for before trusting a voice AI ROI claim? Ask for the raw counts behind the percentage, the measurement window, the population included, and what happened on the calls that failed. A vendor that publishes all four is worth a real conversation. One that publishes only the headline number usually isn't ready for one.
If you want to see our actual numbers instead of a highlight reel, book a discovery call and we'll show you the raw data, not just the ROI slide.
