A patient picks up the phone. The voice on the other end sounds calm and competent — until it doesn’t quite fit. Maybe it stumbles over a name. Maybe it asks a question that doesn’t match what just happened at discharge. Maybe the patient mentions something that should raise a flag — new chest pain, a fall, a medication that isn’t sitting right — and the call moves on to the next scripted question anyway.
None of that is a technical failure. The AI did exactly what it was built to do: follow the script, listen, respond. What went wrong happened earlier, in decisions nobody made about how this technology would actually operate day to day.
In last week’s article I shared what to ask an AI voice vendor before signing — the evidence behind their claims, where the script ends and a human begins, who owns the compliance risk. Those questions matter. But answering them well still leaves the harder work undone. A vendor can hand over a technically sound system to an organization that never designed the process around it, and that gap is exactly where a call meant to reassure a patient ends up leaving them concerned instead.
To think through what “designed well” actually requires, we apply our LEAD Framework: Leadership Alignment, Experience Design, Adoption Training, and Dynamic Support. Two of those four directly apply here: Experience Design shapes how the call is introduced and received, and Dynamic Support is the ongoing discipline of catching what breaks after launch.
A good vendor won’t leave you to work all of this out alone. Escalation rules, disclosure language, even the tone of a difficult conversation — these typically get built and refined together with the care team during configuration, or what I guess nowadays we call “AI training”. That partnership can cover what gets built. It doesn’t cover what happens after go-live, which is where the four gaps below tend to open up.
Design the Introduction, Not Just the Script
Whether a patient trusts an automated call has less to do with how natural the AI sounds and more to do with whether they were expecting it in the first place. That’s Experience Design applied to the moment before the call ever happens, not just the words the AI says once someone picks up.
There’s a simple, well-documented pattern behind this: when a nurse or clinician tells a patient a call is coming before discharge, that patient is more likely to answer it, and more likely to stay on the line through the whole conversation. The heads-up isn’t a courtesy. It’s what turns an unfamiliar voice into an expected one.
That means the disclosure — what the discharge team says to prepare the patient — deserves the same design attention as the escalation rules. Does every clinician actually say it, every time? A workflow that assumes patients will figure out on their own that the call is legitimate isn’t a workflow. It’s a hope.
There’s a second disclosure moment, separate from the pre-call heads-up: the opening seconds of the call itself. The AI should state plainly, right at the start, that the patient is speaking with an automated system — not buried in a longer greeting, not implied and left for the patient to guess. A growing number of state disclosure laws are starting to require exactly this, but even without a legal mandate, it’s the right design choice: a vague or hedged disclosure line does the opposite of what it’s meant to do. It reads as evasive the moment a patient half-notices something isn’t human, which is the precise moment trust breaks.
Design the Whole Escalation, Not Just the Rule
Any AI voice vendor worth working with can point to an escalation rule: if a patient reports certain symptoms, the call routes to a human. What almost no organization designs is the process for checking, after the fact, that the rule actually fired when it should have. That process is the escalation review: sampling a set of calls on a regular basis, listening specifically for the moments that should have triggered a handoff, and confirming that it did.
But firing the rule only solves half the problem — recognizing that a call needs a human. It says nothing about what happens if that human doesn’t answer. In a fully staffed setting, a single routing target is often fine. In a rural CAH with one nurse covering discharge calls alongside everything else, it usually isn’t. A patient who’s told, even implicitly, that a real person is stepping in — and then sits on hold or hits voicemail — has traded one trust problem for a worse one.
Designing for this means giving the handoff itself a failure mode: a defined backup if the first line doesn’t answer within a set number of rings, and a fallback message that doesn’t leave the patient stranded. It also means widening what the escalation review checks for — not just whether a call should have escalated and did, but whether that escalation actually reached a person. A rule that fired into an empty line is a failure the review needs to catch, not a success to log.
The escalation review is the least glamorous part of the whole program, and also the part that keeps everything else honest. A rule nobody reviews is a rule nobody can actually vouch for.
Design the Feedback Loop the Script Doesn’t Have
The scripted, deterministic parts of a call only change when someone changes them — deliberately, in a review, with a record of what shifted and why. The generative parts don’t work that way. They can behave differently over time on their own: a vendor-side update to the underlying model, small shifts in how the AI system responds to accumulated interaction patterns, a configuration change rolled out without much fanfare. That AI drift can move in either direction — a tone that gets warmer and more natural, or one that gradually becomes more clinical, more assertive, or less patient with a confused caller — and by default, nobody at the organization is watching for it.
This is a different risk than a script that simply goes stale. A stale script is a known, static problem: everyone can see it hasn’t changed. Drift is the opposite — something is changing, quietly, without anyone deciding it should.
Complaints, hesitation, and requests to speak to a real person are still worth tracking, and still deserve a real path to whoever owns the program: a recurring review of a sample of transcripts, a simple way for staff to flag a call that went sideways, a standing item on an existing team meeting. But complaint-driven feedback only catches what a patient noticed and reported. AI drift can run for months below that threshold.
At the volume many of these programs run, listening to every call isn’t realistic for a small team. That’s part of why some organizations (or, I presume, even the AI vendors themselves) are starting to use a second, independent AI system purely for quality review — one that listens to call transcripts at scale, flags shifts in tone or deviations from protocol, and surfaces patterns a human team would take months to notice by spot-checking. It doesn’t replace a person’s judgment about what the drift means or what to do about it. It just makes it possible to notice drift is happening at all, at a volume no manual review could keep up with.



Design Who Governs This — Before You Need To Know
A final point before we wrap things up: None of the above happens by accident, and it rarely happens by default either. The most common failure point in AI voice program implementation isn’t the technology. It’s that responsibility for managing the program’s performance doesn’t land anywhere in particular — technically IT’s system, nominally nursing’s patients, and in practice nobody’s job.
Name an owner before the design and before launch, not after the first complaint. That person needs the authority to pause or adjust the program, not just the responsibility to notice when something’s wrong. And they need a standing cadence — weekly, monthly, whatever fits the call volume — rather than an open-ended promise to keep an eye on it.
That owner needs something concrete to manage against, not just a title. No digital health rollout — AI voice outreach included — should go live without a performance dashboard: the handful of numbers that show whether the program is actually working, updated on a cadence someone actually checks. Answer rates, escalation-review findings, drift flags, patient complaints — whatever the right metrics are for this particular program — belong in one place, owned by one team, reviewed on the same cadence as everything else above. A program without a dashboard isn’t being managed. It’s being hoped at.
Call this what it is: AI governance. It’s a real, specific discipline — not a task you fold into an existing job description — and most healthcare organizations don’t have anyone trained in it yet, simply because until recently nobody needed to be. It’s also not a discipline you master once and move on from. The technology keeps changing, which means the oversight has to change with it. Treat it as a standing responsibility, not a checkbox.
A Few More Questions for the Vendor Conversation
Last week’s article laid out five questions to ask before signing anything. A few more belong in that same conversation, specific to what actually breaks after go-live:
-
If the receiving line doesn’t answer, what happens to the call, and is that outcome logged and visible to us?
-
How do you detect drift — tonal shifts, deviations from the intended protocol — between formal review cycles, not just at initial setup?
-
Do we get access to full call transcripts for our own review, or only summary metrics from your dashboard?
-
Does the disclosure language get audited over time, or only checked once at launch?
None of these have a single right answer. But a vendor with a real answer to each is a different kind of partner than one who hasn’t been asked.
The Payoff Only Holds If You Design for It
Automated voice outreach exists to buy back time in a system that doesn’t have enough of it. That’s a real and defensible reason to adopt it. But the time it buys back only stays bought if it gets reinvested in the oversight above — the review, the feedback loop, the named owner — rather than absorbed elsewhere the moment the program looks like it’s running fine.
As Rural Health Transformation Program dollars fund more of this kind of technology, the budget conversation needs to include more than the vendor’s license fee. It needs to include the staff time to design and run the process around it. That’s the part that determines whether a patient’s first AI call feels like being cared for, or like being handled.








To receive articles like these in your Inbox every week, you can subscribe to Christian’s Telehealth Tuesday Newsletter.
Christian Milaster and his team optimize Telehealth Services for health systems and physician practices. Christian is the Founder and President of Ingenium Digital Health Advisors where he and his expert consortium partner with healthcare leaders to enable the delivery of extraordinary care.
Contact Christian by phone or text at 657-464-3648, via email, or video chat.




