Voice AI Warm Transfer: How the Human Handoff Works
A warm transfer decides whether your voice AI feels competent or broken. How the proxy call, briefing, and merge actually work, what the double billing window costs on Vapi and Retell, and the failure modes to design for.

Every voice AI deployment eventually hits a call it cannot finish, and what happens in the next thirty seconds decides whether the whole system feels competent or broken. The fix is a warm transfer: the AI privately briefs a human agent, merges the calls, and steps back, so the caller never repeats a word. This post covers how that handoff actually works under the hood, what the double-billing window costs on current pricing, and the failure modes that quietly ruin it.
If you are still scoping a voice AI rollout, our voice AI customer service implementation guide covers the broader build. This post goes deep on the one moment most builds get wrong.
Cold transfer vs warm transfer
A cold transfer is a routing event. The AI decides it cannot resolve the call, forwards it to a queue, and hangs up its own involvement. The human agent answers knowing nothing beyond maybe a queue name, and the caller starts over from "hi, my name is."
A warm transfer carries context with the call. Before the caller and the human ever hear each other, the AI has already told the agent who is calling, what they want, what has already been tried, and why the call is being escalated. The conversation continues instead of restarting.
This distinction matters more than any measure of how well the AI itself talks. An iQor survey of consumers found that 77% considered repeating information three times on a support call more frustrating than being stuck in airport security, and 81% of those whose information was not carried forward said it delayed their resolution. A cold transfer manufactures exactly that experience. You can have a brilliant AI front-end and still deliver the single most hated interaction in customer service at the moment of handoff.
How a warm transfer actually works
The mechanics are more involved than "forward the call." Bland's warm transfer documentation lays out the full sequence, which is representative of how modern platforms do it:
- Trigger. The AI determines a human is needed and tells the caller it will connect them.
- Proxy call. While the caller waits on hold with music, the AI places a second, separate outbound call to the human agent.
- Briefing. The AI delivers a short spoken briefing to the agent: caller name, issue, what has been tried. This can be a static script or generated dynamically from the conversation with variables like
{{call_summary}}. - Agent preparation. Optionally, the agent can ask the AI clarifying questions or pull up records, then explicitly trigger the merge when ready.
- Merge. The two calls are bridged into a three-way conference. Hold music ends for the caller.
- Introduction. The AI announces the handoff to both parties ("Grace, I have our billing specialist on the line, they are up to speed on your invoice issue").
- Exit. The AI drops off. The caller and agent talk. The caller never repeated anything.
Underneath this, the integration layer follows a SIP and webhook pattern: the AI holds the original call on a bridge using SIP re-INVITE commands, fires a webhook carrying a structured context payload to the CRM or helpdesk, and the receiving agent's softphone gets the incoming call leg and a screen-pop at the same moment. Call and data arrive together. That synchronization is what separates a warm transfer from a forwarded call plus an email summary nobody reads.
One detail worth knowing: Telnyx shipped warm transfers for its Voice AI Agents in September 2025, and the context it passes includes user inputs, assistant responses, detected intents, and collected data. This is now table stakes across serious platforms, not a differentiator. If a vendor cannot describe their merge and briefing flow in this level of detail, treat that as a red flag.
The trigger layer: when the AI should escalate
The hardest design decision is not the transfer mechanics. It is deciding when to fire them. Escalate too readily and you have built an expensive phone menu. Escalate too reluctantly and callers rage-quit while the AI loops.
Production systems use five trigger categories, drawn from Retell's handoff documentation and industry practice:
- Explicit request. The caller asks for a person. Always honor this immediately; arguing with it destroys trust.
- Confidence floor. The AI's confidence in resolving the call drops below a set threshold, or it has misunderstood the caller twice in a row.
- Out-of-scope intent. The request matches a topic the AI is not configured to handle, like a refund above its approval limit.
- Regulatory or sensitivity rules. Hard-coded escalations that override everything else. A healthcare agent might escalate any medication question regardless of confidence.
- Elapsed time. If no resolution within a configured window, escalate rather than loop. This is the trigger teams most often forget.
Benchmarks cited by VoiceInfra recommend keeping total transfer rates below 15 to 20 percent of all calls. If you are consistently above that band, your trigger thresholds are too loose or the AI's resolution scope is too narrow, and the answer is usually to widen what the AI can do, not to make transfers faster.
The context payload: what the human should see
The spoken briefing to the agent covers the first seconds. The written payload covers everything after. Retell's engineering blog gives a good example of the summary style: "Caller attempting to reschedule appointment. No availability found in the requested time window. Escalating to human scheduler." Short, factual, actionable.
The operational standard for the screen-pop is a three-tier layout:
- Top: a two or three sentence summary. Who is calling, what they need, why the AI escalated. The agent reads this while the call connects.
- Middle: a structured data block. Account number, order ID, verified identity status, CRM-ready fields. Pre-populated, not copy-pasted by the agent.
- Bottom: the full transcript. Available on demand, almost never read.
Timing is as important as content. SigmaMind's analysis points out a failure mode teams miss in testing: a summary that generates correctly but lands on the agent's screen three or four seconds late still produces dead air on the line, right at the moment the caller expected to be connected. The payload must be ready before the merge completes, not generated after the trigger fires.
What a warm transfer costs
Here is the part vendor pages skip: during the briefing window, you are paying for two concurrent calls. The caller is on hold on one leg, and the AI is talking to your agent on the second leg. Both bill.
A worked example with current pricing. Say you handle 3,000 inbound calls a month, averaging six minutes, with an 18% transfer rate: 540 transfers. Assume the briefing window runs about 45 seconds per transfer.
- On Vapi: the hosting fee is $0.05 per minute plus provider costs passed through at cost (roughly $0.01 for transcription, $0.02 for the model, $0.02 for the voice, depending on choices), so around $0.10 per minute per leg. Two legs times 0.75 minutes times 540 transfers comes to roughly $80 a month in transfer-window cost.
- On Retell: the pay-as-you-go base is $0.07 to $0.31 per minute all-in, with a typical config around $0.11. Same math: two legs times 0.75 minutes times 540 transfers is roughly $90 a month.
- Telephony: the outbound proxy leg to your agent's number bills at your telephony provider's rates on top, whether that is Vapi or Retell telephony or your own Twilio or Telnyx trunk.
The honest conclusion: the minutes are almost noise. Under $100 a month at meaningful volume. The real cost of the handoff is the human agent's time on the other end, which is exactly why the briefing and payload matter. A confused agent spending three minutes re-gathering context costs more than the entire transfer-minute budget. For the full build and run cost picture beyond the transfer window, see our voice AI implementation cost breakdown.
One pricing gotcha to check before you commit: on Bland, warm transfer is gated behind enterprise accounts, per their own docs. If warm transfer is a day-one requirement, confirm it exists on the plan you are actually buying, not the plan in the demo.
Common failure modes
These are the ways warm transfers break in production, drawn from vendor troubleshooting docs and the mechanics above:
- The AI says it will transfer, then does nothing. The model generated the acknowledgment speech but never invoked the transfer tool. Vapi's transfer docs call this out explicitly: the prompt must require the acknowledgment and the tool invocation in the same response.
- Nobody answers, and there is no fallback. The agent is at lunch, the proxy leg rings out, and the caller is left on hold until they hang up. Every transfer needs a timeout (Bland supports 30 to 3,600 seconds) and a defined fallback: voicemail with a summary attached, a callback offer, or routing back into the AI's flow.
- Caller ID screening kills the proxy call. Your agent sees an unknown number and declines it. Configure the proxy number to be a familiar one; note that on Bland it must come from the same Twilio account as the parent call.
- The AI hangs up on the phone tree. The human agent sits behind an IVR or extension. You need DTMF sequences (
w54means wait half a second, then dial extension 54) and queue-patient behavior: answering machine detection and wait-for-greeting, which Bland enables by default for exactly this reason. - Wrong destination. With multiple transfer destinations, the assistant picks one based on each destination's description. Overlapping descriptions produce misroutes. Keep descriptions distinct and mirror them in the system prompt.
- Silence during the window. Dead air reads as a dropped call. Hold music plus a two-part transition message from the AI, a reason and a time anchor ("I am connecting you with a specialist, this takes about 30 seconds"), keeps callers on the line.
When warm transfer is the wrong choice
Warm transfer assumes a human is available, awake, and worth the interruption. Design something else when that assumption fails:
- After-hours calls. If nobody is rostered, every transfer fails by definition. Take a structured message with a callback commitment instead.
- Low-stakes, high-volume calls. A caller asking for opening hours does not need a human even if the AI stumbles. Retry with simpler language before escalating.
- When the escalation is really a task. If the caller needs a refund processed or an appointment moved, the better fix is often giving the AI the tool to do it, not a faster path to a person.
- Very small teams. If your one agent is also the person who answers the door, a callback queue with a rich context payload beats a live transfer that always goes to voicemail.
A well-designed voicemail or callback fallback is not a consolation prize. For many callers, "we have all your details and Sarah will call you back within the hour" beats five minutes on hold.
A pre-launch checklist
Before you put a warm transfer in front of real callers, verify each of these with a live, unscripted call:
- Trigger test: each of your five trigger categories fires when it should, and only then.
- Transfer rate: in testing, transfers stay under 20% of calls. If not, widen AI scope before launch.
- Payload timing: the screen-pop arrives before the merge completes, not after.
- No-answer path: ring timeout, voicemail, and callback all behave as designed.
- Caller experience: transition message, hold music, and a working merge announcement.
- Logging: transfer states (started, merged, no answer, timed out) flow to your CRM so you can measure abandonment and resolution.
- Kenya and East Africa note: destination numbers must be in E.164 format (
+254...for Kenyan mobiles), and if you are transferring into an office PBX, use a SIP URI destination and test DTMF extension dialing early. Local number portability and caller ID presentation behave differently across carriers, so test the proxy leg against the actual phones your team carries.
If you would rather have this engineered for you, handoff design is a core part of our custom voice and automation builds, and it slots into the broader AI voice agent programs we write about elsewhere.
Frequently asked questions
What is the difference between a warm transfer and a cold transfer? A cold transfer forwards the caller with no context, so the agent starts blind and the caller repeats everything. A warm transfer briefs the human agent first, passes a written context payload, and merges the calls only when the agent is ready.
How long does a warm transfer take? The briefing window typically runs 30 to 60 seconds: dial, agent answer, briefing, optional agent questions, merge. Set a hard timeout (60 to 90 seconds is common) with a fallback, because an unanswered proxy leg should never leave the caller holding indefinitely.
Does warm transfer cost extra? The feature itself is usually included, though Bland gates it behind enterprise plans. The real cost is the double-billing window: while the AI briefs your agent, both call legs bill per minute. At 540 transfers a month that works out to roughly $80 to $90 on current Vapi or Retell pricing, plus telephony for the outbound leg.
Can the AI transfer to another AI instead of a human? Yes. Vapi has a separate handoff tool for assistant-to-assistant moves, and Telnyx supports chaining specialized AI agents with context passed along. Useful for routing from a generalist to a specialist agent, though human escalation remains the common case.
What happens if the human agent does not answer? Whatever you designed, which is the problem if you designed nothing. Configure a hold timeout, then fall back to voicemail with the AI's summary attached, a scheduled callback, or a return to the AI flow. Log the failed transfer so you can staff against demand.
Do callers mind being transferred by an AI? They mind repeating themselves and being dropped into silence. Research consistently shows the frustration is about lost context, not automation itself. A warm transfer with a clear transition message and a briefed agent removes both causes.
Where to go from here
Warm transfer is a small feature with an outsized effect on whether callers trust your voice AI at all. Get the triggers, the payload, and the no-answer fallback right, and escalation stops being a failure and becomes part of the service. If you want a handoff built and tuned rather than theorized about, talk to us about our AI employee and voice agent work, or start with the implementation guide linked above.
About AI Agents Plus Editorial
AI automation expert and thought leader in business transformation through artificial intelligence.



