Voice AI Latency on Phone Calls: Where the Milliseconds Go
A voice agent that answers correctly but late feels broken. Here is the full latency budget with real numbers, the East Africa network problem, how to measure p95 mouth-to-ear delay, and the fixes ranked by payoff.

A voice agent that answers correctly but two beats late feels broken, and callers hang up on it. The fix is rarely a bigger model. It is knowing your latency budget: the milliseconds each stage of the pipeline adds between the caller finishing a sentence and the agent starting its reply. This post gives you that budget with real numbers, shows why callers in East Africa face a steeper bill than callers in Ohio, and walks through measuring and fixing each hop.
The short answer: target 800ms to 1.1 seconds of mouth-to-ear latency, measure it at the 95th percentile, and spend your effort on endpointing and network topology before you touch the model.
What callers actually experience
Latency on a voice call has a precise definition: mouth-to-ear delay, the time from when the caller stops speaking to when the first byte of the agent's audio reply reaches their ear. Everything else, transcription time, model time, synthesis time, is a component of that single number.
Telephony has a benchmark for this. The ITU-T G.114 recommendation puts one-way delay under 150ms as generally acceptable for conversation and anything over 400ms as unacceptable. That standard was written for person-to-person calls, where both sides are human and forgiving. For a voice agent, Twilio's engineering guidance on core latency in AI voice agents frames the practical target differently: a response gap of roughly 800ms to 1.1 seconds keeps the conversation feeling natural. Beyond about 1.2 seconds, callers start repeating themselves, saying "hello?", or talking over the agent's reply, which then derails the next turn.
Note the gap between those two numbers. Humans tolerate a 150ms pause in a person-to-person call, yet a well-built voice agent gets away with 800ms because callers unconsciously adjust their expectations when they know they are talking to a machine. They do not adjust twice. You get one slow turn forgiven; the third slow turn ends the call.
The latency budget, hop by hop
A standard cascaded voice agent has four stages, and each one bills you. Using the ranges published in Speko's breakdown of voice agent latency:
- Endpointing: 200 to 700ms. The system waits to confirm the caller has actually finished speaking before it acts on the transcript. This is voice activity detection with a silence threshold, and it is the most underestimated cost in the whole pipeline.
- Speech-to-text finalization: often folded into endpointing, but streaming STT still needs time to commit a final transcript after the last audio frame.
- LLM time to first token: 300 to 800ms, depending on model size, prompt length, and provider load.
- Text-to-speech time to first byte: 150 to 400ms before the first audio chunk exists.
- Transport: 50 to 150ms for the audio to travel between telephony edge, your server, and back to the caller.
A worked example with mid-range numbers:
- Endpointing: 500ms
- LLM first token: 600ms
- TTS first byte: 250ms
- Transport: 150ms
- Total: 1,500ms. Callers feel this. They start talking over the reply.
The same pipeline, tuned:
- Endpointing: 250ms (semantic endpointing instead of a fixed silence timer)
- LLM first token: 350ms (smaller model, shorter prompt, streaming enabled)
- TTS first byte: 180ms (streaming synthesis, cached voice)
- Transport: 80ms (media and inference in the same region)
- Total: 860ms. This sits comfortably inside the natural range.
Same models, same vendors. The difference is configuration and topology, which is why the first optimization is never "buy a faster model."
The geography tax: phone calls from East Africa
Everything above assumes your caller, telephony edge, and inference are close together. If your callers are in Nairobi or Kampala, that assumption fails at the network layer.
Look at Twilio's published list of edge locations: the United States, Ireland, Frankfurt, Tokyo, Singapore, Sydney, Sao Paulo, and a handful more. There is no African edge. A call from a Safaricom number enters the provider's network at the nearest edge, which for East Africa means Europe. Your media has already crossed a continent before your agent sees a single audio frame.
Then consider where the inference runs. Most STT, LLM, and TTS endpoints sit in US regions. A typical deployment for a Kenyan caller looks like this: audio travels Nairobi to Frankfurt to be ingested, then to a US region for transcription, back to a US region for the model, to the TTS provider, and the reply audio retraces the whole path to the caller's phone. As a rule of thumb, budget roughly 100ms of round-trip time for each intercontinental leg and verify with your own measurement using the procedure below. Two or three such legs per conversational turn can quietly add 300ms or more to every response, before any model has done any work.
Mobile networks add jitter on top. A caller on 4G in a matatu has a variable last-mile link, and jitter buffers trade delay for audio quality automatically.
The mitigations, in order of impact:
- Terminate media as close to the caller as possible. Pick the telephony edge nearest to your callers, not nearest to your server.
- Run inference in the same region as that edge. If media lands in Frankfurt, your STT, LLM, and TTS should be in Europe too. Many providers let you pin a region; OpenAI's realtime models connect over WebRTC, WebSocket, or SIP, so the transport choice is yours to make deliberately.
- Use WebRTC where you can. For app-based or web-based calling, WebRTC's congestion handling beats a PSTN hairpin. Frameworks like LiveKit's open-source agents stack are WebRTC-native and bridge to phone networks over SIP when you need real phone numbers.
- Accept a floor. For PSTN calls from East Africa into European infrastructure, a total budget under about 1 second is achievable but takes the tuning described above. Plan for 1.1 to 1.3 seconds as the realistic starting point, then optimize down.
Speech-to-speech models: what you gain and what you pay
Speech-to-speech models collapse the cascaded pipeline. Audio goes in, audio comes out, no separate STT and TTS stages. OpenAI's gpt-realtime is the current general-availability example, and it removes two vendor round trips plus the STT finalization wait from every turn.
The cost math is public. OpenAI's realtime pricing documentation bills audio at 1 token per 100ms of caller audio and 1 token per 50ms of assistant audio, at $32 per million input audio tokens and $64 per million output audio tokens for gpt-realtime. A worked minute:
- Caller speaks 30 seconds: at 600 tokens per minute that is 300 tokens at $32 per million, about $0.01.
- Agent speaks 30 seconds: 600 output audio tokens at $64 per million, about $0.04.
- Text tokens for instructions and tool calls add a smaller amount on top.
Independent derivations land in the same place: Forasoft's breakdown of realtime API pricing puts a typical agent minute at roughly $0.06 to $0.11 on the full model and $0.02 to $0.05 on the mini once prompt caching is working. Compare that with a cascaded stack where STT, a text LLM, and TTS are billed separately and can run well under $0.05 per minute combined. Deepgram's comparison of its STT pricing against alternatives is a useful reference point for the cascaded side.
The trade-off is not just price. Speech-to-speech gives you lower latency and far better handling of interruptions, tone, and barge-in. What you lose is control: no independent choice of STT language model for Swahili or Sheng code-switching, no separate TTS voice vendor, and a text transcript you have to request rather than one you get for free. For Kenyan deployments where STT accuracy on local accents is the binding constraint, a cascaded pipeline with a chosen STT can outperform a speech-to-speech model that treats accent handling as an internal detail.
How to measure before you optimize
You cannot fix what you have not timed. Cresta's engineering write-up on real-time voice latency makes the key measurement point: track percentiles, not averages. A p50 of 700ms with a p95 of 2.4 seconds means one caller in five has a bad experience, and averages hide that completely.
A practical measurement procedure:
- Log four timestamps per turn: last audio frame from caller, transcript finalized, first LLM token, first TTS audio byte sent. Every serious framework exposes these events; LiveKit's agent framework emits them as metrics you can export.
- Compute stage durations: endpointing is timestamp two minus one, LLM is three minus two, synthesis is four minus three. The remainder to the caller's ear is transport, which you estimate from SIP or WebRTC round-trip stats.
- Record p50, p95, and p99 per stage, per hour of day. Provider endpoints slow down under load; your 2am numbers flatter you.
- Run test calls from the real network. If your callers are on Kenyan mobile networks, test from a Kenyan SIM, not from a browser in Frankfurt. The last mile is part of the budget.
- Set an alert on p95 mouth-to-ear over 1.4 seconds. That is the threshold where abandonment starts climbing.
One day of this instrumentation tells you exactly which hop to fix first. Without it, teams guess, and they usually guess "the model," which is wrong about half the time.
The fixes, ranked by effort and payoff
Work through this list in order. Each item is cheaper than the one below it.
- Enable streaming everywhere. Streaming STT, streaming LLM tokens into streaming TTS. If any stage waits for the full output of the previous one, you are paying hundreds of milliseconds for nothing. This is usually a configuration change, not code.
- Tune endpointing. Shorten the silence threshold where the domain allows it, or move to semantic endpointing that uses what was said, not just how long the pause was. Speko's latency guide shows endpointing is often the single largest hop.
- Colocate media and inference. Same region for telephony edge, STT, LLM, and TTS. For East African callers, that region is usually in Europe. This is an architecture decision, so make it before you build, not after.
- Shrink the prompt. Every 500 tokens of system prompt adds to time to first token. Move reference material into retrieval or cached context.
- Pick a faster model for the first response. Have the agent open with a fast model and escalate to a stronger one only for turns that need it. Most turns in a support call are simple.
- Consider speech-to-speech if interruptions and naturalness matter more than per-minute cost and vendor control.
- Last resort: change vendors. Only after the above, because vendor swaps are weeks of work and the problem is usually in your topology.
Common mistakes
These show up in almost every first deployment:
- Measuring averages. Averages hide the p95 tail that actually hangs up calls.
- Testing from the wrong network. Your office fiber in Westlands is not your customer's 4G link in Kisumu.
- Ignoring endpointing. Teams spend weeks on model choice and leave a 700ms silence timer untouched.
- Letting the agent monologue. Long replies multiply TTS and transport time per turn. Design for short turns.
- Hairpinning audio through your app server. Relay audio directly between telephony and inference where the platform allows it. A TwiML media stream looks like this, and the WebSocket URL should point at infrastructure near the caller:
<Response>
<Connect>
<Stream url="wss://your-agent-endpoint.example/media" />
</Connect>
</Response>
- Assuming the demo number is the production number. Vendor demos run on warm caches, quiet hours, and nearby regions.
When a voice agent is the wrong choice
Honesty saves money. Skip voice, or keep a human in the loop, when:
- The task needs long, precise output. Reading out a 12-digit reference number or a list of eight options is painful at any latency. Send an SMS summary instead and keep the call short.
- Your callers have very poor connectivity. If the last mile adds 400ms of jitter, no pipeline tuning rescues the experience. Chat or USSD will serve them better.
- The conversation is high-stakes and adversarial. Disputed insurance claims or collections calls with angry customers punish every awkward pause. Latency reads as evasion.
- Volume is tiny. Under a few hundred calls a month, the engineering cost of getting latency right exceeds the cost of a person answering.
In those cases a chat-based AI employee or a human with AI assistance is the better build.
Frequently asked questions
What is good latency for a voice AI agent? Aim for 800ms to 1.1 seconds of mouth-to-ear delay at the median, and keep the 95th percentile under 1.4 seconds. Twilio's engineering guidance puts that range at the edge of what feels like natural conversation. Beyond it, callers talk over the agent and abandonment climbs.
Why does my voice bot pause so long before answering? The pause is almost never one thing. The usual suspects, in order: a long endpointing silence threshold, an LLM with a heavy prompt and no streaming, a TTS call that waits for the full sentence, and intercontinental network hops between telephony and inference. Instrument the four timestamps described above and the slow hop will identify itself within a day.
Is speech-to-speech faster than STT plus LLM plus TTS? Yes, by design. Collapsing three stages into one model removes two vendor round trips per turn. The trade-offs are higher per-minute cost, less control over transcription and voice, and weaker handling of language-specific needs like Swahili code-switching.
Can a voice agent work for callers in Kenya? Yes, but plan for the geography. There is no African edge location on major telephony platforms, so media routes through Europe and every intercontinental hop adds to the budget. Terminate media in the nearest edge, colocate inference in the same region, and test from real Kenyan mobile networks.
What does a low-latency voice agent cost per minute? On speech-to-speech models like gpt-realtime, roughly $0.06 to $0.11 per agent minute at current pricing, or $0.02 to $0.05 on the mini with caching. Cascaded stacks billed as separate STT, LLM, and TTS calls can come in lower but pay for it in extra round trips. Telephony itself is additional, billed per minute by the carrier or CPaaS.
Is WebRTC better than a normal phone call for voice AI? Where you control the endpoint, yes. WebRTC gives you congestion control, better jitter handling, and no PSTN hairpin. Phone calls still win on reach: everyone has a phone number. Most production deployments bridge both, using SIP to connect the phone network to a WebRTC-native stack like LiveKit.
Where to go from here
If you take one thing from this post, take the measurement procedure: four timestamps per turn, p95 per stage, tested from your callers' real network. That single afternoon of instrumentation tells you whether your problem is endpointing, the model, or the geography, and each of those has a different fix with a different price tag. We build voice agents and the surrounding automation as part of our custom builds work, and our voice AI implementation cost breakdown covers the budget side of the same decision. For the wider picture of what these agents can do once they respond fast enough to be trusted, see how businesses are using AI voice agents and our voice AI customer service implementation guide. If you want a second pair of eyes on a slow agent, talk to an engineer.
About AI Agents Plus Editorial
AI automation expert and thought leader in business transformation through artificial intelligence.


