All perspectives

Voice AI

What 'sub-second latency' actually means on a customer call

The gap between 300ms and 1.5 seconds decides whether a call feels human or robotic. Here's where the delay hides in a voice AI pipeline, and why it's felt before it's heard.

What 'sub-second latency' actually means on a customer call
Voice AI / xAIa
Author
The xAIa Team
Published
18 June 2026
Reading time
4 min read

Pick up a phone, say "hello," and count how long you'll wait for a reply before it starts to feel wrong. It isn't long. Somewhere around a second of silence and the other person feels absent. You say "hello?" again. You wonder if the line dropped. On a call, silence is loud, and it's the one thing you can't paper over.

This is why latency is the whole ballgame for voice AI, and why "sub-second" isn't marketing garnish. The gap between the customer finishing their sentence and the AI beginning its reply is the single thing that decides whether the call feels like a conversation or like talking to a machine that's buffering.

The number that decides everything

Human conversation runs on a tight clock. In natural back-and-forth, the gap between one person stopping and the next starting is roughly two hundred milliseconds. A fifth of a second. We don't notice it because it's the water we swim in. It's just what "normal" sounds like.

Stretch that gap and things degrade fast. At half a second the pause is noticeable but forgivable. At a full second the caller starts to wonder. At a second and a half they either repeat themselves or talk over the reply, and now two voices collide and the whole exchange stumbles.

So the target isn't "fast." The target is a specific window: get the response started inside roughly a second, ideally well under, because that's the threshold where a human brain stops hearing "delay" and starts hearing "conversation."

Latency isn't a technical metric the customer never sees. It's the first thing they feel, before they've registered a single word.

Where the milliseconds actually go

The reason sub-second is hard is that a lot has to happen in that window, in sequence, before a single word comes back. Roughly three stages:

  1. Speech to text. The customer's audio has to be transcribed into words the system can reason about. This runs while they're still speaking, but it can't fully finish until they stop.
  2. Reasoning. The system has to understand what they meant, check whatever it needs to check (an account, a balance, a booking), and decide what to say. This is where an AI that has to call out to a slow database or a distant server quietly bleeds time.
  3. Text to speech. The reply has to be turned back into natural, spoken audio, in the right voice and dialect, and start playing.

Every one of those stages costs time, and naively they stack. Add them up in a straight line and you're well past a second before anyone hears anything. Getting under the threshold means overlapping them: starting to reason before the sentence is fully done, beginning to speak before the whole reply is generated, and cutting the distance the data has to travel.

That last point matters more than people expect. If the audio has to fly to a server on another continent and back for each stage, physics alone can eat your budget. Round trips add up. Holding the processing in-region isn't only a data-residency nicety in the GCC. It's shorter wires, and shorter wires are fewer milliseconds.

Why the phone is the hardest place to hide

You can get away with lag in a chat window. The customer types, glances at their phone, looks back, and a two-second delay reads as "thinking," not "broken." Voice gives you none of that cover.

On a live call there's no typing indicator, no spinner, no "AI is composing." There's just sound, or the absence of it. And absence, on a phone, means something specific and bad: the line went dead, the person left, something is wrong. The medium turns every extra half-second into anxiety.

It's harder still in the Gulf's real conditions, where callers code-switch between Arabic and English mid-thought and interrupt freely. A system that needs a long beat to process each turn falls apart the moment a caller changes their mind halfway through a sentence, which, in a real conversation, they do constantly.

The delay you can't point to but can feel

Here's the subtle part. A caller will almost never say "your latency is 1.4 seconds." They don't have the vocabulary and they shouldn't need it. What they'll say instead is that it felt "robotic," or "awkward," or that they "couldn't tell if it heard me." That's latency wearing a disguise.

Get the timing right and the same technology gets described in completely different words. Callers say it felt "natural." They forget, mid-call, that they're talking to software. Not because the words got smarter, but because the rhythm got human. Two hundred milliseconds is invisible. And invisible, on a phone call, is exactly the goal.


Want to hear sub-second voice AI in your customers' own dialect? Book a demo.

AI for every department. Starting with yours.

A quick chat to answer questions and see if we can help.

Book a consultation