The latency budget for a voice agent
A caller notices silence at about 800ms. Transcription, inference and speech all have to fit inside that, and most of the budget is already spent before the model runs.
In a phone conversation, a pause longer than roughly 800ms reads as the other party having stopped listening. That is the entire budget, end to end, and it is not negotiable — a broker will talk over you or hang up.
Where it goes#
caller stops speaking
├─ endpointing 150–300ms deciding they are actually done
├─ final transcript 50–100ms flushing the ASR
├─ inference 200–600ms time to first token
├─ speech synthesis 100–200ms time to first audio
└─ network/jitter 50–150ms round trip
Add the floor and you are at ~550ms before the model has produced anything useful. Add the ceiling and you are well past the point the caller notices.
Endpointing is the first real decision#
Waiting for confident silence is the safe choice and the expensive one. Too eager and you interrupt; too patient and every turn feels slow.
The thing that helped most was not tuning the threshold. It was making endpointing semantic: a trailing "and the pickup is..." is obviously unfinished regardless of how long the pause runs.
Streaming changes the arithmetic#
Nothing should wait for a complete prior stage:
| Stage | Waits for | Should wait for |
|---|---|---|
| Inference | Final transcript | Partial transcript, revised as it firms up |
| Synthesis | Full completion | First clause |
| Playback | Full audio | First chunk |
Streaming all three turns a sum into a maximum, which is the difference between a 900ms turn and a 400ms one with identical components.
The model being fast is less important than nothing in the pipeline blocking on something it did not strictly need.
What we stopped optimising#
Token throughput. Once the first clause is out and audio is playing, generation is comfortably ahead of speech. Spending effort there bought nothing a caller could hear.