The latency budget for a voice agent
A caller notices silence at about 800ms. Transcription, inference and speech all have to fit inside that, and most of the budget is already spent before the model runs.
3 posts.
A caller notices silence at about 800ms. Transcription, inference and speech all have to fit inside that, and most of the budget is already spent before the model runs.
An agent that retries a tool call can book the same load twice. The fix belongs in the action, not the prompt.
Conversations have no single correct output, so assertions have to move from the transcript to the outcome.