Engineering

Giving an avatar real tools

Tool calling in a live voice conversation is a latency problem before it is a capability problem. Here is how we cover the gap.

April 28, 2026
6 min read
Engineering

Text agents can afford to think. A voice agent cannot, three seconds of dead air on a video call is read as a dropped connection, not as deliberation.

We handle this with two mechanisms. First, tools are declared with an expected duration, and anything slow enough triggers a spoken acknowledgement, a short line generated from the tool's name rather than a canned filler, while the call resolves in the background.

Second, a tool result is injected mid-turn rather than starting a new one, so the model receives it as a continuation of the utterance it is already speaking and the answer arrives as one sentence instead of two.