Full-Duplex Speech-to-Speech Architecture for Tool Use
September 18, 2026
This architecture enables low-latency, full-duplex speech models to execute tool calls by using a duplex STT frontend that emits delegation tokens to a text-based LLM backend. The system achieves 92-97% tool-call recall and 81.2% rejection accuracy while maintaining natural turn-taking and interruption handling.
HOW THIS AFFECTS YOU
●
builderYou can integrate complex tool-calling capabilities into voice agents without sacrificing duplex conversational latency.
●
designerThis allows for more natural, interruptible voice interactions that can still perform functional tasks.