PACT-SLM Evaluates Spoken Agent Action Timing and Accuracy
October 1, 2026
The PACT-SLM framework assesses when streaming spoken agents possess sufficient speech evidence to act. Using 1,600 prefix predictions and a WavLM-based probe, the study reveals that models trigger actions on 18.99% of prefixes before sufficient semantic evidence is present.
HOW THIS AFFECTS YOU
●
builderThis helps you diagnose if your voice agents are acting prematurely on noisy or incomplete audio inputs.
●
researcherYou can use this framework to measure the latency-accuracy trade-off in streaming speech models.