SPAR-K Accelerates Spoken Language Model Inference via Early Exit
August 28, 2026
SPAR-K reduces speech decoding depth in interleaved models like GLM-4-Voice using a modality-aware alternating-depth schedule. The framework maintains ASR accuracy within a 0.82% margin while significantly decreasing average decoding depth for speech tokens.
HOW THIS AFFECTS YOU
●
builderYou can reduce latency and compute costs for speech-to-speech models without sacrificing significant accuracy.
●
researcherThe modality-aware exit schedule offers a new way to handle interleaved text and speech token distributions.