[arXiv]score: 0.14
Towards Audio Token Compression in Large Audio Language Models
August 21, 2026
Reducing audio token density from 25 tokens/s using unsupervised segmentation and uniform average pooling decreases input sequences by up to 3x. Integrating low-rank adapters during finetuning recovers performance in ASR and speech-to-speech translation tasks, maintaining accuracy near frame-level models while lowering attention complexity for LLM decoders.
DAILY DIGEST
you don't check 9 sources — we do. one email every morning, read in 2 min. free. unsubscribe anytime. privacy