RT-OPD Improves Acoustic Grounding in Audio-Language Models
September 25, 2026
Reward-Tilted On-Policy Distillation (RT-OPD) uses log-probability contrast between teacher predictions with and without audio to strengthen acoustic reliance in compact audio-language models. This method prevents models from relying on textual shortcuts to answer audio-based questions.
HOW THIS AFFECTS YOU
●
researcherYou can use this technique to reduce linguistic bias in multimodal distillation tasks.