Audio Language Models Show Divergent Feature Representations for Speech and Text
September 25, 2026
Analysis of six models across 15 languages reveals that audio and text streams rarely represent distinctive phonemic features in the same direction. Only the Qwen2.5-Omni models showed consistent representation for the voicing feature across both modalities.
HOW THIS AFFECTS YOU
●
researcherYou should account for misalignment between acoustic and linguistic embeddings when training unified multimodal decoders.