CLAP Metric Fails to Capture Attribute Binding in Music-Text Models
September 28, 2026
Testing reveals that the standard CLAP score fails to reliably distinguish between music descriptions when instrument attributes are swapped. While large audio-language models perform better, contrastive music-text models struggle to capture specific bindings like timbre or lead versus accompaniment.
HOW THIS AFFECTS YOU
●
researcherYou should look beyond simple cosine similarity scores when evaluating how well music models follow complex text prompts.