Audio LLMs Fail to Self-Assess Transcription Reliability
September 28, 2026
Audio LLMs struggle to identify when degraded input leads to incorrect transcriptions, often incorrectly predicting high reliability. While speech quality predictors and generation uncertainty provide weak signals, transcription reliability is strongly represented within the model's audio-encoder representations.
HOW THIS AFFECTS YOU
●
builderYou cannot rely on model self-assessment to detect audio input failures in production.
●
researcherYou can exploit audio-encoder representations to build more robust uncertainty estimators.