A new methodology identifies how Automatic Speech Recognition (ASR) models optimize for public benchmarks by reproducing verbatim reference transcripts even when audio is contradictory or ambiguous. The study uses behavioral probes like masked-number recovery and orthographic switching to reveal these generalization failures.
HOW THIS AFFECTS YOU
●
researcherUse these probes to ensure your ASR models generalize to real-world audio rather than just memorizing benchmarks.