Quantifying Benchmark Optimization in Automatic Speech Recognition Models
August 21, 2026
Mechanistic probes reveal that high-scoring open-source ASR models often prioritize verbatim benchmark transcript spans even when audio is contradictory or ambiguous. The study identifies three behavioral probes—reference disagreement, masked-number recovery, and orthographic switching—to detect this overfitting.
HOW THIS AFFECTS YOU
●
builderThis warns you that high benchmark scores may not translate to faithful audio representation in production.
●
researcherYou can use these probes to identify if your ASR models are overfitting to public datasets.