Whisper Fine-Tuning Enables Text Recovery From Low-Pass Filtered Speech
October 9, 2026
Fine-tuning the Whisper model on the 12 lowest Mel bins (450Hz cutoff) achieves a 36% Word Error Rate for prosody-to-text prediction. The method recovers 10% of utterances perfectly and reaches a 79% next-token accuracy when provided with the correct prefix.
HOW THIS AFFECTS YOU
●
builderYou can explore using extremely low-bandwidth audio signals for lightweight text-based intent or emotion signaling.
●
researcherThis establishes a new baseline for studying the lexical information contained in low-frequency speech prosody.