SpeechConversationBench Evaluates Multi-Turn Reasoning in Speech-to-Speech Models
October 1, 2026
SpeechConversationBench measures spoken mathematical reasoning across incremental information shards using 103 GSM8K problems. Results show commercial systems lose 5.0-25.3% accuracy when information is disclosed across turns, while the proprietary LEGO pipeline achieves 77.5% accuracy.
HOW THIS AFFECTS YOU
●
builderThis highlights the performance gap in speech systems when handling information revealed over multiple turns.
●
researcherYou can use this to study how conversational context management affects spoken reasoning performance.