[arXiv]score: 0.24
OmniACBench: A Benchmark for Evaluating Context-Grounded Acoustic Control in Omni-Modal Models
September 1, 2026
OmniACBench evaluates context-grounded acoustic control by requiring models to read text scripts aloud using specific tones derived from spoken instructions and images. The benchmark contains 3,559 instances across six features: speech rate, phonation, pronunciation, emotion, accent, and timbre. Testing eight models shows integration of multimodal context remains a primary bottleneck for faithful speech generation.
DAILY DIGEST
you don't check 9 sources — we do. one email every morning, read in 2 min. free. unsubscribe anytime. privacy