Multimodal Models Fail Acousto-Kinematic Word Inference Benchmark
August 18, 2026
The Unwritten Benchmark evaluates models on inferring written words from only the audio of pen scratches and video of hand movements. While humans achieve over 80% accuracy, GPT-4o and Gemini 2.5-Pro fail to surpass 10% performance.
HOW THIS AFFECTS YOU
●
researcherCurrent multimodal architectures lack the abstract perceptual reasoning required for dynamic, generative process inference.