NARU Benchmark Evaluates Narrative and Cultural Nuance in Japanese Video
August 14, 2026
The NARU benchmark introduces 1,481 questions across 146.8 hours of Japanese long-form video to test narrative reasoning and cultural understanding. It uses a hierarchical memory-based annotation pipeline to evaluate how models track evolving social meanings and implicit contexts.
HOW THIS AFFECTS YOU
●
researcherYou can now evaluate how well long-context video models handle high-context, non-English cultural nuances.