NARU Benchmark for Japanese Long-Form Video Narrative and Culture
August 12, 2026
NARU provides a dataset of 1,481 questions across 155 Japanese videos totaling 146.8 hours to evaluate narrative evolution and cultural nuance. The benchmark uses a hierarchical memory-based annotation pipeline to structure event and social context information.
HOW THIS AFFECTS YOU
●
researcherYou can use this to evaluate how well long-context models handle high-context, non-English cultural reasoning.