HappyWorld-Bench Evaluates Consistency in Video and Spatial World Models
September 20, 2026
HappyWorld-Bench assesses the reliability and responsiveness of generated worlds through a hierarchical framework of six capabilities. The benchmark includes 1,138 video prompts, 300 spatial scenes, and 254 embodied test cases to evaluate video, spatial, and embodied world models via human A/B comparisons.
HOW THIS AFFECTS YOU
●
researcherYou can use this to evaluate if your world models maintain consistency during agent interaction.