Introduces a training source and benchmark using 36,076 examples from video-pretrained generative models to teach visual transition reasoning. It aims to bridge the deficit in spatial, embodied, and temporal reasoning in MLLMs.
HOW THIS AFFECTS YOU
●
builderThis provides a pathway to more capable embodied AI and robotic agents.
●
researcherYou can use this dataset to train MLLMs on physical and temporal causality.