CoVA-SFT Dataset Enables Visual Abstraction Reasoning in Multimodal Models
August 28, 2026
CoVA-SFT provides a structured corpus of 51.9K samples with 222K reasoning steps to teach models how to build internal visual workspaces. It uses five layout families across 17 complex tasks to bridge the gap between text-only chain-of-thought and multimodal spatial reasoning.
HOW THIS AFFECTS YOU
●
builderYou can leverage these structured reasoning steps to improve the spatial intelligence of your multimodal applications.
●
researcherThis dataset offers a new way to train models on multi-step spatial and layout-based reasoning.