PANORAMA: Panoptic Grounded Captioning with Mask Selection
September 15, 2026
PANORAMA introduces PanoCaps, a benchmark for panoptic grounded captioning that requires models to describe both foreground objects and background regions with pixel-level masks. This addresses the current gap between fluent captioning and accurate spatial grounding.
HOW THIS AFFECTS YOU
●
researcherThis provides a new metric for evaluating how well VLMs associate text with specific image segments.
●
designerThis capability is essential for building tools that require precise visual-to-text spatial understanding.