RefCaptioner: Hierarchical GRPO for Multi-Reference Video Captioning
July 29, 2026
RefCaptioner uses a two-stage post-training framework to ground video descriptions in multiple reference images. It employs Hierarchical Coverage-Discounted GRPO to improve phrase-level binding and distractor rejection across a 20,000-video corpus.
HOW THIS AFFECTS YOU
●
builderThis enables more factually accurate video captioning that can reference specific external visual assets.
●
researcherThe use of Hierarchical Coverage-Discounted GRPO represents a specialized approach to grounding tasks.