Identifying the Latent Evidence-Credit Gap in Visual Reasoning
September 27, 2026
Latent visual reasoning in MLLMs suffers from a gap where latent tokens respond weakly to image perturbations that change the final answer. This suggests that standard GRPO training with final-answer rewards provides insufficient guidance for supervising latent-token behavior.
HOW THIS AFFECTS YOU
●
researcherYou should look beyond final-answer rewards when supervising latent reasoning steps.