The Answerable Working Memory (AWM) framework introduces a diagnostic to ensure VLM agents retain enough evidence in working memory to answer questions without the original page context. It incorporates this signal into GRPO rewards.
HOW THIS AFFECTS YOU
●
builderImplementing AWM-style rewards can help reduce 'hallucinated' answers in long-document agents.
●
researcherThis targets a critical blind spot in how we evaluate agentic memory.