A last-write resolver achieves 80.25% accuracy on MemoryAgentBench, suggesting many perceived selective forgetting errors are actually reachability failures. Long-context models score significantly higher on reachable items than the benchmark's BM25 agent, which only hit 6.0% on specific multi-hop failures.
HOW THIS AFFECTS YOU
●
researcherYou should re-evaluate MemoryAgentBench results by distinguishing between memory erasure and graph reachability issues.