TWIST Benchmark for Conversational Memory Intervention Quality
September 25, 2026
TWIST evaluates how memory systems handle belief changes through four tracks: tension detection, draft vetting, supersession history, and sensitive recall. The benchmark uses surface-matched hard negatives to prevent models from gaming scores via over-detection.
HOW THIS AFFECTS YOU
●
builderUse this to measure if your agent's memory system correctly updates user beliefs without losing historical context.
●
researcherThis provides a more rigorous evaluation of memory interventions beyond simple retrieval and recall metrics.