RUMBA Benchmark Evaluates Long-Term Conversational Memory in Russian and English
July 24, 2026
RUMBA introduces a fine-grained taxonomy for testing long-term memory through timestamped user-assistant dialogues. It requires models to perform retrieval, combination, and reasoning across multiple sessions using temporal information and varying semantic scopes.
HOW THIS AFFECTS YOU
●
researcherYou can use this to diagnose specific failures in temporal reasoning and cross-session memory retrieval.