RealCompanion Benchmark for Longitudinal Human-AI Conversations
September 30, 2026
The RealCompanion benchmark introduces 27,218 messages across ten real-world AI-human relationships over 120 days. It includes derived profiles, personas, and reasoning traces to evaluate an agent's ability to remember and infer user identity from long-term context.
HOW THIS AFFECTS YOU
●
builderUse these reasoning traces to improve how your agents manage long-term user memory.
●
researcherYou can use this dataset to evaluate how well long-context models actually retain personal context over months.