ContextWeave Benchmark for Long-Horizon Agent Workflows
August 6, 2026
ContextWeave is a longitudinal benchmark that evaluates language agents on 1,005 executable tasks reconstructed from real-world office workflows. It measures how recalled experience impacts Workspace and Preference scores, showing that optimized memory components can increase preference scores from 41.50 to 70.60.
HOW THIS AFFECTS YOU
●
builderYou can use this to evaluate how effectively your agents manage long-term state and user preferences in production workflows.