●builderYou can use this benchmark to measure how well your agents handle complex, end-to-end assistant workflows rather than simple tool calls.
●researcherThis provides a more rigorous evaluation for agentic synthesis and spatial reasoning than current retrieval-only benchmarks.