FM-Bench Evaluates Long-Horizon Decision Making in LLM Agents
August 20, 2026
FM-Bench is a new benchmark for long-horizon agentic management, requiring models to run a football club for 20 simulated years. It uses a deterministic engine to evaluate cumulative decisions across 400 decision stops without relying on an LLM judge.
HOW THIS AFFECTS YOU
●
builderUse this benchmark to test if your agents can maintain consistency in complex, multi-step environments.
●
researcherThis provides a more rigorous evaluation of agentic persistence and cumulative error rates than short-horizon benchmarks.