Claude Fable 5.1 Leads FrontiersWE v2 Long-Horizon Benchmark
September 29, 2026
The updated FrontiersWE v2 benchmark evaluates models on 34 ultra-long-horizon tasks. Claude Fable 5.1 currently holds the top position, followed by GPT-5.6 and GLM-5.3.
HOW THIS AFFECTS YOU
●
builderPrioritize Claude Fable 5.1 for applications requiring extended task execution and complex reasoning chains.
●
researcherUse these long-horizon metrics to evaluate agentic reasoning capabilities beyond short-context tasks.