NanoGPT Speedrun benchmarks agent performance on human record gap
August 22, 2026
Comparative benchmarks show Fable 5 closing 81.7% of the human record gap in approximately 2.7 days, outperforming Opus 5 and Kimi K3. The dataset tracks various models including GPT-5.6 Sol and DeepSeek V4 Pro across different trajectory types and agent configurations.
HOW THIS AFFECTS YOU
●
researcherYou can use these comparative agentic benchmarks to evaluate model performance on complex, multi-day tasks.