FrontierMath Erdös Benchmark Reveals Low AI Math Proficiency
September 23, 2026
The FrontierMath Erdös (FME) benchmark tests AI on 68 open mathematical conjectures within the Lean proof assistant. Under a $300 per-problem budget, GPT-6 Astra achieved only a 3% success rate, while all other tested models scored 0%.
HOW THIS AFFECTS YOU
●
researcherThis sets a high-bar evaluation metric for formal mathematical reasoning capabilities.
●
investorThe results suggest that even frontier models still face significant scaling hurdles in formal automated reasoning.