Open-Weight Models Match Frontier Performance in Math Proof Grading at 100x Lower Cost
August 4, 2026
Small models including GPT-OSS 120B, DeepSeek-V4 Flash, and Gemma-4 31B achieve human-level agreement on IMO-GradingBench math proofs. Using a unanimous agreement consensus rule provides high precision at a fraction of the cost of Claude Opus 4.7 or Gemini 3.1 Pro.
HOW THIS AFFECTS YOU
●
builderYou can replace expensive frontier LLM judges with cheaper open-weight ensembles for math reasoning evaluation.
●
founderThis allows you to scale automated evaluation pipelines while significantly reducing operational overhead.