Small Models Grade Open-Ended Exams via Explicit Rubrics
August 19, 2026
Small language models achieve grading reliability comparable to frontier models when using explicit rubrics. In the any-to-bench framework, judge identity explains only 0.2% of score variance, whereas answer identity accounts for 95.6%, suggesting low-cost models can replace expensive ones for rubric-based assessment.
HOW THIS AFFECTS YOU
●
builderYou can significantly reduce inference costs by using small models for rubric-based evaluation tasks.
●
founderThis enables you to build high-scale automated grading products with much higher margins.