LLM Judges Fail to Distinguish Tutoring From Answer-Giving
July 31, 2026
An audit of tutor models shows that general-purpose helpfulness rubrics cannot distinguish between pedagogical guidance and direct answer-giving. While Claude Opus 4.8 scores policies similarly for helpfulness, the pedagogy rubric achieves perfect rank separation with a Cliff's delta of 1.0.
HOW THIS AFFECTS YOU
●
builderAvoid relying on standard LLM-as-a-judge helpfulness scores for fine-tuning tutoring applications.
●
researcherYou should use specialized pedagogy rubrics rather than general helpfulness metrics when evaluating educational agents.