CalibratedRubric uses Bayesian filtering and Item Response Theory (IRT) to build task-adaptive rubric banks for LLM evaluation. This framework improved human-gold agreement on JudgmentBench from $\kappa=0.604$ to $0.743$ by filtering for rubric measurability.
HOW THIS AFFECTS YOU
●
builderYou can implement this to build more reliable automated evaluation pipelines for your generative outputs.
●
researcherThis offers a mathematically grounded method for scaling expert-level qualitative evaluations.