Censored Rating Scales Manufacture Bias in LLM-Judge Audits
August 28, 2026
A pre-registered audit reveals that difference-in-differences designs on bounded rating scales can manufacture artificial effects. Observed statistics often confound actual preference with differential attenuation caused by unequal distances from the scale's endpoints.
HOW THIS AFFECTS YOU
●
researcherYou must account for scale censorship when using LLM-as-a-judge to certify model biases.
●
policyBe cautious of audit results that claim significant bias based on bounded scoring systems.