Label-Free Strategies Fail to Fix MCQ Evaluation Bias
August 11, 2026
Testing multiple-choice question (MCQ) benchmarks using generation-then-matching and isolation-scoring reveals that preventing models from seeing option labels does not improve accuracy. The findings suggest the primary bottleneck in MCQ evaluation is withholding the options themselves.
HOW THIS AFFECTS YOU
●
researcherYou should be skeptical of MCQ scores as reliable measures of model knowledge due to inherent positional and structural biases.