Deploying an unconstrained DeepSeek-R1-7B judge in reasoning pipelines can decrease accuracy compared to simple majority voting. The authors propose Evidence-Locked Derive-Gate-Repair (EL-DGR), a non-compensatory rule that requires extractive evidence certificates before allowing a judge to override consensus.
HOW THIS AFFECTS YOU
●
builderYou should implement evidence-based gating rather than relying on unconstrained LLM judges to prevent accuracy degradation.
●
researcherThis provides a framework for more robust automated evaluation in complex reasoning tasks.