Harness-IF Benchmark Exposes Instruction-Following Deficiencies in Coding Agents
August 13, 2026
Harness-IF evaluates coding agents by scoring rules against execution evidence across five surfaces. The new Against-Prior Accuracy (AP-Acc) metric reveals that frontier models drop 3.6 to 7.4 points in accuracy when instructions conflict with their unprompted defaults.
HOW THIS AFFECTS YOU
●
builderYou should test your agents against rules that oppose their default behaviors to ensure true compliance.
●
researcherThis provides a more rigorous metric for distinguishing intentional instruction following from coincidental task success.