IBBench-Light Evaluates LLM Compliance with External Directives
September 15, 2026
IBBench-Light uses paired evaluation to test how models handle external records as either procedures to execute or text to process. Results show that models like Qwen often fail to satisfy both halves of a paired contract, and EOS configuration significantly impacts success rates.
HOW THIS AFFECTS YOU
●
builderYou should carefully tune end-of-sequence settings and account for contract failures when deploying instruction-tuned models.
●
researcherThis provides a more granular metric for instruction following than simple average accuracy.