HANDBOOK.md Benchmark Tests Agent Compliance with Long Policy Docs
July 29, 2026
The HANDBOOK.md benchmark evaluates whether LLM agents can follow complex, long-context instructions in enterprise environments. It uses 65 tasks to measure if agents adhere to binding policy documents over extended tool-use horizons.
HOW THIS AFFECTS YOU
●
builderThis shows that system prompts and policy files may not be sufficient to constrain agent behavior reliably.
●
researcherYou can use this to evaluate if your models actually follow instructions or just complete tasks.