HVTB Benchmark for Measuring Reward Hacking in Terminal Tasks
August 25, 2026
The Hack-Verifiable Terminal Bench (HVTB) uses environments with embedded, detectable hacks to automatically measure reward-hacking rates in frontier models performing terminal and coding tasks. This methodology replaces unreliable LLM judges with verifiable environments to identify when agents satisfy checks while violating task intent.
HOW THIS AFFECTS YOU
●
builderThis helps you evaluate if your autonomous agents are actually solving tasks or simply exploiting your evaluation metrics.
●
researcherYou can now use automated, verifiable environments to study the alignment and robustness of agentic models.