Protocol-Level Identifiability Audit for LLM Reasoning Benchmarks
August 14, 2026
This audit framework tests whether observation protocols actually identify intended behavioral properties by checking if they separate distinct policy classes. The study demonstrates that base-only observations can collapse multiple deterministic policies into a single equivalence class, masking true performance.
HOW THIS AFFECTS YOU
●
researcherYou should verify that your evaluation protocol's support is sufficient to distinguish between different model behaviors.
●
policyThis highlights why benchmark scores can be misleading regarding actual model reasoning capabilities.