BACKTRACE framework detects gap between agent skill attribution and actual influence
July 31, 2026
The BACKTRACE framework and BACKROOMBench evaluation test whether language agents actually utilize declared skills or merely exhibit 'skill theater.' By comparing skill-conditioned answers against no-skill counterfactuals through intervention on skill identity and content, the study identifies a systematic gap between stated reasoning and actual decision influence.
HOW THIS AFFECTS YOU
●
builderDo not trust an agent's self-reported reasoning when evaluating the effectiveness of tool-use integrations.
●
researcherUse this framework to evaluate if your agent's reasoning traces are actually driving its outputs.