TRACE Benchmark for Human-AI Controller Coordination Failures
August 10, 2026
TRACE is a multi-layer benchmark designed to localize how drift and failures propagate through the human-AI control loop. Using 1,918 drifted traces from the ALFRED benchmark, it tracks errors across state, observation, decision, rules, and control layers.
HOW THIS AFFECTS YOU
●
researcherYou can now diagnose exactly which layer of a control loop is responsible for system-wide coordination failures.
●
policyYou can use these time-aligned traces to better understand the safety implications of drift in cyber-physical systems.