ReFract Benchmark for Agent Perspective Awareness in Physical Environments
October 5, 2026
ReFract introduces a 150-entry benchmark to evaluate whether LLM agents can infer user roles and respect knowledge boundaries. This prevents agents from performing actions or providing information that exceeds a specific user's authority or capability in high-stakes physical settings.
HOW THIS AFFECTS YOU
●
builderYou can use this to test if your agents act appropriately when interacting with different user permission levels.
●
policyThis helps identify safety risks when agents perform irreversible physical actions without proper role calibration.