Option-Channel Attack bypasses typed decision model guardrails
October 9, 2026
Evaluating seven open-weight models used as agent guardrails reveals accuracy between 36% and 72% for screening prompt injections and tool calls. The study identifies fail-open vulnerabilities where models allow prohibited actions, often triggered by minimal text input.
HOW THIS AFFECTS YOU
●
builderYou should not rely on small typed decision models as sole security layers for agent tool calls.
●
policyThis highlights significant reliability gaps in using LLMs as automated governance layers for agentic systems.