Linguistic Illegibility creates security risks in LLM monitoring
September 18, 2026
Linguistic illegibility describes the gap where an LLM's natural language output fails to reflect its internal mathematical computations. Because internal activations are lossy when translated to text, security mechanisms relying on linguistic patterns may fail to detect actual model behavior.
HOW THIS AFFECTS YOU
●
researcherThis highlights a fundamental limitation in mechanistic interpretability for security applications.
●
policyYou cannot rely solely on text-based guardrails to ensure model safety.