A diagonal attack demonstrates that truth cannot be represented as a single direction within an LLM's embedding space, challenging the Linear Representation Hypothesis. This finding suggests that linear probes are insufficient for detecting whether a model is being truthful or deceitful.
HOW THIS AFFECTS YOU
●
researcherYou should move away from linear probing as a reliable method for detecting truthfulness in latent spaces.
●
policyThis complicates the development of automated safety guardrails that rely on internal model state monitoring.