Gemma 3 4B Distinguishes Impossible from True Statements via Activations
August 14, 2026
An activation study of Gemma 3 4B shows that while the model conflates contingent falsehoods with contradictions in text, its internal representations differ. A linear truth probe achieves 0.93 AUC for separating impossible from true statements, while an impossibility probe hits 1.00 AUC on held-out families.
HOW THIS AFFECTS YOU
●
researcherYou can exploit specific layer activations, particularly around layer 15, to detect logical impossibility.
●
policyThis suggests that model outputs can be deceptive regarding the model's internal understanding of truth.