Chat templates trigger self-referential disclaimers in LLMs up to 9B parameters
September 27, 2026
Chat templates act as a functional switch that modulates whether models use disclaimer-heavy language or experiential phrasing. Analysis of 8 open-source instruct models shows that activation steering can reproduce these shifts in self-referential voice caused by template changes.
HOW THIS AFFECTS YOU
●
researcherYou can use activation steering to manipulate how models report their own identity or capabilities.