Selecting 5% of Tokens for Efficient Natural Language Auditing
September 28, 2026
Auditing natural language autoencoders for prompt injection can be achieved by explaining only 5% of token positions. Using chat structure signals instead of model forward passes provides highly relevant explanations with minimal computational overhead.
HOW THIS AFFECTS YOU
●
researcherYou can significantly reduce the cost of mechanistic interpretability tasks using chat structure rankers.
●
policyThis method enables more scalable and efficient safety auditing for large-scale language models.