SAEVerbalizer Generates Natural Language Explanations for Sparse Autoencoder Features
August 14, 2026
SAEVerbalizer injects SAE decoder directions into LLM representations and fine-tunes downstream layers to generate natural language explanations. This method generalizes to unseen features and transfers across different SAE dictionaries and LLM architectures using a lightweight adapter.
HOW THIS AFFECTS YOU
●
builderYou can implement lightweight adapters to interpret hidden representations in production models.
●
researcherYou can use this to move beyond behavioral observation for feature interpretability.