Sparse Readout Prism explains logit-lens via sparse features
August 31, 2026
Sparse Readout Prism (SRP) decomposes the unembedding matrix to explain token logits through sparse readout features rather than relying on corpus-dependent token readings.
HOW THIS AFFECTS YOU
●
researcherThis provides a more robust method for mechanistic interpretability that is independent of the fitting corpus.