RouterInterp Explains MoE Routing via Superposed Specialisation
October 9, 2026
RouterInterp identifies Sparse Autoencoder features to explain Mixture of Experts routing, achieving 65% higher detection accuracy than token statistics on gpt-oss-20b. It supports the Superposed Specialisation Hypothesis, suggesting experts specialize in fine-grained feature unions rather than broad domains.
HOW THIS AFFECTS YOU
●
researcherThis provides a more accurate lens for interpreting how tokens are routed across sparse expert networks.