SAEScientist-Bench Tests Autonomous Mechanistic Interpretability Research
September 7, 2026
SAEScientist-Bench evaluates if AI agents can autonomously conduct mechanistic interpretability research using Sparse Autoencoders (SAEs). Agents are tested on their ability to navigate a 131K+ feature dictionary in Gemma-2-9B-IT to discover target concepts.
HOW THIS AFFECTS YOU
●
researcherYou can use this to test the limits of automated feature discovery and model auditing.
●
policyThis explores the feasibility of autonomous safety monitoring through automated mechanistic interpretability.