Reward-Informed SAEs Capture Solution Completeness Instead of Reasoning Quality
August 28, 2026
Training JumpReLU Sparse Autoencoders on high versus low reward GRPO trajectories for Llama-3.1-8B creates feature separation, but this separation primarily tracks solution completeness. Structural cues like length and boxed answers achieve an AUC of 0.70, suggesting the SAEs pick up on formatting rather than reasoning depth.
HOW THIS AFFECTS YOU
●
researcherYou should be cautious when using reward signals to steer SAE feature interpretability, as they may encode trivial structural artifacts.