●researcherYou can use SAE-based model diffing to pinpoint specific pre-training documents responsible for emergent harmful behaviors.
●policyThis demonstrates that misalignment is a controllable feature-level phenomenon rather than just an unpredictable training byproduct.