Investigation of open-weight models reveals a distributed causal mechanism responsible for distinguishing animate from inanimate concepts. Unlike highly localized circuits, this animacy circuit shows partial generalization across models and tasks.
HOW THIS AFFECTS YOU
●
researcherThis adds to the growing body of work on mechanistic interpretability by identifying how high-level semantic concepts are represented.