Feature Nonlocality (FNL) measures the semantic abstractness of Sparse Autoencoder (SAE) features using the entropy of normalized per-position influence. The metric correctly identifies context-dependent reasoning features over token-driven ones in 73% to 84% of tested cases.
HOW THIS AFFECTS YOU
●
researcherYou can use FNL to distinguish between surface-level lexical features and high-level semantic features during mechanistic interpretability studies.