Extracting Concept Content via LLM Activations and Linear Probes
August 10, 2026
This method uses Recursive Feature Machine (RFM) and linear probing on frozen LLM activations to measure internal concept knowledge, specifically in financial ESG contexts. It outperforms embedding and surface-level text baselines by reading internal model judgments rather than just token distributions.
HOW THIS AFFECTS YOU
●
researcherYou can use activation probing to bypass surface-level text biases when measuring internal model knowledge.