STEREODISCO Framework Identifies Stereotypical Semantic Axes in LLM Activations
July 31, 2026
STEREODISCO uses WordNet antonym synsets to construct 2,000 semantic axes, recovering them as geometric axes in LLM activation spaces via probing. The framework identifies stereotypical associations in Llama-3-8B and Mistral-7B through statistical tests on concept projections.
HOW THIS AFFECTS YOU
●
researcherYou can use this method to probe internal representations for latent biases beyond standard semantic axes.
●
policyThis provides a more systematic way to audit model safety and social alignment.