Validating Human Value Geometry in LLM Steering Space
September 4, 2026
This study investigates whether activation steering vectors encode coherent semantic structures related to human values using a 26K-sample benchmark. It analyzes distribution-driven methods like CAA and ODESteer to determine if they reflect Schwartz's Theory of Basic Human Values.
HOW THIS AFFECTS YOU
●
researcherYou can evaluate if steering vectors exploit behavior-specific shortcuts or capture actual semantic value structures.
●
policyThis helps determine the reliability of lightweight, inference-time behavioral controls in sensitive contexts.