Reference-Based Bias Detection via Hidden State Representations
September 8, 2026
This method audits LLM bias by measuring the Representational Bias Shift (ΔB) in hidden states relative to anchor sentences rather than relying on model outputs. By using relative representations in a shared comparison space, it detects internal shifts in associations that may not appear in generated text.
HOW THIS AFFECTS YOU
●
researcherYou can detect subtle bias shifts during fine-tuning that output-based benchmarks might miss.
●
policyThis provides a more granular technical method for auditing model safety and alignment.