[arXiv]score: 0.18
Reference-Based Bias Detection in LLMs via Relative Representations of Hidden States
September 10, 2026
Detecting bias through hidden-state geometry allows for auditing model variants without relying on costly output benchmarks. By encoding sentences via similarities to fixed anchor points, this method measures Representational Bias Shift to track how target groups associate with attributes. The metric correlated with output-level bias changes in 15 of 18 tested settings.
DAILY DIGEST
you don't check 9 sources — we do. one email every morning, read in 2 min. free. unsubscribe anytime. privacy