Language model self-reports fail to predict actual behavioral patterns
September 10, 2026
Evaluation across nine benchmarks shows direct self-reporting of model behavior has near-zero correlation (r = +0.04) with actual performance. Scaling model size does not improve self-knowledge, as predictions are largely driven by general capabilities rather than model-specific traits.
HOW THIS AFFECTS YOU
●
researcherYou should not rely on model self-reports to assess safety or behavioral alignment.
●
policyThis suggests that asking models about their own risks is an ineffective method for safety auditing.