Internal Feature Analysis Reveals Overconfidence Bias in Qwen3-4B
August 20, 2026
An interpretability study on Qwen3-4B shows that uncertainty is implemented as a sparse override of a broad coalition of certainty-generating features. This mechanism leads to verbalized overconfidence, particularly when the model is prompted to provide numeric confidence scores.
HOW THIS AFFECTS YOU
●
researcherYou can use these transcoder feature identification methods to study the causal drivers of model uncertainty.
●
policyThis highlights a fundamental reliability risk in how models communicate confidence to end-users.