Abliteration increases model optimism without improving accuracy
July 29, 2026
Removing censorship through abliteration techniques correlates with increased model confidence and optimism, specifically in stock market prediction tasks. Testing on Gemma and Qwen showed that while confidence levels shifted, task accuracy remained at baseline coinflip levels.
HOW THIS AFFECTS YOU
●
researcherYou should account for shifts in model personality and confidence when applying safety-removal techniques.
●
policyYou must recognize that uncensored models may present incorrect information with higher confidence.