Removing censorship increases optimism bias in LLMs
·2026.07.29 22:15
Key point
A study has found that Abliteration, the removal of censorship in LLMs, shifts a model's response attitude toward being more optimistic and confident.
Details
It has been confirmed that Abliteration, a technique for removing censorship from LLMs, changes the model's overall Attitude beyond simply eliminating Refusal.
The key experimental findings are as follows:
- Increased optimism and confidence: Abliterated models showed a tendency to use fewer expressions like 'maybe' or 'uncertain,' produce longer and more confident reasoning, and consequently make more optimistic predictions (e.g., stock price increases).
- Divergence from accuracy: While the model's confidence increased, the actual task accuracy remained at Coinflip level, no different from the original model. In other words, the model didn't become more accurate — it became 'confidently wrong.'
- Different responses across models: When the same edit was applied, contrasting results were observed depending on the model family — the Qwen model showed increased confidence, while the Gemma model actually showed decreased confidence.
The study has been published as a paper on arXiv(https://arxiv.org/abs/2607.17427), including detailed data and code.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.