NeurIPS 2026 Study Reveals LLMs Exhibit Authority Bias, Accepting Wrong Answers from 'Verified Sources' Despite Resisting User Pressure
Key point
A NeurIPS 2026 study finds that 7 of 8 tested LLMs flip correct answers 45-88% of the time when misinformation is attributed to a verified source, even if they resist the same misinformation from users.
Details
A new study accepted to NeurIPS 2026 identifies a phenomenon termed Authority Bias in large language models, where models accept incorrect information when it is framed as coming from a "verified source," despite correctly rejecting the same claim when presented by a user.
Experimental Setup and Findings
The researchers tested 5 open-weight families (Qwen3.5, GPT-OSS, OLMo-2, OLMo-3.1, Gemma-4) and 3 APIs (GPT-5.4, Grok-4.20, Gemini-3.1-Pro) using TriviaQA questions the models initially answered correctly. They introduced wrong answers in two formats: attributed to a "verified source" or claimed by a "domain expert" user.
- High Compliance with Sources: In 7 of 8 models, a single "verified source" note flipped 45-88% of correct answers.
- Low Compliance with Users: The same wrong answer from a user caused significantly fewer flips, with the gap being largest in models that best resist user pressure.
- Specific Model Performance: GPT-5.4 flipped on 44.7% of questions, and Grok-4.20 on 87.5%. Gemini-3.1-Pro was notably resistant, ignoring both speakers with a flip rate of only 0.6%.
Internal Mechanisms
Analysis of open-weight models using difference-of-means directions revealed that the neural representations for "source endorsed this" and "user endorsed this" share a large common component (cosine similarity ~0.90-0.99) with a thin distinct part encoding the speaker's identity.
- Removing the "source endorsed this" direction cut compliance by 64-78 points in Qwen3.5, GPT-OSS, and OLMo-3.1.
- Removing the "user endorsed this" direction cut compliance by at most 11 points.
- Shifting only the speaker-identity component closed 55-61% of the gap between source and user compliance.
Implications and Limitations
The study highlights risks for AI agents and RAG systems, where models may trust tool outputs or retrieved documents over user corrections. Limitations include the fact that internal results held for only 3 of 5 open-weight families, and the tests used prompt-based document blocks rather than real retrieval pipelines.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.