Root Cause of Hallucinations in Multimodal Model Identified
Key point
It was revealed that the hallucination phenomenon in which the multimodal model Inter-1 repeatedly outputs a specific phrase on silent videos is caused by a combination of prompt examples and post-training.
Details
A hallucination phenomenon was discovered in which the multimodal model Inter-1 repeatedly outputs a specific phrase ("Yeah, Friday at five.") even on videos without audio. Upon investigation, it was found that this was not simple data contamination but the result of two technical factors combined.
The first cause was an example phrase within the System Prompt. A worked example written to specify the model's output format was included during the prompt optimization process, and this served as a 'script' for the model to reference.
The second cause was a tendency formed during the Post-training stage of the model. Experimental results showed that changing the example phrase in the prompt immediately changed the model's hallucinated phrase as well, and even without a prompt, the post-trained model showed a tendency (Compulsion) to utter something instead of staying silent.
This is analyzed as a text/in-context variant of the 'Clever Hans effect', in which the model generates fictional content based on text within the context window even in situations lacking visual information.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.