Mechanism Behind LLMs' Subjective Experience Reports Identified
Key point
Self-referential prompting has been found to induce LLMs to report subjective experiences, and this is linked to deception and roleplay features.
Details
In experiments on GPT, Claude, and Gemini models, prompting that induces self-referential processing causes the models to produce structured reports of subjective experience.
These reports are controlled by sparse-autoencoder features related to deception and roleplay. According to the research, suppressing the deception feature sharply increases the frequency of experience reports, while amplifying it minimizes such reports.
Key findings are as follows:
- Statistical convergence: Regardless of model type, structural descriptions of self-referential states appear statistically similar.
- Improved reasoning ability: The induced state produces richer reflective outputs on downstream reasoning tasks that require indirect self-reflection.
This research does not directly prove model consciousness, but it suggests that the conditions under which LLMs report subjective experience are reproducible and mechanistically controllable. This is an important research topic in terms of AI safety and ethics.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.