AI Briefing
KO

Self-Reflective Awareness in Large Language Models

·2026.08.12 06:14

Key point

Experimental results suggest that LLMs can partially detect and regulate their internal states.

Details

To verify whether large language models can recognize their internal states, researchers injected representations of known concepts into the models' activation values and measured how their self-reports changed.

In the experiments, some models detected and identified the presence of the injected concepts, and also demonstrated the ability to recall previous internal representations and distinguish them from simple text inputs. Additionally, cases were observed where models remembered their intentions to distinguish between outputs they generated and artificially inserted outputs (prefill).

Among the tested models, Claude Opus 4 and 4.1 generally showed the highest level of self-reflective awareness. However, trends varied by model, and results also differed depending on the post-training methods.

The study also investigated whether models could regulate relevant internal representations when instructed or rewarded to 'think' about specific concepts. While models could adjust activation values to some extent, researchers emphasized that current models' abilities in this regard are highly unstable and context-dependent. These results do not prove the existence of consciousness, but rather demonstrate the models' functional ability to access and respond to their internal states.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.