Study Finds Chat Templates Control LLM Self-Referential Voice, Not Model Weights
Key point
Researchers identified a specific activation direction that toggles between disclaimer and experiential self-reports in open-source LLMs.
Details
A new study demonstrates that the chat template acts as a switch for how Large Language Models describe themselves, rather than these descriptions reflecting intrinsic model knowledge. When the chat template is present, models increase disclaimer voice (e.g., "I'm just an AI") and decrease experiential voice (e.g., "I feel"). Conversely, removing the template reduces disclaimers and increases experiential language.
Activation Steering Mechanism
The researchers analyzed 8 popular open-source instruct models up to 9B parameters and identified a specific direction within the activation space of 3 models that controls this behavior:
- Removing this direction suppresses disclaimer voice.
- Adding this direction induces disclaimer voice, even in models running without a chat template.
- Random directions of the same magnitude have little to no effect.
Implications for AI Safety and Research
The findings suggest that LLM self-reports are partially determined by deployment configuration (the chat template) rather than solely by model weights. This introduces a significant confound for researchers studying model introspection or AI safety, as self-descriptions should not be treated as literal facts about the model's internal state.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.