AI Briefing
KO

Changes in an LLM's internal state determine its answering tendencies

·2026.06.24 01:34

Key point

A study has found that depending on the content of the input text, an LLM's internal hidden state shifts, changing the tendency of its answers.

Details

It has been confirmed that an LLM's internal state (hidden state) can function as an alignment policy traversal vector that determines the model's answering tendencies.

According to the study, even when the content of the question itself does not change, the nature of the neutral pre-token text provided immediately before the question causes the model's internal state to shift into different regions, which directly affects how the model answers (direct answers vs. avoidance/hedging).

The main characteristics are as follows:

  • Invisible change: No change is detected at the level of the Input and Output text, but a measurable change occurs in the model's internal hidden state.
  • Weights unchanged: This is a different mechanism from fine-tuning or jailbreaking, where the model's weights change.
  • State transition: The act of reading a certain text induces the model into a specific state, like an 'assistant's room,' which determines how it responds to subsequent questions.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.