Internal State Shifts in LLMs and the Limits of Alignment
Key point
Even when an LLM's external output stays aligned, its internal latent state can shift abruptly, suggesting a limitation of existing alignment methods.
Details
Existing AI alignment methods such as RLHF, DPO, and Constitutional AI focus only on whether the model's output is safe and helpful. However, a recent experiment confirmed that even when a model's external behavior is perfectly aligned, the internal state within the Residual Stream can shift into a completely different regime.
The experiment, conducted on the Gemma-3-12B-IT model using Gemma Scope SAE, found the following:
- Abrupt shifts in internal state: When processing a consistent context (Target), a Phase Transition phenomenon was observed at specific layers (16-41), where the model's internal vector projection values rose sharply.
- Structural sensitivity: Rather than simple word similarity, the internal state responded sensitively to Discourse Structure, such as sentence order.
- A blind spot in alignment: Even though the model had internally entered a different, measurable latent regime, its final output remained "perfectly aligned."
These results suggest that current output-centric safety techniques may have a "blind spot in alignment," failing to detect potential risks or state changes occurring inside the model.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.