Golden Gate Claude
Key point
Anthropic has published interpretability research that identifies internal features within Claude and demonstrates that adjusting them changes the model's behavior.
Details
Anthropic released a new research paper aimed at understanding the internal workings of Claude 3 Sonnet. The team confirmed that within the model's 'mind' exist millions of concepts, or features, that activate when processing text or images.
One of these, the 'Golden Gate Bridge' feature, emerges when a specific combination of neurons activates. The researchers demonstrated that by adjusting the activation strength of this feature, they could directly control Claude's behavior.
For example, amplifying this feature causes Claude to mention the Golden Gate Bridge or generate related content in its answers, even when it has no direct relevance to the question. This is not simple prompt manipulation or fine-tuning, but a method of precisely modifying the model's internal activation states.
This technique could in the future be used to control safety-related features—such as those tied to generating dangerous code or deceptive behavior—helping to make AI models safer.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.