AI Briefing
KO

Gemma 3 NLA released

·2026.05.08 10:44

Key point

Anthropic released NLA interpretability research for Gemma 3.

Details

Anthropic's interpretability research Natural Language Autoencoders (NLA) has been released. It converts LLM token-level activations into natural language explanations, making the internal signals at the moment of token generation human-readable.

The setup has two components.

  • Auto Verbalizer (AV): converts residual-stream activations into natural language explanations.
  • Activation Reconstructor (AR): reconstructs the explanations back into vectors to check consistency.

On Hugging Face, AV/AR checkpoints fine-tuned from google/gemma-3-27b-it have been uploaded, with the extraction point at block 41 residual stream output. These models are not for general conversation but exclusively for activation decoding. On Neuronpedia, you can enter a question and click on tokens to view interpretations via explain; as an example, a screenshot was shared showing that for the input "I am Elon musk," the early tokens were classified as fabricated and satirical.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.