AI Briefing
KO

What Do Your Logits Know? (The Answer May Surprise You!)

·2026.04.20 09:00

Key point

Even the final **logits** alone can leak unnecessary information about an image query.

Details

Using a Vision-language model as a testbed, the study compared what information remains as internal representations are compressed. The benchmarks were the information-rich residual stream, low-dimensional projections obtained via tuned lens, and the final top-k logits that have the greatest influence on the model's answer.

The key finding is that even the easily accessible top logit values alone can leak task-irrelevant information contained within image-based queries. In some cases, these logits revealed as much information as direct projections of the entire residual stream.

In other words, just looking at the model's final generated output can expose internal information that users originally assumed was inaccessible. This shows that probing a model's internals can go beyond simple interpretation and lead to information leakage risks.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.