AI Briefing
KO

Truth Is Not a Direction: A Tarskian Attack on LLM Probes

·2026.07.27 21:56

Key point

Mathematical logic proves that no specific direction in an LLM's embedding space can represent Truth.

Details

In the embedding space of LLMs, there exists a 'Linear Representation Hypothesis' claiming that specific concepts (gender, emotion, etc.) exist as specific Directions. Recent studies argue that 'Truth' can also be captured as a specific direction within the embedding space, and attempt to use this to detect whether a model is being deceptive.

However, this article points out that such attempts are fundamentally impossible, using Tarski's undefinability theorem and the Diagonal argument.

  • The Paradox of Self-Reference: LLMs process input natural language by converting it into vectors, and the model's input itself can describe the model's internal state or properties.
  • Attack Mechanism: If a truth probe $t(s)$ exists, then inputting the sentence "This sentence's truth probe score is FALSE" produces a logical contradiction.
  • Conclusion: If a language is sufficiently expressive, there cannot exist a probe within that language that perfectly determines truth values. In other words, the geometric structure of the embedding space alone cannot fully capture a model's truthfulness.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.