AI Briefing
KOSign in

AlphaGo Co-Developer Argues LLMs Lack True Reasoning, Advocates for Epistemic State Architecture

·2026.10.02 17:00

Key point

Thore Graepel, a core member of the AlphaGo team, left Google DeepMind to pursue research into systems that maintain explicit, auditable epistemic states rather than relying on next-token prediction.

Details

Thore Graepel, chair of machine learning at University College London and former core member of the AlphaGo team, argues that current Large Language Models (LLMs) do not possess genuine reasoning capabilities. He contrasts LLMs with AlphaGo, which combined neural network intuition with explicit search machinery to evaluate future consequences. While AlphaGo’s policy network provided fast, intuitive hunches (System 1), its search mechanism performed slow, deliberative evaluation (System 2) by constructing and weighing a game tree of possible futures.

Limitations of Chain-of-Thought

Graepel contends that techniques like chain of thought in LLMs are insufficient because they rely on the same next-token prediction process, merely iterated for longer. He identifies three specific shortcomings that prevent LLMs from qualifying as true reasoners:

  • Lack of Persistent Epistemic State: LLMs do not maintain an explicit, inspectable ledger of hypotheses, confidence levels, and unresolved questions that can be systematically revised.
  • Interwoven Knowledge and Reasoning: In neural networks, knowledge and reasoning logic are inseparable within the weights, lacking an independent, explicitly represented set of beliefs.
  • Post-Hoc Rationalization: Research indicates that LLMs often generate reasoning chains after reaching an answer, reporting a path that does not reflect how the conclusion was actually derived.

The Case for Auditable Reasoning

The author emphasizes that in high-stakes fields like medicine and scientific research, the process of arriving at a conclusion is as critical as the conclusion itself. To achieve trustworthy AI, Graepel proposes a new architecture inspired by AlphaGo’s game tree. This system would maintain an epistemic state representing settled facts, doubts, and open questions. Reasoning would then be defined as a sequence of moves that update this state to reduce uncertainty, with an independent evaluator ensuring that belief changes are backed by evidence.

Graepel recently left his position at Google DeepMind to pursue this approach, arguing that scaling current LLMs only sharpens intuition without adding deliberative capability. He asserts that future systems must produce auditable sequences of evidence and inference to generate novel insights in complex, open-world domains.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.