Ten Years After Move 37: an AlphaGo Veteran Says LLMs Still Don't Reason
An AlphaGo veteran argues that chain-of-thought LLMs still only run System 1 pattern completion—and that trustworthy AI needs AlphaGo-style search over an explicit, auditable epistemic state.
Thore Graepel, a core member of the AlphaGo team at DeepMind and now chair of machine learning at University College London, draws a direct line from Move 37 in the 2016 match against Lee Sedol to what today's LLMs are missing. Move 37—a fifth-line stone that expert commentators read as a glitch—was not pure intuition, he writes. AlphaGo's policy network rated it roughly one in 10,000 as a human expert move. What selected it was the search machinery: an explicit game tree with thousands of branches, each representing a possible future, annotated with judgments from the neural networks and updated as reasoning progressed.
Graepel contrasts this with chain-of-thought prompting. The gains are real, especially in mathematics and coding, but the intermediate steps are produced by the same next-token prediction process, iterated longer. He identifies three structural failures. First, LLMs maintain no explicit, persistent, inspectable epistemic state—no ledger of hypotheses, confidence levels, evidence, and open questions that gets revised as new information arrives. Second, knowledge and reasoning are interwoven in the network weights, with no independent, explicitly represented set of beliefs. Third, research shows models often concoct chains of thought after the fact, reaching an answer by one route but reporting another.
His proposed alternative borrows AlphaGo's architecture. A general reasoning system should maintain an epistemic state representing what it holds as settled, what it doubts, what it has ruled out, and which questions remain open. Reasoning becomes a sequence of moves that change that state: deducing consequences, decomposing problems, and deciding what question to ask or experiment to run next. LLMs can contribute by suggesting approaches and interacting with tools via APIs or code, but an independent component must evaluate each move by how much it actually resolves uncertainty, updating beliefs only when backed by evidence.
Graepel left Google DeepMind to pursue this direction. He frames the stakes around high-stakes applications—medicine, engineering, scientific research—where it matters not only what a system concludes but how it arrived there, and where mistakes must be traceable to faulty reasoning, invalid evidence, or incorrect assumptions.