AivexaNewsSearch
AI news for builders and product teamsChecked every hour

Don’t be fooled—LLMs don’t reason

Collected Oct 2, 2026

Thore Graepel, chair of machine learning at University College London and a former core member of the AlphaGo team at Google DeepMind, writes in MIT Technology Review that large language models do not truly reason, and that today's AI lacks the reasoning powers that produced AlphaGo's celebrated Move 37 against Lee Sedol in Seoul in March 2016.

Graepel describes AlphaGo as two systems: a policy network trained to predict strong human moves, and search machinery that built and searched a game tree of thousands of branches representing possible futures. The intuitive network rated Move 37 as nothing special, with roughly a one in 10,000 chance of being played by an expert human. The search machinery chose it by weighing future consequences. Graepel compares this split to Daniel Kahneman's System 1 and System 2 modes of thought. AlphaGo won the five-game match 4-1, and Lee said afterwards that he had changed his mind and that AlphaGo was creative.

By contrast, Graepel writes, an LLM repeatedly picks the next token, which he characterizes as System 1 in action. Chain-of-thought generation, in which models produce intermediate steps before answering, has yielded real gains, above all in mathematics and coding, but he says it is not a genuinely separate reasoning mechanism because the intermediate reasoning is still produced by the same next-token prediction process.

He cites three shortcomings: models typically maintain no explicit, persistent, inspectable epistemic state; they lack a clean separation between what the system knows and how it manipulates that knowledge; and research has shown chains of thought are often produced after the fact, with the model reaching an answer by one route and reporting another. He says this matters in high-stakes fields such as medicine, engineering, and scientific research, where it matters not only what a system concludes but how.

Graepel says he recently left his position at Google DeepMind. He proposes drawing on AlphaGo's architecture, in which a game tree records considered variations annotated with neural network judgments and is updated as reasoning progresses. General reasoning systems, he argues, should maintain an epistemic state of what is settled, doubted, ruled out, or open, with reasoning as a sequence of moves that advances knowledge and reduces uncertainty. He notes open-world reasoning is harder than board games because the state of affairs is only partially known, available actions are large and variable, and consequences are stochastic or unknown. LLMs, he adds, can suggest approaches and interact with tools via APIs or code, and an independent part of the system should evaluate each move by how much it resolves uncertainty, updating beliefs only when backed by evidence. He concludes that trustworthy machine intelligence will not come from making System 1 bigger, and that insights in drug discovery, materials, climate, and diagnosis require systems whose conclusions arise from an auditable sequence of evidence, inference, and belief revision.

Read at MIT Technology Review · AI

Based on reporting from the original publisher. Visit the source for full context and later updates.

Publisher excerpt

On an afternoon in Seoul in March 2016, I watched a program I helped build put a stone on the fifth line of a Go board in what looked like a gift to its human opponent. Move 37 in game two of the five-game match looked so absurd that some commentators thought it was a…