Don't Be Fooled—LLMs Don't Reason

Don't Be Fooled—LLMs Don't Reason


Ten years after AlphaGo's match against Go champion Lee Sedol, today's AI still isn't tapping into the machinery that made that win possible.


By Thore Graepel | October 2, 2026




A decade has passed since AlphaGo defeated Go champion Lee Sedol in a five-game match that many considered a watershed moment for artificial intelligence. That victory hinged on a sophisticated architecture that combined deep neural networks with Monte Carlo tree search—a system capable of lookahead, systematic exploration, and strategic planning. It was, in a meaningful sense, a machine that could reason about a problem.


Today, large language models (LLMs) dominate the AI landscape. They can write code, summarize research, pass professional exams, and hold eerily fluent conversations. Yet despite their impressive surface capabilities, they are not doing what AlphaGo did. They are not reasoning.


The Fundamental Difference


LLMs are, at their core, next-token predictors. Given a sequence of text, they compute probability distributions over the next token and sample from them. This process, scaled to hundreds of billions of parameters and trained on vast swaths of internet text, produces outputs that appear thoughtful. But appearance is not mechanism.


When an LLM answers a math problem or drafts a legal brief, it is not performing deliberate, step-by-step inference. It is pattern-matching against statistical regularities learned during training. It has no internal model of the world, no capacity to simulate counterfactuals, and no mechanism for systematic search over possible solutions. It produces plausible continuations, not reasoned conclusions.


AlphaGo, by contrast, explicitly searched through a tree of possible future board states. It evaluated positions using a value network, proposed moves using a policy network, and used Monte Carlo rollouts to estimate outcomes. That is reasoning: the deliberate manipulation of internal representations to explore consequences and guide decisions.


Why the Confusion Persists


Part of the problem is that LLMs are extraordinarily good at imitating the outputs of reasoning. They have read countless explanations, proofs, and arguments. When prompted, they can generate text that resembles a chain of logical steps. But this is mimicry, not cognition.


Recent "chain-of-thought" prompting techniques encourage models to produce intermediate steps before a final answer. This can improve performance on some benchmarks, but it does not mean the model is reasoning. It is still predicting tokens—just more of them. The intermediate steps are not grounded in an internal world model; they are plausible-sounding text that often happens to lead to correct answers because similar patterns appeared in training data.


Benchmarks like GSM8K or MATH measure accuracy on specific problem sets, but they do not distinguish between genuine reasoning and sophisticated pattern recognition. A model can score highly by memorizing templates or exploiting statistical shortcuts, without ever engaging in deliberate inference.


The Missing Machinery


What would it take for an LLM to truly reason? It would need at least three things:


  1. An internal world model—a structured representation of entities, relations, and dynamics that can be queried and updated.
  2. A search mechanism—the ability to explore alternative trajectories, evaluate them, and select actions based on predicted outcomes.
  3. A grounding in causality—an understanding of interventions and counterfactuals, not just correlations.

  4. Current LLMs have none of these in a robust form. They are trained on observational data, not on interactive experience. They do not build explicit models; they compress statistical patterns into weights. They cannot reliably simulate the consequences of actions they have never seen.


    Some researchers are attempting to bolt reasoning capabilities onto LLMs—by coupling them with external tools, symbolic solvers, or reinforcement learning environments. These hybrid approaches are promising, but they are not the same as an LLM reasoning on its own. The LLM remains a component, not the reasoner.


    The Stakes


    Why does this matter? Because if we believe LLMs can reason, we will trust them in domains where they are fundamentally unreliable. We will deploy them in high-stakes settings—medicine, law, finance, autonomous systems—where their lack of genuine understanding can cause real harm. We will mistake fluency for competence and confidence for correctness.


    We will also misdirect research. If we think scaling up LLMs will eventually produce reasoning, we may pour resources into the wrong architecture. The history of AI is littered with approaches that seemed promising until their limitations became undeniable. We should not repeat that mistake.


    A Way Forward


    None of this means LLMs are useless. They are powerful tools for generation, summarization, translation, and retrieval. They can augment human productivity in countless ways. But they are not reasoners, and we should not pretend otherwise.


    If we want machines that reason, we should look to the architectures that actually do it: the AlphaGo lineage of systems that combine learning with search, perception with planning, and pattern recognition with explicit models. We should invest in hybrid systems that use LLMs as interfaces or memory stores while relying on other mechanisms for inference.


    Ten years after AlphaGo, we have faster hardware, larger datasets, and more sophisticated training techniques. But we have not solved the core problem of machine reasoning. We have, in some ways, retreated from it—trading explicit search and world models for scale and statistical mimicry.


    It is time to be honest about what LLMs can and cannot do. They do not reason. And until we build systems that do, we should not claim otherwise.

    via MIT Tech Review AI

Related