LLMs Aren't Reasoning. They're Predicting - And Why That Matters More Than People Think.
Thore Graepel just published a piece in MIT Technology Review that deserves attention.
His argument is simple: Modern large language models look smart, but they aren't reasoning. They aren't evaluating evidence. They aren't maintaining a stable internal state. They aren't revising beliefs. They're predicting the next token.

This isn't a criticism of the technology; it's a reminder of what the technology actually is.
Graepel compares LLMs to AlphaGo, which combined intuition with a real reasoning system. AlphaGo explored possible futures, tracked uncertainty, and updated its beliefs as it evaluated positions. It had a clear separation between knowledge and reasoning. It had an epistemic state that could be inspected and audited.
LLMs don't have that. Even when they produce chain of thought, they're still using the same pattern mechanism. The reasoning isn't separate from the output. It's just more tokens.
This matters because people are starting to treat LLMs as if they're capable of scientific or engineering reasoning. They're not. They can help with discovery, but they can't replace the structured thinking that science and engineering require. They don't necessarily maintain a reliable record of how they arrived at an answer. They can't show their work in a way that can be trusted.
Graepel argues that trustworthy AI for high stakes domains will need architectures that can actually think. That means systems that maintain an epistemic state, revise beliefs based on evidence, and show their reasoning steps in a way that can be audited. Scaling intuition isn't enough. We need systems that reduce uncertainty in a structured way.
This is where the conversation intersects with critical thinking. People often confuse fluency with thought. They confuse confidence with correctness. They confuse pattern recognition with understanding. LLMs amplify these mistakes because they produce fluent answers that look like reasoning. They make it easy to skip the hard work of thinking.
In my book Critical Thinking, I talk about how people fall into traps when they rely on surface cues instead of deeper analysis. LLMs are the ultimate surface cue. They sound right. They sound confident. They sound thoughtful. But they aren't reasoning. They're predicting.
If we want AI systems that support science, medicine, engineering, and public sector decision making, we need architectures that can handle uncertainty, evidence, and revision. We need systems that can explain how they arrived at an answer. We need systems that can be audited.
This isn't a limitation. It's an opportunity. We're at the beginning of a shift from pattern machines to reasoning machines. The next generation of AI will need to combine intuition with structured thinking. It will need to separate knowledge from reasoning. It will need to maintain an epistemic state that can be inspected.
Until then, we should treat LLMs as powerful tools for language, synthesis, and exploration. We shouldn't treat them as reasoning systems. And we shouldn't let fluency trick us into thinking they're doing something they're not.
Related Thoughts
Perspectives sharing related architectures, models, and domain context.
The Illusion of Autonomy: Why AI Breakthroughs Still Require Human Oversight
A fascinating debate recently broke out on LinkedIn that cuts right to the heart of how we evaluate technological...
AI Can Now P‑Hack at Scale. The Real Risk Isn't Cheating. It's Obedience.
For years, p‑hacking was a slow, human problem. A researcher could sift through enough models and enough specifications...
Building GeoAI Systems That People Can Trust
The latest edition of the GeoAI and the Law Newsletter lays out a clear message for anyone working at the intersection...