How Reasoning Models Work: Chain of Thought Explained

Reasoning models are language models that think before they answer. Instead of predicting the next token immediately from the input, they generate an extended internal monologue — working through a problem step by step, checking their own logic, backtracking when something does not hold, and only then producing a final response. This is what “chain of thought” means in practice: not a prompt technique, but an architectural behaviour baked into the model. Understanding how this works explains why reasoning models are dramatically better at hard problems and also why they are slower and more expensive than standard models.

The Limitation of Standard Language Models

A standard language model — GPT-4o, Claude 3.5 Sonnet, Gemini 1.5 Pro — generates each token in a single forward pass through the network. When it answers a maths problem, it is effectively retrieving a pattern similar to ones it saw during training, not performing the actual computation. This works well for problems well-represented in training data, but fails on novel multi-step reasoning because the model has no mechanism to check whether the intermediate steps are actually correct. It commits to each token as it produces it, with no way to go back. The result is that standard models get easy reasoning problems right (because similar problems appeared in training) and systematically fail on hard ones (because reasoning chains of five or more steps require each step to be correct, and errors compound).

Chain of Thought: The Core Idea

The original chain-of-thought research (Wei et al., 2022) showed that prompting a model to write out its reasoning steps before answering significantly improved accuracy on arithmetic, common-sense, and symbolic reasoning tasks. The key insight was that generating intermediate steps gave the model more “tokens to think with” — each intermediate token the model generates becomes part of the context for predicting the next token, effectively allowing the model to condition later steps on earlier computed values rather than trying to do everything in one forward pass. Writing “let me work through this step by step” is not just style — it genuinely changes what computation the model performs.

From Prompting to Training: How Reasoning Models Are Built

Early chain-of-thought was a prompting technique applied at inference time to existing models. Reasoning models like o1 and DeepSeek R1 take this further by training the model to reason as a default behaviour, not as a prompted behaviour. The key advancement is reinforcement learning from verifiable rewards: the model is trained on problems with checkable correct answers (mathematics problems, coding problems where solutions can be run and tested), using RL to reward correct final answers regardless of the path taken. Through this training, the model learns that extended reasoning before committing to an answer leads to more rewards. It discovers reasoning strategies — decomposing problems, checking sub-answers, using analogies — because these strategies consistently improve its reward signal. The crucial insight from DeepSeek’s R1 paper is that this reasoning behaviour emerged from the RL process without requiring manually labelled chain-of-thought examples as training data.

Figure 1 — Standard LLM vs reasoning model: inference time compute

Standard LLM Input prompt Single pass Answer (immediate) Reasoning model Input prompt Thinking tokens (extended) Answer (verified)

What Happens During the Thinking Phase

The internal reasoning of a model like DeepSeek R1 (which makes its thinking visible) shows a recognisable pattern on hard problems. The model typically starts by restating the problem in its own words to check its understanding. It then proposes an initial approach, works through it step by step, and often pauses to verify intermediate results — “let me check this: if X = 3 then Y = 6, which means…” When it reaches a contradiction or uncertain step, it backtracks and tries a different approach, sometimes multiple times. For very hard mathematics problems, R1’s thinking traces can run to thousands of tokens, with multiple failed attempts before a successful solution. This self-correction behaviour — the willingness to abandon a reasoning path and start over — is what distinguishes reasoning models from standard models and explains the performance gap on problems that require genuine multi-step inference.

Test-Time Compute Scaling

The insight underlying reasoning models is that scaling compute at inference time (spending more tokens thinking) can improve output quality in ways analogous to scaling compute at training time (using more parameters or more data). This is called test-time compute scaling. For a given problem, a reasoning model with a larger thinking budget will generally produce a more accurate answer than one with a smaller budget — up to a point. This property enables the configurable “reasoning effort” settings in models like OpenAI o3/o4-mini: you can dial the compute spend up for hard problems and down for easy ones, trading cost for quality dynamically. The scaling law for test-time compute is still being characterised, but the general principle — more thinking tokens means better reasoning on hard problems — is well established. A key finding from OpenAI’s scaling work is that test-time compute scaling shows similar returns to training-time compute scaling: doubling the thinking budget improves performance comparably to doubling the model size in certain regimes, which has significant implications for how AI development should be structured going forward. The ability to trade cost for quality dynamically at inference time is a fundamentally different capability than anything previous model generations offered.

Why Reasoning Models Underperform on Simple Tasks

Reasoning models are not always better than standard models. On simple, well-defined questions — “what is the capital of France?”, “summarise this paragraph”, “write a cover letter” — standard models are faster, cheaper, and often produce equally good or better output. The extended thinking process adds latency and cost without improving quality when the task does not require multi-step inference. For creative writing, extended thinking can actually hurt output quality by making the model overly analytical about a task that benefits from fluency. The practical takeaway: use reasoning models when the task genuinely requires multi-step logic, verification, or precision (mathematics, coding, planning, complex analysis), and use standard models for tasks where fast, fluent generation is the primary value.

The Self-Verification Loop

One of the most valuable properties of reasoning models is self-verification — the ability to check their own answers during the thinking process. On a mathematics problem, the model can derive an answer, then substitute it back into the original equation to verify correctness. On a coding problem, it can mentally trace through the code execution path to check for edge cases. This is qualitatively different from asking a standard model to “check your answer” — the reasoning model does this spontaneously as part of its trained behaviour, not because it was prompted to. The self-verification loop is why reasoning models make fewer systematic errors on problems with checkable answers: incorrect intermediate steps are more likely to be caught and corrected before the final answer is committed to.

Reasoning models represent a genuine change in how language models handle hard problems — not just more parameters or more training data, but a different computation pattern at inference time. The thinking process is not a trick or a prompt technique; it is a trained behaviour that emerges from reinforcement learning on verifiable outcomes. Understanding this mechanism explains both the capabilities (multi-step reasoning, self-correction, better accuracy on hard problems) and the limitations (slower, more expensive, no benefit on simple tasks) of reasoning models, and helps you choose when to use them and when a standard model is the better tool.

Figure 1 — Standard LLM vs reasoning model: how inference differs

Standard LLM Input Single pass Answer (immediate) Reasoning model Input Thinking tokens (extended) Answer (verified)

Reasoning models represent a genuine change in how language models handle hard problems — not just more parameters or more training data, but a different computation pattern at inference time. The thinking process is not a trick or a prompt technique; it is a trained behaviour that emerges from reinforcement learning on verifiable outcomes. Understanding this mechanism explains both the capabilities (multi-step reasoning, self-correction, better accuracy on hard problems) and the limitations (slower, more expensive, no benefit on simple tasks) of reasoning models, and helps you choose when to use them and when a standard model is the better tool.

The practical consequence for anyone building AI-powered products: match the tool to the task. Hard reasoning problems with verifiable answers — reasoning models. Fast, fluent, high-volume generation — standard models. That simple routing rule captures most of the quality-cost optimisation available in modern AI deployment.

Leave a Comment