Chain of Thought vs Zero-Shot Reasoning

Chain-of-thought and zero-shot prompting sit at opposite ends of the reasoning depth spectrum. Zero-shot asks a model for a direct answer with no examples or reasoning scaffolding. Chain-of-thought prompts the model to work through its reasoning before producing a final answer. For standard language models, the choice between them has a measurable impact on accuracy for reasoning-intensive tasks. For reasoning models, the distinction blurs — the model reasons by default. Understanding when each approach applies, and what the research says about why, makes you a more effective prompt engineer and helps you avoid spending tokens on reasoning scaffolding when a direct question works just as well.

Zero-Shot Prompting

Zero-shot prompting means presenting a model with a task and asking it to solve it directly, with no examples and no explicit instruction to reason. “What is 127 × 43?” or “Summarise this article in three sentences” are zero-shot prompts. The model draws on patterns from training to generate an answer in a single forward-pass style. For tasks where the answer pattern is well-represented in training data — translation, summarisation, factual recall, standard code generation — zero-shot works very well. The model has seen millions of examples of these tasks and can pattern-match effectively. For tasks requiring multi-step inference that was not directly present in training data — novel mathematical problems, logical reasoning with novel premises — zero-shot performance drops because the model has no retrieval pattern to match and no mechanism to verify whether its immediate answer is correct.

Chain-of-Thought Prompting

Chain-of-thought prompting, introduced by Wei et al. in 2022, showed that prompting a model to “think step by step” dramatically improved accuracy on arithmetic, common-sense, and symbolic reasoning tasks. The key finding was that generating intermediate steps — even without providing example steps — unlocked reasoning capabilities that zero-shot elicited poorly. The two main forms are few-shot chain-of-thought (providing examples of problems with worked-out reasoning steps), and zero-shot chain-of-thought (simply appending “Let’s think step by step” to the prompt). Both work significantly better than zero-shot on hard reasoning tasks, and the performance gap grows with problem difficulty.

Why Chain-of-Thought Works: The Mechanism

The explanation for why chain-of-thought improves performance connects to how autoregressive models work. When a model generates a token, it conditions on all previous tokens. By generating intermediate reasoning steps, the model effectively creates a scratchpad in the context window — intermediate results are written down and become part of the conditioning context for subsequent steps. A model that writes “127 × 43 = 127 × 40 + 127 × 3 = 5080 + 381 = 5461” can condition the final answer on each verified intermediate step. Without this, the model must compute the entire answer in a single forward pass through the network, which is fundamentally more difficult for multi-step problems because the network architecture is not designed for long-range numerical computation in a single pass.

Zero-Shot Chain-of-Thought: Just Add “Think Step by Step”

Kojima et al. (2022) showed that simply appending “Let’s think step by step” to a prompt — with no examples — substantially improved performance on reasoning benchmarks. This zero-shot CoT approach works because it activates reasoning behaviour that the model has learned from training on text that includes worked examples. The instruction is a signal to the model to produce the kind of step-by-step output pattern it has seen in training rather than jumping to a conclusion. More specific instructions often work even better: “Let’s work through this carefully, checking each step” or “Think through this systematically, identifying any assumptions” direct the model toward more structured reasoning than the generic step-by-step instruction.

Figure 1 — Zero-shot vs chain-of-thought: when each approach fits

Task Zero-shot Chain-of-thought Translation / summarisation✓ PreferredWasteful Factual Q&A✓ PreferredNo benefit Multi-step mathsOften wrong✓ Large improvement Logical reasoningOften wrong✓ Large improvement Code generation (standard)✓ FineMarginal benefit Complex debuggingInconsistent✓ Helpful

Few-Shot Chain-of-Thought

Few-shot CoT provides worked examples before the actual question — typically 2–8 examples of problems with full reasoning traces. This is more powerful than zero-shot CoT because the examples demonstrate the specific reasoning format, level of detail, and verification steps the model should use. For a custom domain where the reasoning patterns are unusual — a specific mathematical notation, a proprietary analytical framework, a domain-specific verification procedure — few-shot CoT lets you teach the model the exact reasoning style you want. The cost is the token overhead of the examples, which can be significant for long reasoning traces. For most standard tasks, zero-shot CoT (“let’s think step by step”) is efficient enough that few-shot CoT’s additional token cost is not justified.

Chain-of-Thought with Reasoning Models

Reasoning models like o3, o4-mini, and DeepSeek R1 use chain-of-thought internally as a trained behaviour rather than a prompted behaviour. Appending “think step by step” to a prompt sent to a reasoning model is typically unnecessary — the model already does this. More importantly, it is possible to interfere with the model’s internal reasoning by over-constraining it with explicit step-by-step instructions. The best practice for reasoning models is to describe the problem clearly and let the model decide how to reason through it, rather than dictating a specific reasoning format. Where explicit reasoning structure helps with reasoning models is in defining the output format of the final answer — “show your work, then give a final answer in a box” — which separates presentation from reasoning without constraining the reasoning process itself.

Self-Consistency: Taking the Best of Multiple Chains

Wang et al. (2022) introduced self-consistency as an extension of chain-of-thought: generate multiple independent reasoning chains for the same problem and take a majority vote on the final answers. If you generate 10 chains and 7 reach the same answer, that answer is more likely correct than any single chain’s answer. Self-consistency typically improves accuracy by 5–15 percentage points over single chain-of-thought on hard benchmarks. The cost is generating 10× the tokens, which makes it expensive for production use but practical for high-stakes tasks where accuracy justifies the cost. Modern reasoning models implicitly implement something similar — they explore multiple approaches in a single extended thinking trace rather than generating separate chains.

When to Use Each in Practice

The practical decision rule is straightforward. For any task where you would classify the problem as “hard reasoning” — mathematics, logic, complex debugging, multi-step planning — use chain-of-thought prompting (or route to a reasoning model that does it automatically). For tasks where pattern recall and fluency are the primary value — writing, summarisation, translation, factual Q&A — use zero-shot prompting to save tokens and reduce latency. For tasks in the middle — moderate difficulty reasoning, standard coding — try zero-shot first and only add CoT if you observe consistent errors that chain-of-thought would likely catch. Adding “let’s think step by step” costs you perhaps 200–500 tokens in reasoning overhead; on a simple question, that is wasteful; on a hard reasoning question, it can be the difference between a wrong answer and a correct one.

The Token Cost of Chain-of-Thought

The practical cost of chain-of-thought needs to be kept in perspective. A zero-shot answer to a mathematics problem might be 50 tokens. A chain-of-thought answer to the same problem might be 300–600 tokens. At GPT-4o pricing (~$2.50/M output tokens), that difference is $0.00063 per query — essentially nothing at low volume, but meaningful at millions of queries per day. At reasoning model pricing (o4-mini at ~$1.10/M, o3 at ~$15/M), the thinking tokens add up more significantly — a 3,000-token thinking trace on o3 costs $0.045 per query, which at 100,000 queries per day is $4,500 per day. Cost discipline on thinking budget matters at scale, which is why configurable reasoning effort settings are practically important rather than just a product feature.

Zero-shot is the right default for most tasks — it is fast, cheap, and accurate enough for the large majority of queries that do not require multi-step inference. Chain-of-thought is a targeted intervention for tasks where the model’s zero-shot accuracy is inadequate, specifically because the task requires logical chaining that benefits from intermediate steps being committed to the context window. For reasoning models, the prompting question shifts from “should I add chain-of-thought?” to “how much thinking compute should I allocate?” — which is essentially the same question expressed at a different level of abstraction.

The research arc from zero-shot to chain-of-thought to reasoning models is a single story about unlocking latent reasoning capability through inference-time computation — each step in the progression gives the model more tokens to think with, and more tokens consistently translates to better accuracy on the tasks that matter most for AI utility.

Leave a Comment