Test-time compute scaling is the observation that spending more compute during inference — having a model think longer rather than answer immediately — improves output quality in ways that parallel the improvements from spending more compute during training. It is the key insight behind reasoning models: instead of making the model larger or training it on more data, you give it more tokens to think with before committing to an answer. Understanding test-time compute scaling explains why o3, DeepSeek R1, and similar models behave the way they do, why they get configurable “thinking budgets,” and where the limits of this approach lie.
The Two Ways to Scale AI Performance
Before reasoning models, the primary lever for improving AI performance was training-time compute: train a larger model on more data with more GPU-hours. The scaling laws described by Kaplan et al. and later refined by Hoffmann et al. (the Chinchilla paper) quantify this precisely — given a compute budget, there is an optimal model size and dataset size that maximises performance. Every major model generation from GPT-2 to GPT-4 followed this curve: more training compute, better models. The insight that prompted the development of reasoning models is that there is a second scaling axis — inference-time compute — that had been largely unexplored. Rather than spending all the compute budget at training time and then generating answers in a fixed number of forward passes at inference, what if you allowed the model to spend more compute at inference time by generating more tokens? The result, if the model uses those tokens productively, is better answers without retraining.
Why More Tokens Means Better Reasoning
Standard language models generate each token conditioned on all previous tokens. When a model generates a reasoning chain — “first, let me note that… then, this implies… therefore…” — each subsequent token is conditioned on the accumulation of the reasoning so far. This means the model can effectively perform multi-step computation by committing intermediate results to the token stream and conditioning later steps on them. A model with 100 thinking tokens available can express a two-step argument; a model with 1,000 thinking tokens available can express a ten-step argument; with 10,000 tokens it can explore multiple approaches, verify intermediate steps, and backtrack. The quality improvement comes not from the model being “smarter” in any fundamental sense — the underlying weights are identical — but from the model having more context to work with and more opportunities to self-correct before producing a final answer.
The Scaling Law for Test-Time Compute
OpenAI’s work on test-time compute scaling showed that the relationship between inference compute and performance follows a power law similar to the training-time scaling law. Doubling the thinking budget improves performance on hard benchmarks by a predictable amount. Importantly, the research found that for a fixed total compute budget, there is a crossover point where spending more compute at inference time (with a smaller base model) outperforms spending all compute at training time (with a larger base model). A smaller model given more thinking tokens can match or exceed the performance of a larger model given fewer tokens, on reasoning-intensive tasks. This has significant implications for AI deployment: it means reasoning capability can be bought cheaply at inference time rather than expensively at training time, and it changes the economics of AI development.
How Models Use Test-Time Compute
There are several distinct strategies a model can use to spend thinking tokens productively. Sequential reasoning chains: work through a problem step by step, where each step is conditioned on previous steps — this is the most common pattern in deployed reasoning models. Exploration and backtracking: try an approach, recognise it is not working, discard it, and try a different approach — observed consistently in QwQ and R1’s thinking traces on hard mathematics. Verification: derive an answer, then verify it independently — substitute an answer back into an equation, trace through code execution, check a proof step by step. Decomposition: break a hard problem into subproblems, solve each, and recombine — effective for complex multi-constraint problems. Self-critique: generate an answer, then critique it as if reviewing someone else’s work, then revise — a pattern that improves output quality on tasks with subjective quality dimensions. Each of these strategies requires tokens to implement, which is why more thinking tokens enables better reasoning.
Figure 1 — Test-time compute scaling: thinking tokens vs benchmark accuracy
Configurable Thinking Budgets in Practice
The practical manifestation of test-time compute scaling is the configurable thinking budget exposed in modern reasoning model APIs. OpenAI’s API supports a reasoning_effort parameter (low, medium, high) for o3 and o4-mini. DeepSeek’s API allows configuring whether extended thinking is enabled. Anthropic’s Claude 3.7 Sonnet accepts a thinking block with a configurable budget_tokens parameter. These controls let developers trade cost and latency for quality dynamically — using a low thinking budget for easy queries and a high one for hard ones. The optimal budget depends on the specific task and model: for competition mathematics, more thinking tokens consistently improve accuracy up to tens of thousands of tokens; for simple factual recall, the thinking budget makes almost no difference.
The Relationship to Training-Time Scaling
A key finding from the test-time compute scaling research is that the two scaling axes are somewhat substitutable. For a given performance target on a hard reasoning benchmark, you can reach it either by training a large model for a long time and running it with minimal inference compute, or by training a smaller model more quickly and running it with more inference compute. The OpenAI scaling work suggested that these paths have roughly equivalent total compute costs up to a point — the question is when you want to spend the compute (upfront at training time, or at query time distributed across inference calls). This substitutability is economically significant: it means that a well-funded lab with massive training compute does not have an insurmountable advantage over an organisation that can afford inference compute but not frontier training runs.
Limits of Test-Time Compute Scaling
Test-time compute scaling is not unlimited. There is a diminishing returns curve — doubling the thinking budget always helps but the marginal improvement shrinks as the budget grows. There is also a capability ceiling: no amount of thinking tokens can allow a model to recall information it was never trained on, or to solve a problem that exceeds the model’s fundamental capacity to represent the solution. Extended thinking also helps much more on some tasks than others. For problems with verifiable intermediate steps (mathematics, code), the model can use thinking tokens to check its work and correct errors — these tasks benefit enormously. For problems without checkable intermediate steps (creative writing, subjective judgments), extended thinking provides much smaller benefits because the model has no way to verify whether its intermediate reasoning is leading in the right direction.
Best-of-N Sampling as an Alternative
Before the extended chain-of-thought approach became standard, a simpler form of test-time compute scaling was best-of-N sampling: generate N independent answers to the same question and select the best one using a verifier or reward model. This approach also improves accuracy as N increases — it is essentially searching over multiple candidate answers rather than refining a single answer through longer reasoning. Best-of-N is less efficient than extended chain-of-thought for hard reasoning tasks (it does not benefit from self-correction within a single answer) but is simpler to implement and still widely used for code generation (generate 5 implementations, run tests, return the first that passes).
Implications for AI Development
The practical implications of test-time compute scaling are significant for how teams build AI applications. First, model selection is not just about capability at a fixed inference cost — it is about capability across different inference budgets. A smaller model with more thinking tokens may outperform a larger model with fewer. Second, the appropriate inference budget varies by task, and hardcoding it to a fixed value is suboptimal — dynamic routing based on estimated query difficulty is more cost-effective. Third, the economics of test-time compute favour use cases where correctness on hard problems justifies high per-query cost, rather than high-volume low-difficulty queries where training-time scaling and standard models are more efficient. The broader implication is that reasoning capability has become a resource you can purchase at inference time, which changes the build-vs-buy and model-selection calculus for anyone deploying AI at scale.
Test-time compute scaling is the theoretical foundation that explains why reasoning models work. More thinking tokens do not make a model smarter in the sense of knowing more — they give it more computational steps to work with, which translates directly to better performance on tasks that require logical chaining, self-verification, and structured exploration. The configurable thinking budgets in current model APIs are a direct product of this understanding: they let you dial the inference compute up or down per query, paying only for the reasoning depth each task actually requires.