← Back to blog
AI Research·September 27, 2026·8 min read

Test-Time Compute Scaling: Why Letting a Model "Think Longer" Sometimes Beats Making It Bigger

A 2024 DeepMind study found that spending more compute at inference time — not training time — can let a small model match one many times its size on math problems. Here's the mechanism, and why it only works in a narrow band of difficulty.

Between 2020 and 2023, almost every capability gain in language models traced back to the same lever: more parameters, more training tokens, more pretraining compute. Then OpenAI shipped o1 in September 2024, and the story changed. o1 wasn't a bigger model — it was roughly the same size class as existing GPT-4-tier models, but it spent far more computation after the prompt arrived, generating long internal chains of reasoning before committing to an answer. On math and coding benchmarks, it beat larger, non-reasoning models by a wide margin.

The finding underneath this shift, made explicit in DeepMind's "Scaling LLM Test-Time Compute Optimally" (Snell, Lee, Xu, and Kumar, 2024), is worth stating plainly: for a meaningful band of problem difficulty, a smaller model given more inference-time compute — extra reasoning tokens, multiple attempts, a search process — can match the accuracy of a model many times its size trained the conventional way. The paper's headline result put a small model with well-allocated test-time compute in the same performance range as a model on the order of 14 times larger, on curated math benchmarks. That number is specific to their setup and shouldn't be treated as a universal exchange rate, but the qualitative point held up across follow-up work: extra thinking time is a real substitute for extra parameters, within limits.

Those limits are the interesting part. This isn't "just let the model think longer and everything gets better." It's a compute-allocation problem with a fairly narrow sweet spot, and understanding where that sweet spot is — and where it collapses — is the difference between using reasoning models well and burning tokens for nothing.

How test-time compute is actually spent

"Thinking longer" is not one technique. In practice it decomposes into a handful of distinct strategies, each with a different cost profile and a different failure mode:

StrategyHow it worksRelative costWhere it helpsWhere it breaks down
Self-consistency / majority voteSample N independent chains of thought, take the most common final answerLinear in NProblems with a small set of discrete possible answersTies don't resolve; useless when the model is consistently wrong
Best-of-N with a verifierSample N completions, score each with a separate reward or verifier model, keep the top-scoring oneLinear in N, plus verifier costAny domain with a cheap correctness check (unit tests, symbolic math checkers)Only as good as the verifier; a weak verifier picks confidently wrong answers
Process-reward-guided search (beam search / tree search over reasoning steps)Score intermediate reasoning steps, not just the final answer, and prune low-scoring branches earlyHigher — needs a step-level reward modelMulti-step problems where early mistakes compoundProcess reward models are expensive to train and easy to overfit to superficial reasoning patterns
Sequential revisionModel critiques and rewrites its own answer iteratively in a single chainRoughly linear in revision roundsCases where the first draft has a locatable, fixable errorModels often "fix" a correct answer into a wrong one, or loop without converging
Native long chains-of-thought (o1/o3, DeepSeek-R1 style)The model itself was trained via reinforcement learning to generate long internal reasoning before answering — no external search scaffoldBaked into inference; cost scales with reasoning-token length, not with re-samplingGeneral reasoning tasks, without needing a separate verifier at inference timeReasoning length isn't guaranteed to correlate with correctness; longer isn't always righter

The first four are things you can bolt onto almost any base model at inference time. The fifth — what o1, o3, and DeepSeek-R1 (released January 2025) actually do — is different in kind: the long reasoning trace is a learned behavior, produced by reinforcement learning against verifiable rewards (mostly math and code, where "correct" is checkable automatically), not a search scaffold wrapped around a static model. That's the technical shift that made "reasoning models" a distinct product category rather than a clever prompting trick.

Why the gains don't scale forever

The DeepMind paper's actual contribution wasn't "more compute helps" — that's not surprising. It was showing that the optimal way to spend a fixed test-time compute budget depends heavily on how hard the question is relative to the model's baseline competence, and that naive strategies (like always sampling a fixed N) waste most of that budget.

Three regimes show up consistently across this line of research:

This is the part that gets flattened in "just add reasoning" marketing framing: test-time compute is not a dial that trades cleanly against model size across the whole difficulty spectrum. It's most valuable in the middle band, and that band shifts depending on the base model's underlying capability — a stronger base model pushes more problems into "easy" (where extra compute is wasted) and shrinks the number of problems that are hopeless regardless of compute.

What changed with RL-trained reasoning models

Before o1, most test-time compute research used a fixed base model plus an external scaffold — sample-and-vote, or search guided by a separately trained verifier. That approach has an obvious ceiling: the scaffold can only select among answers the base model was already capable of producing.

o1 and DeepSeek-R1 took a different approach: train the model itself, via reinforcement learning, to produce longer and more self-correcting reasoning traces before answering, using outcome-verifiable domains (math with checkable answers, code with runnable tests) as the reward signal. DeepSeek's R1 paper in particular showed this could be done with large-scale RL and comparatively little hand-labeled reasoning data, which mattered because process-reward-model training had previously been seen as a data bottleneck.

The practical effect: reasoning ability that used to require an external search scaffold at inference time got partially absorbed into the model's own weights and decoding behavior. You don't need to run best-of-16 with a separate verifier to get a chunk of the benefit — the model does something functionally similar inside a single generation. That's a real efficiency win, but it comes with the same underlying tradeoff pushed into a different place: a reasoning model spends more tokens (and therefore more latency and cost) per query than a non-reasoning model of similar size, whether or not the extra reasoning was actually necessary for that particular question.

What this does and doesn't change if you're building on top of these models

For anyone choosing between a bigger general model, a reasoning model, or a sample-and-verify pipeline of your own, a few things follow directly from the mechanism rather than from vendor marketing:

The takeaway

Test-time compute scaling is real and it's not a trick — a 2024 result showing a much smaller model matching a far larger one on math benchmarks, purely by spending inference compute more intelligently, held up and helped motivate an entire generation of reasoning-model products. But it's a substitution effect with a narrow effective range, not a free multiplier: it does the most good on problems the model can almost solve reliably, does nothing on problems it already solves or problems it fundamentally can't, and it always costs latency and tokens in exchange for accuracy. The engineering question worth asking before reaching for a reasoning model or a resample-and-verify pipeline isn't "will this help" — it almost always will, some — but "is this query actually in the band where extra thinking pays for itself."

#test-time-compute#reasoning-models#ai-research#chain-of-thought#llm-scaling#inference-scaling

Related reading

AI Research
Your Model's Reasoning Trace Is Not a Log: What Chain-of-Thought Faithfulness Research Actually Shows