Home /note

Test-Time Scaling

Test-Time Scaling is a series of techniques that improve a model's answer by spending more compute at inference time, rather than by training a bigger model. Typically it refers to improving a model's answer by spending more reasoning/thinking tokens before responding.

It was popularised by OpenAI's o1 models, which were trained to reason before answering and let developers control how much reasoning effort to spend (openaiO1Model). See LLM Reasoning.

There are two broad approaches:

  • Sequential: the model thinks for longer, generating a longer chain of thought before it answers.
  • Parallel: the model generates several answers and one is picked, either by majority vote, as in self-consistency (Wang et al., 2022), or by a verifier that scores each candidate.

Parallel approaches only help if you can reliably pick the right answer. Heavy Thinking: A Test-Time Scaling Pattern for Hard Problems combines the two: subagents reason in parallel, then another LLM deliberates over their answers sequentially.

References

O1 Model. https://developers.openai.com/api/docs/models/o1. ↩

Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc Le, Ed Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou. Self-Consistency Improves Chain of Thought Reasoning in Language Models. 2022. doi:10.48550/ARXIV.2203.11171. ↩