Skip to content
Preprint

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility

Aug 2026 · 5 citations · ⚡ 1 influential
Computer Science

TL;DR

This empirical study covers broad knowledge, symbolic reasoning, and competition mathematics, and it introduces an evaluation profile whose coordinates and simple functionals recover or bound common repeated-sampling metrics, and require compute accounting and uncertainty estimates that match the protocol.

Abstract

Large language models can solve harder reasoning problems with more inference-time compute. The term"test-time scaling,"however, covers several inference algorithms: extending deliberation along one trajectory, sampling completed candidates and aggregating them by voting or verification, and searching over partial states. These algorithms differ in statistical structure, compute requirements, and failure modes. Treating them as interchangeable under a scalar"budget,"or reporting accuracy without specifying the inference protocol, makes results difficult to compare across studies. We study test-time scaling along three axes. First, we formalize it as budgeted inference over the implicit prefix tree of an autoregressive model and distinguish single-trajectory sequential scaling, leaf-level scaling with terminal reduction, and prefix-level scaling. Second, we treat the full inference system as the evaluated object and separate end-to-end performance from candidate-bank diagnostics. We introduce an evaluation profile whose coordinates and simple functionals recover or bound common repeated-sampling metrics, and require compute accounting and uncertainty estimates that match the protocol. Third, we distinguish exact replay from distributional reproducibility and state the requirements for each. We also organize open-weight reasoning models by model-side and interface mechanisms. Our empirical study covers broad knowledge, symbolic reasoning, and competition mathematics, and we publicly release 1,403,520 sampled model attempts. The project website is available at https://mohsenhariri.github.io/scorio/tts. The released datasets are Trace (https://huggingface.co/datasets/harimo/scorio-trace), Lite (https://huggingface.co/datasets/harimo/scorio-lite), Math (https://huggingface.co/buckets/harimo/scorio-math), and SuperGPQA (https://huggingface.co/buckets/harimo/scorio-gpqa).

View source

Similar papers

When LLM Meets Tree Search: A Systematic View of Inference as Search in Large Language Models

This survey systematizes recent progress in tree-search-based reasoning, viewing inference as instance-specific optimization rather than decoding, and introduces a Unified Design Space spanning search topology, evaluation signals, and control dynamics to unify a fragmented literature.

Jia-Qi Wei, Xiang Zhang, Yue-Jin Yang et al. · 0 citations
#artificial intelligence Preprint Aug 2026

Test-Time Scaling in the Wild: Why Exploitation, Not Exploration, Is the Bottleneck

The first compute-normalised comparison of five TTS families across five open-ended generation benchmarks spanning medicine, law, finance, general chat, and creative writing is conducted - grounded in a unified framework that decomposes the effectiveness of each method's token budget into exploration and exploitation.

Davide Romano, Kanak Raj, Jerrod Parker et al. · 0 citations
Preprint Aug 2026

Thought-Level Beam Search for Reasoning

By periodically pruning unpromising trajectories and immediately branching from high-quality prefixes, Gambit dynamically concentrates compute onto the most promising reasoning traces via a light-weight scorer probing hidden states while maintaining continuous high hardware utilization.

Lijie Yang, Hongyin Luo, Jia-Wei Zhao et al. · 0 citations
#machine learning Preprint Oct 2026

How Much Can Language Models Gain from Test-Time Computation?

How much can test-time computation improve a language model, and at what cost? Test-time scaling is widely proposed as a substitute for larger models, but existing comparisons mostly evaluate one domain at a time and rarely charge selection to the budget. We introduce SELF-POT, a benchmark and evaluation framework that...

Bang Yang, Jing-Yuan Li, Jia-Jun Fan et al. · 0 citations
#artificial intelligence Preprint Sep 2026

It's the Problem, Not the Path: Budget and Difficulty Confounds in LLM Reasoning Trajectories

Reasoning traces of large language models are widely read as containing"breakthrough"moments and early-legible fates. Both readings rest on measurements missing a counterfactual control at the level of the claim; we supply both controls. First, a restart-controlled truncation probe separates when a solution fits the co...

Yigit Utku Bulut · 0 citations
Preprint Aug 2026

Funnel of Thoughts: Efficient Test-Time Scaling via Early Voting and Rollout Pruning

Funnel of Thoughts (FoT) is introduced, an inference-time method that preserves the full 32-trajectory voted accuracy while halving its attention FLOPs, a 28.8% reduction in full-model inference cost.

Chanhee Park, Sun Han, Jeongho Yoon et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.