← 피드로
[Submitted on 4 Aug 2026 (v1), last revised 31 Aug 2026 (this version, v2)]
Authors:Mohsen Hariri, Weicong Chen, Nahal Shahini, Vikash Singh, Kai Ye, Amirhossein Samandar, Debargha Ganguly, Sreehari Sankar, Yanyan Zhang, Shouren Wang, Jerry Peng, Biyao Zhang, Michael Hinczewski, Vipin Chaudhary
Abstract:Large language models can solve harder reasoning problems with more inference-time compute. The term “test-time scaling,” however, covers several inference algorithms: extending deliberation along one trajectory, sampling completed candidates and aggregating them by voting or verification, and searching over partial states. These algorithms differ in statistical structure, compute requirements, and failure modes. Treating them as interchangeable under a scalar “budget,” or reporting accuracy without specifying the inference protocol, makes results difficult to compare across studies. We study test-time scaling along three axes. First, we formalize it as budgeted inference over the implicit prefix tree of an autoregressive model and distinguish single-trajectory sequential scaling, leaf-level scaling with terminal reduction, and prefix-level scaling. Second, we treat the full inference system as the evaluated object and separate end-to-end performance from candidate-bank diagnostics. We introduce an evaluation profile whose coordinates and simple functionals recover or bound common repeated-sampling metrics, and require compute accounting and uncertainty estimates that match the protocol. Third, we distinguish exact replay from distributional reproducibility and state the requirements for each. We also organize open-weight reasoning models by model-side and interface mechanisms. Our empirical study covers broad knowledge, symbolic reasoning, and competition mathematics, and we publicly release 1,403,520 sampled model attempts. The project website is available at this https URL. The released datasets are Trace (this https URL), Lite (this https URL), Math (this https URL), and SuperGPQA (this https URL).
Submission history
From: Mohsen Hariri [view email]
[v1]
Tue, 4 Aug 2026 17:57:20 UTC (553 KB)
[v2]
Mon, 31 Aug 2026 18:11:17 UTC (639 KB)
추출 본문 · 출처: arxiv.org · https://arxiv.org/abs/2608.04001