Source Document
Bangji Yang, Jingyuan Li, Jiajun Fan, Yi Evie Zhang, Ruihan Guo, Hongba Ma, Neil He, Chumeng Liang, Qinglong Zheng, Zhanghan Ni, Ge Liu, "How Much Can Language Models Gain from Test-Time Computation?", arXiv:2610.01110 [cs.LG], submitted 2026-10-01 (v1), DOI 10.48550/arXiv.2610.01110. Affiliation: University of Illinois at Urbana-Champaign, University of Washington, Tsinghua University. The full text was checked against a primary-source snapshot taken 2026-10-04T22:10:07Z (session egress was fully blocked — three retries at 10-minute intervals before falling back).
This is a preprint without peer review. The text shows no venue submission marker, and code/data release is not confirmed — unlike other papers reviewed in this series, no external replication path is stated. All authors are university-affiliated, and neither the body nor the references disclose a funding source or a conflict-of-interest statement. No overlap was found between the authors' affiliations and the vendors of the evaluated models (OpenAI, Alibaba, Zhipu AI, MiniMax, DeepSeek, Anthropic).
Study Overview
Three questions drive the paper: under a fixed budget multiple, which parallel and sequential test-time methods raise accuracy, and in which domains; on static tasks, where is opportunity lost among generation, selection and revision; and can a low-cost model with added test-time compute match a stronger model at equal cost.
The authors built SELF-POT, an evaluation framework that seals 350 tasks across mathematics (APEX, Omni-MATH; 60 tasks), competitive programming (LiveCodeBench AtCoder; 100 tasks), app workflows (AppWorld; 160 tasks) and terminal operation (Terminal-Bench; 30 tasks) before any run, so tasks cannot be cherry-picked after seeing results. Five low-cost reasoning models (DeepSeek V4.1 Flash, Qwen3-Next-80B-A3B, GLM-5, GPT-OSS-120B, MiniMax M2.5) were run under four protocols — Direct, Parallel@2, Parallel@4 and Revise@4 — with Claude Opus 5.5's Direct answer as the anchor reference. Every model call, including selection and critique, is charged, so the budget is denominated in dollars rather than tokens — the design choice that sets this apart from prior comparisons.
Key Results
Accuracy (by task count) and mean cost, by domain and model.
| Domain (tasks) | Model | Direct acc. / cost | Parallel@4 acc. / cost |
|---|---|---|---|
| Math (60) | Opus 5.5 (anchor) | 95.0% / $0.103 | not run |
| DeepSeek Flash | 76.7% / $0.051 | 86.7% / $0.176 | |
| Code (100) | Opus 5.5 (anchor) | 93.0% / $0.028 | not run |
| GLM-5 | 66.0% / $0.103 | 55.0% / $0.377 | |
| Workflow (160) | Opus 5.5 (anchor) | 92.5% / $0.097 | not run |
| DeepSeek Flash | 48.7% / $0.020 | 50.6% / $0.026 | |
| Terminal (30) | Opus 5.5 (anchor) | 43.3% / $1.930 | not run |
| DeepSeek Flash | 10.0%‡ / $0.307 | 10.0% / $0.438 |
The mean paired change (Δ, against Direct) across the five-model low-cost panel runs in opposite directions by domain.
| Domain | Parallel@2 Δ (95% interval) | Parallel@4 Δ (95% interval) | Revise@4 Δ (95% interval) |
|---|---|---|---|
| Math | +0.7pp [-0.7,+2.0] | +4.0pp [+0.3,+7.7] | +0.8pp [-1.7,+3.3] |
| Code | +6.2pp [+4.2,+8.2] | −8.6pp [-11.8,-5.4] | −0.5pp [-1.5,+0.0] |
| Workflow | −4.2pp [-7.2,-1.4] | −4.6pp [-7.4,-1.9] | −7.2pp [-11.2,-3.4] |
| Terminal | 0.0pp [-5.6,+5.6] | 0.0pp [-4.4,+4.4] | −5.0pp [-10.0,+0.0] |
Code's −8.6pp under Parallel@4 traces to selection failure, not generation failure. On the same candidate pools (500 scheduled cells), simply retaining the first available candidate when judging returns no choice raised correct submissions from 376 to 437, flipping the change from −8.6pp to +3.6pp; switching to public-example-based selection raised it further to 453 (+6.8pp), while saving 12–49% of API cost across models. On mathematics, judging with the same fallback yielded 186 correct submissions on the same 292 retained pools versus 182 for majority voting, while voting saved 12–21% of cost — the 95% interval for that gap crosses zero.
Credibility Assessment
What earns trust: the model, information and grading are held fixed while only the protocol varies, and all 350 tasks were sealed before evaluation to block post-hoc task selection. Replaying the same retained candidate pools under alternative selection rules isolates the effect of the selection rule itself, so protocol effects and selection-rule effects are not conflated.
What to weigh: this is an unreviewed v1 preprint, and code/data release is not stated in the text, leaving external replication weak. The panel of five low-cost models plus one anchor is a small sample. Workflow and terminal domains show high rates of unfinished protocol execution (117/160 stage failures for GPT-OSS, 159/160 for MiniMax in workflows) — the authors themselves note these numbers often measure "interface reliability" rather than model capability. Mathematics Parallel@4's +4.0pp has a 95% interval of [+0.3, +7.7] that barely clears zero, and Flash's math gain (46→52) is nominal p=0.031 but not significant after Holm correction.
Related Work
- Snell et al. (2024), "Scaling LLM Test-Time Compute Optimally Can Be More Effective Than Scaling Model Parameters" — ICLR 2025 — prior work. Comparing smaller/larger models from the same family at matched FLOPs, it first showed added compute can let a small model beat one 14× larger, but mainly in mathematics and in FLOPs units. SELF-POT extends this to four domains and dollar cost, showing the conclusion does not transfer as-is to coding or workflows.
- Wu et al. (2024), "Inference Scaling Laws: An Empirical Analysis of Compute-Optimal Inference for LLM Problem-Solving" — ICLR 2025 — prior work. Showed a tree-search Llemma-7B beating majority-vote Llemma-34B, again in one domain (math) and FLOPs units. SELF-POT's selection-rule replay reconfirms the same direction — that search/selection method matters more than model size — in coding as well.
- Stroebl, Kapoor & Narayanan (2024/2026), "The Limits of Inference Scaling Through Resampling" — ICLR 2026 — contradiction/complement. Showed a theoretical ceiling: when a verifier has a nonzero false-positive rate, resampling alone cannot raise accuracy past that ceiling. SELF-POT's RQ3 finding (Opus Direct dominating every low-cost model and protocol on workflow and terminal tasks) reads as an empirical instance of that same limit.
All three were identified independently via WebSearch (session egress blocked the full text, so titles, authors and venue were confirmed via search results) and cross-checked against SELF-POT's own bibliography.
Reviewer's Judgement
First, the most practically useful result here is not an average like +4.0pp or −8.6pp but the replay finding that fixing one selection failure was cheaper and more effective than buying one more candidate. Going from 376 to 437 correct submissions in coding at no added cost means: check your failure-handling logic before increasing the test-time scaling budget.
Second, the "low-cost model plus test-time compute replaces a stronger model" narrative should not be taken at face value. It holds only in parts of math and coding; in workflows (148/160 versus the best low-cost score of 78/160) and terminal tasks (13/30 versus 3/30), Opus Direct dominated every low-cost model under every protocol. The pattern reads as: the more a task mixes irreversible actions, the smaller the gain from test-time compute — and the low completion rates in workflows (117/160 stage failures for GPT-OSS, 159/160 for MiniMax) are the same story from a different angle, measuring harness and interface reliability rather than model competence.
Putting It to Work
- Fix selection rules first — before buying more inference, replace "discard when the judge returns nothing" with "keep the first valid candidate." The coding case flipped −8.6pp to +3.6pp at no added cost.
- Vary the test-time strategy by domain — the same budget multiple moves accuracy in opposite directions in math/coding versus workflows/terminal tasks.
- Track cost in dollars, not tokens — include selection and critique calls in the per-task cost, or the accuracy-per-cost picture is distorted.
- Measure completion rate first in agent workflows — check whether the protocol even finished executing before tallying accuracy, to separate harness defects from model capability.
Conclusion
SELF-POT shows, with a controlled cost-accuracy design, that the payoff from test-time scaling depends heavily on domain and on failure handling. The same added budget that averaged −8.6pp in coding flipped to +3.6pp with a single selection-rule fix, while in workflows and terminal tasks no low-cost model/protocol combination beat the stronger model's single Direct answer.
That said, this is an unreviewed preprint without a confirmed code/data release, leaving external replication weak, and the five-model panel is a limited sample. The conclusion that checking selection and failure-handling logic before raising the inference budget pays off more per dollar lines up with the operating principle in Splitting Reasoning Budgets by Task Type.
References
- How Much Can Language Models Gain from Test-Time Computation? — arXiv abstract (source)
- Full HTML text of the same paper — used to verify tables and cost figures
- Splitting Reasoning Budgets by Task Type — sunny34.com blog
Ask AI about this review
The assistant has read this review and the numbers it verified. Ask anything — it answers from the text and says so when something isn't in it.
Loading the chat…