One Kept Candidate Flipped −8.6pp: Reviewing SELF-POT on Test-Time Scaling
We check whether test-time compute lets cheap models match expensive ones, against a controlled cost-accuracy benchmark across 5 models and 4 domains.
- Coding Parallel@4 correct submissions were 376/500 scheduled cells (−8.6pp), but keeping the first valid candidate on selection failure alone raised it to 437 (+3.6pp), and public-example selection to 453 (+6.8pp) — saving 12–49% of cost per model.
- The same five low-cost models under the same budget multiple moved in opposite directions by domain — math Parallel@4 averaged +4.0pp (95% CI [+0.3,+7.7]), coding −8.6pp ([−11.8,−5.4]), workflow −4.6pp ([−7.4,−1.9]).
- On workflow and terminal tasks, Claude Opus 5.5's single Direct answer beat every low-cost model under every test-time protocol (workflow 148/160 vs. best 78/160; terminal 13/30 vs. best 3/30).
























































