Can step selection beat thinking?
Thinking solves 3.3× as many AIME problems as Jev step selection.
Qwen3-1.7B × Jev 1.13.0 · AIME 2025 · N = 8
Accuracy
30 problems · three fixed seeds · 90 trials per method · exact integer match
Loading completed trials. No simulated data.
Accuracy versus resources
Higher and further left is better. The line joins non-dominated measured settings, not an exhaustive frontier.
Waiting for measured accuracy and resource use.
Background
Can a small model solve harder problems by picking better steps instead of thinking longer? We gave Qwen3-1.7B two ways to spend extra compute: use its built-in thinking mode, or generate eight possible next steps and let a judge pick one.
Methodology
We ran all 30 AIME 2025 problems with three fixed seeds per method. The original four methods were thinking, no thinking, Qwen-picked steps, and Jev-picked steps. A later 90-run follow-up let Jev reject a batch and ask for up to two more. Every method used the same Qwen weights and exact-answer grading; search selected steps, not complete solutions.
Results
Thinking scored 36.7%, versus 11.1% for Jev step selection, 10.0% without thinking, and 6.7% for Qwen step selection. Rerolls reached 12.2%, still only a third of thinking’s score. Thinking also cost less per problem than either Jev method. Extra candidates did not close the gap.
Discussion
The gap between thinking and Jev was clear here: 25.6 percentage points, with a 95% problem-cluster interval of 15.6 to 35.6 points. Jev’s smaller lead over Qwen judging was unresolved. This tests one search recipe with unequal compute budgets on a small public dataset, not every way to use a judge. The reroll follow-up changed both the judging rule and the retry policy, so it cannot isolate either effect.
Conclusion
Turn on native thinking before building a step-selection loop for this model. Jev judging and extra rerolls added work without catching up. A better search policy needs to beat that simple baseline before it earns a place in the stack.
Data and run details
Loading experiment status…
Full result summary
Follow-up: rerolls
Loading 90 trials.
Jev accepts a step at ≥0.8. Otherwise it tries up to two more batches of eight and keeps the best seen. Generator, seeds, limits, and grading stay fixed.
This changes judging and rerolls together, so neither effect is isolated. The original 360 trials stay unchanged; costs use the original GPU rate.
Follow-up protocolFollow-up sourceCombined resultsFollow-up auditFollow-up expensesOriginal resultsOriginal traces
Did Jev win?
Paired 95% intervals resample 30 problems, keeping all three seeds together.
What we measured
Qwen proposes eight next steps; Qwen or Jev picks one. Compare that with Qwen thinking on and off. Judges select steps, not complete answers.
30 problems × three seeds. Search allows 32 steps of 384 tokens and one 2,048-token finalization; baselines cap output at 38,912 tokens. Sampling settings and stopping rules are in the protocol.
Budgets, timing, and caveats
Only the judge changes between search methods: Qwen makes a constrained, non-thinking A–H choice; Jev makes one typed Choice over the same state structure and rubric.
Non-reasoning means Qwen's thinking mode is disabled, not that it cannot show working. Both baselines receive the same request to reason step by step and put the final answer in a box.
All methods run on the same rented RTX 4090 with vLLM 0.10.2, bfloat16, no quantization, and prefix caching disabled. Four problems run concurrently; each proposal call batches eight continuations. Three fixed seeds (2025, 2026, 2027) cover all 30 questions. Method order rotates between seeds. Sampling seeds do not guarantee bitwise reproducibility under dynamic batching.
Runtime is observed mean end-to-end problem latency under that load, including judge calls. Cost allocates each method's occupied GPU batch wall time at $0.411111/hour across its 30 trials, then adds Jev input-token charges at $0.042/million. GPU waiting for Jev is included. Setup, warmup, downloads, and this static website are excluded. This is amortized rental cost, not a token-provider quote or isolated-request latency.
Search and baseline compute budgets differ. The two search judges share generator settings, limits, rubric, and seeded candidate ordering; their paths diverge after selections. The self-judge has no hidden thinking or free-form critique, so this does not establish a comparison against every possible Qwen judge design. There is no random-selector or N = 1 stepwise ablation.
Gold answers are never passed to generators or judges. Prompts were smoke-tested on synthetic arithmetic, then frozen before AIME evaluation. Missing or invalid final answers count wrong. For the thinking baseline, only the answer after the closing think tag is graded. No prompt or seed is selected using AIME accuracy.
The first batch hit a 300-second HTTP-client timeout after 29 non-thinking completions. That incomplete, unscored batch was archived and restarted with streamed transport; prompts and seeds were unchanged. Its spend is excluded from method-level inference costs.
AIME 2025 is small and public. Training contamination is unknown. Problem-cluster bootstrap intervals describe this experiment, not performance on unseen competition math. The frontier compares four configurations on one serving setup, not a full N or compute sweep.
Token limits and stopping
Invalid or missing boxed integers count wrong. A step-limit hit triggers one finalization call. Candidate steps can themselves hit the 384-token cap; every finish reason is in the full traces.
Inspect the run
A Jev trace logging bug required labeled reconstructions. Raw logs and scores are unchanged. Trace note · Original audit · Parser check · Expense estimate
Original protocolOriginal sourceLive results JSONLive trials JSONFull tracesSource hashes
Results by problem
Sources: math-ai/aime25 · Qwen3-1.7B · Jev model and pricing