Can step selection beat thinking?
Qwen3-1.7B × Jev 1.13.0 · AIME 2025 · N = 8
Loading experiment status…
Abstract
We compared four methods: an LLM, a reasoning LLM, LLM-as-a-judge, and Jev-as-a-judge. Loading measured results.
Follow-up: quality-gated rerolls
Preparing 90 additional trials.
Jev rates each candidate step independently. Accept the highest score at ≥0.8; otherwise draw a fresh batch of eight, up to two extra batches per step. If all three batches fail, retain the highest-scored candidate seen. The generator, three seeds, accepted-step cap, finalizer, and grading stay fixed.
This follow-up changes both the judge's scoring method and rerolling, so it cannot isolate the effect of rerolling alone. Scores estimate step acceptability, not calibrated chances of solving the problem. The original 360 trials remain unchanged. The new run uses a separate RTX 4090 rental; plotted GPU costs use the original hourly rate. Live results favor faster finishes and remain provisional.
Follow-up protocolFollow-up sourceCombined resultsFollow-up auditFollow-up expensesOriginal resultsOriginal traces
Accuracy
30 problems · three fixed seeds · 90 trials per method · exact integer match
Loading completed trials. No simulated data.
Accuracy versus resources
Higher and further left is better. The line joins non-dominated measured settings, not an exhaustive frontier.
Waiting for measured accuracy and resource use.
Did Jev win?
Paired 95% intervals resample the 30 problems as clusters, keeping all three seeds together. Repeated attempts are not 90 independent problems.
What we measured
Qwen generates eight possible next steps. A judge picks one, and the process repeats. We compare Qwen and Jev as judges against Qwen with thinking on and off. This selects steps, not complete solutions.
Thinking and non-thinking baselines use Qwen's recommended sampling settings and a 38,912-token output cap. Search uses temperature 1.0, up to 32 steps of 384 tokens each, and one 2,048-token finalization if needed. A selected boxed final answer ends search early. N = 8 is a practical power-of-two choice, not a universal research standard.
Budgets, timing, and caveats
Only the judge changes between search methods: Qwen makes a constrained, non-thinking A–H choice; Jev makes one typed Choice over the same state structure and rubric.
Non-reasoning means Qwen's thinking mode is disabled, not that it cannot show working. Both baselines receive the same request to reason step by step and put the final answer in a box.
All methods run on the same rented RTX 4090 with vLLM 0.10.2, bfloat16, no quantization, and prefix caching disabled. Four problems run concurrently; each proposal call batches eight continuations. Three fixed seeds (2025, 2026, 2027) cover all 30 questions. Method order rotates between seeds. Sampling seeds do not guarantee bitwise reproducibility under dynamic batching.
Runtime is observed mean end-to-end problem latency under that load, including judge calls. Cost allocates each method's occupied GPU batch wall time at $0.411111/hour across its 30 trials, then adds Jev input-token charges at $0.042/million. GPU waiting for Jev is included. Setup, warmup, downloads, and this static website are excluded. This is amortized rental cost, not a token-provider quote or isolated-request latency.
Search and baseline compute budgets differ. The two search judges share generator settings, limits, rubric, and seeded candidate ordering; their paths diverge after selections. The self-judge has no hidden thinking or free-form critique, so this does not establish a comparison against every possible Qwen judge design. There is no random-selector or N = 1 stepwise ablation.
Gold answers are never passed to generators or judges. Prompts were smoke-tested on synthetic arithmetic, then frozen before AIME evaluation. Missing or invalid final answers count wrong. For the thinking baseline, only the answer after the closing think tag is graded. No prompt or seed is selected using AIME accuracy.
The first batch hit a 300-second HTTP-client timeout after 29 non-thinking completions. That incomplete, unscored batch was archived and restarted with streamed transport; prompts and seeds were unchanged. Its spend is excluded from method-level inference costs.
AIME 2025 is small and public. Training contamination is unknown. Problem-cluster bootstrap intervals describe this experiment, not performance on unseen competition math. The frontier compares four configurations on one serving setup, not a full N or compute sweep.
Token limits and stopping
Invalid or missing boxed integers count wrong. A step-limit hit triggers one finalization call. Candidate steps can themselves hit the 384-token cap; every finish reason is in the full traces.
Inspect the run
The original four-method audit found a stored-history logging bug in Jev requests. Labeled reconstructions are included alongside unchanged raw logs; inference and scores are unchanged. Trace note · Original audit · Independent parser check · Original expense estimate
Original protocolOriginal sourceLive results JSONLive trials JSONFull tracesSource hashes
Results by problem
Sources: math-ai/aime25 · Qwen3-1.7B · Jev model and pricing