Benchmarks

Can step selection beat thinking?

Qwen3-1.7B × Jev 1.13.0 · AIME 2025 · N = 8

Loading experiment status…

Abstract

We compared four methods: an LLM, a reasoning LLM, LLM-as-a-judge, and Jev-as-a-judge. Loading measured results.

Accuracy

30 problems · three fixed seeds · 90 trials per method · exact integer match

Loading completed trials. No simulated data.

Accuracy versus resources

Higher and further left is better. The line joins non-dominated measured settings, not an exhaustive frontier.

Waiting for measured accuracy and resource use.

What we measured

Qwen generates eight possible next steps. A judge picks one, and the process repeats. We compare Qwen and Jev as judges against Qwen with thinking on and off. This selects steps, not complete solutions.

Thinking and non-thinking baselines use Qwen's recommended sampling settings and a 38,912-token output cap. Search uses temperature 1.0, up to 32 steps of 384 tokens each, and one 2,048-token finalization if needed. A selected boxed final answer ends search early. N = 8 is a practical power-of-two choice, not a universal research standard.

Budgets, timing, and caveats

Only the judge changes between search methods: Qwen makes a constrained, non-thinking A–H choice; Jev makes one typed Choice over the same state structure and rubric.

Non-reasoning means Qwen's thinking mode is disabled, not that it cannot show working. Both baselines receive the same request to reason step by step and put the final answer in a box.

All methods run on the same rented RTX 4090 with vLLM 0.10.2, bfloat16, no quantization, and prefix caching disabled. Four problems run concurrently; each proposal call batches eight continuations. Three fixed seeds (2025, 2026, 2027) cover all 30 questions. Method order rotates between seeds. Sampling seeds do not guarantee bitwise reproducibility under dynamic batching.

Runtime is observed mean end-to-end problem latency under that load, including judge calls. Cost allocates each method's occupied GPU batch wall time at $0.411111/hour across its 30 trials, then adds Jev input-token charges at $0.042/million. GPU waiting for Jev is included. Setup, warmup, downloads, and this static website are excluded. This is amortized rental cost, not a token-provider quote or isolated-request latency.

Search and baseline compute budgets differ. The two search judges share generator settings, limits, rubric, and seeded candidate ordering; their paths diverge after selections. The self-judge has no hidden thinking or free-form critique, so this does not establish a comparison against every possible Qwen judge design. There is no random-selector or N = 1 stepwise ablation.

Gold answers are never passed to generators or judges. Prompts were smoke-tested on synthetic arithmetic, then frozen before AIME evaluation. Missing or invalid final answers count wrong. For the thinking baseline, only the answer after the closing think tag is graded. No prompt or seed is selected using AIME accuracy.

The first batch hit a 300-second HTTP-client timeout after 29 non-thinking completions. That incomplete, unscored batch was archived and restarted with streamed transport; prompts and seeds were unchanged. Its spend is excluded from method-level inference costs.

AIME 2025 is small and public. Training contamination is unknown. Problem-cluster bootstrap intervals describe this experiment, not performance on unseen competition math. The frontier compares four configurations on one serving setup, not a full N or compute sweep.

Inspect the run

The original four-method audit found a stored-history logging bug in Jev requests. Labeled reconstructions are included alongside unchanged raw logs; inference and scores are unchanged. Trace note · Original audit · Independent parser check · Original expense estimate

Original protocolOriginal sourceLive results JSONLive trials JSONSource hashes

Sources: math-ai/aime25 · Qwen3-1.7B · Jev model and pricing