{
  "complete": true,
  "live": false,
  "matchedTrials": 90,
  "generatedAt": "2026-09-22T09:13:22.490Z",
  "expected": 450,
  "recorded": 450,
  "problems": 30,
  "seeds": [
    2025,
    2026,
    2027
  ],
  "config": {
    "version": "aime25-step-search-v1",
    "created": "2026-09-21",
    "model": "Qwen/Qwen3-1.7B",
    "jevModel": "jev-1.13.0",
    "dataset": "math-ai/aime25",
    "problems": 30,
    "seeds": [
      2025,
      2026,
      2027
    ],
    "methods": [
      "nonreasoning",
      "reasoning",
      "self",
      "jev"
    ],
    "concurrency": 4,
    "n": 8,
    "maxSteps": 32,
    "stepTokens": 384,
    "finalTokens": 2048,
    "baselineTokens": 38912,
    "reasoningSampling": {
      "temperature": 0.6,
      "top_p": 0.95,
      "top_k": 20
    },
    "nonreasoningSampling": {
      "temperature": 0.7,
      "top_p": 0.8,
      "top_k": 20
    },
    "candidateSampling": {
      "temperature": 1,
      "top_p": 0.95,
      "top_k": 20
    },
    "judgeSampling": {
      "temperature": 0,
      "top_p": 1,
      "top_k": -1
    },
    "jevUsdPerMillionInputTokens": 0.042,
    "gpu": "RTX 4090 24 GB",
    "gpuUsdPerHour": 0.4111111111111111,
    "instanceId": 51930498,
    "inference": "vLLM 0.10.2, bfloat16, max model length 40960, prefix caching disabled",
    "methodNotes": "Both search methods sample eight non-thinking next-step continuations, select exactly one, and append it to the retained path. No lookahead, answer voting, oracle grading, or execution tools. Same generator, rubric, ordering, budgets, and seeds. Qwen self-judge uses non-thinking constrained categorical output. Jev uses a single Choice over the same candidates. Paths diverge after judge choices. A selected final boxed integer ends search. At step limit, one non-thinking finalization call is used, without judging or resampling.",
    "measurement": "Three fixed seeds over all 30 problems (90 trials per method). Process one method and seed at a time with four concurrent problems. Rotate method order by seed. Time each trajectory end-to-end and each method batch separately. Cost = occupied GPU batch wall time at rental rate plus Jev input-token charges, allocated across completed trajectories. Includes GPU waiting during API calls, excludes installation, warmup, and analysis. Throughput cost and loaded latency are not isolated-request latency or token-list-price estimates.",
    "grading": "Last boxed integer in the final answer, normalized to 0..999; missing/invalid answers and output-budget exhaustion without a final answer count wrong. In thinking mode only the content after </think> is graded. Gold answers are loaded only by summarize.ts, not inference or judges.",
    "uncertainty": "Accuracy is mean of 90 binary trial outcomes. 95% percentile intervals and paired differences use 10000 bootstrap resamples of the 30 problem clusters, retaining all three seeds per problem. Report paired differences rather than a claimed win from small point differences. Only four measured settings, not an exhaustive frontier.",
    "limitations": "Small public benchmark, possible training contamination, three stochastic repeats, one GPU and serving configuration. Search and baseline compute budgets differ; both judges' search budgets match. N=8 is a practical power-of-two search width, not a universal research standard. No accuracy-driven prompt tuning on AIME 2025. Synthetic arithmetic smoke tests only. Never claim expected Jev improvement unless measured.",
    "safety": "Six-hour maximum rental lifetime. Stop and preserve partial data on repeated API errors. Destroy only instance 51930498 after data is synced. No Brex purchase needed unless an account funding problem blocks execution.",
    "modelRevision": "70d244cc86ccca08cf5af4e1e306ecf908b1ad5e",
    "datasetRevision": "563bb8404243c5f09de6ec262f2db674fe5bce9b",
    "infrastructureNote": "The first non-thinking batch aborted after 29 completions when an HTTP client timed out at 300 seconds on a long response. That entire unscored incomplete batch is archived separately, excluded from primary metrics, and rerun with the same seeds and prompts after changing transport to SSE streaming. No AIME accuracy was computed before this transport-only fix. Old rental 51928761 was destroyed. Aborted/setup spend is reported separately."
  },
  "methods": [
    {
      "method": "reasoning",
      "label": "Reasoning",
      "count": 90,
      "correct": 33,
      "accuracy": 0.36666666666666664,
      "ci": [
        0.23333333333333334,
        0.5111111111111111
      ],
      "perSeed": [
        {
          "seed": 2025,
          "count": 30,
          "correct": 10
        },
        {
          "seed": 2026,
          "count": 30,
          "correct": 11
        },
        {
          "seed": 2027,
          "count": 30,
          "correct": 12
        }
      ],
      "meanLatencySeconds": 211.63247825644441,
      "medianLatencySeconds": 193.52464301850029,
      "p95LatencySeconds": 383.7443843956504,
      "gpuSeconds": 5007.3365149149995,
      "activeGpuSeconds": 0,
      "costProvisional": false,
      "gpuCost": 0.5718254662094289,
      "jevCost": 0,
      "totalCost": 0.5718254662094289,
      "costPerProblem": 0.006353616291215876,
      "qwenInputTokens": 17796,
      "qwenOutputTokens": 1613802,
      "jevInputTokens": 0,
      "jevOutputTokens": 0,
      "meanSteps": 0,
      "meanJevSeconds": 0,
      "missingAnswers": 8,
      "outputLimitRuns": 1,
      "stepLimitRuns": 0,
      "truncatedCalls": 1,
      "retries": 0
    },
    {
      "method": "nonreasoning",
      "label": "Non-reasoning",
      "count": 90,
      "correct": 9,
      "accuracy": 0.1,
      "ci": [
        0.03333333333333333,
        0.18888888888888888
      ],
      "perSeed": [
        {
          "seed": 2025,
          "count": 30,
          "correct": 3
        },
        {
          "seed": 2026,
          "count": 30,
          "correct": 3
        },
        {
          "seed": 2027,
          "count": 30,
          "correct": 3
        }
      ],
      "meanLatencySeconds": 24.29410478697774,
      "medianLatencySeconds": 14.61786413099896,
      "p95LatencySeconds": 44.737310422550976,
      "gpuSeconds": 927.6211834500027,
      "activeGpuSeconds": 0,
      "costProvisional": false,
      "gpuCost": 0.10593204872731511,
      "jevCost": 0,
      "totalCost": 0.10593204872731511,
      "costPerProblem": 0.0011770227636368345,
      "qwenInputTokens": 18156,
      "qwenOutputTokens": 285690,
      "jevInputTokens": 0,
      "jevOutputTokens": 0,
      "meanSteps": 0,
      "meanJevSeconds": 0,
      "missingAnswers": 17,
      "outputLimitRuns": 2,
      "stepLimitRuns": 0,
      "truncatedCalls": 2,
      "retries": 0
    },
    {
      "method": "self",
      "label": "Best-of-8 · Qwen judge",
      "count": 90,
      "correct": 6,
      "accuracy": 0.06666666666666667,
      "ci": [
        0,
        0.15555555555555556
      ],
      "perSeed": [
        {
          "seed": 2025,
          "count": 30,
          "correct": 1
        },
        {
          "seed": 2026,
          "count": 30,
          "correct": 2
        },
        {
          "seed": 2027,
          "count": 30,
          "correct": 3
        }
      ],
      "meanLatencySeconds": 156.84938332304446,
      "medianLatencySeconds": 62.90670913749956,
      "p95LatencySeconds": 573.849406564,
      "gpuSeconds": 3854.200034789999,
      "activeGpuSeconds": 0,
      "costProvisional": false,
      "gpuCost": 0.4401401274297221,
      "jevCost": 0,
      "totalCost": 0.4401401274297221,
      "costPerProblem": 0.004890445860330246,
      "qwenInputTokens": 11542314,
      "qwenOutputTokens": 2748176,
      "jevInputTokens": 0,
      "jevOutputTokens": 0,
      "meanSteps": 10.877777777777778,
      "meanJevSeconds": 0,
      "missingAnswers": 12,
      "outputLimitRuns": 1,
      "stepLimitRuns": 17,
      "truncatedCalls": 831,
      "retries": 0
    },
    {
      "method": "jev",
      "label": "Best-of-8 · Jev judge",
      "count": 90,
      "correct": 10,
      "accuracy": 0.1111111111111111,
      "ci": [
        0.03333333333333333,
        0.2
      ],
      "perSeed": [
        {
          "seed": 2025,
          "count": 30,
          "correct": 2
        },
        {
          "seed": 2026,
          "count": 30,
          "correct": 5
        },
        {
          "seed": 2027,
          "count": 30,
          "correct": 3
        }
      ],
      "meanLatencySeconds": 247.52522047419995,
      "medianLatencySeconds": 79.23904134349945,
      "p95LatencySeconds": 811.6240026325506,
      "gpuSeconds": 5929.522450233,
      "activeGpuSeconds": 0,
      "costProvisional": false,
      "gpuCost": 0.6771368230204352,
      "jevCost": 0.46097931600000003,
      "totalCost": 1.1381161390204353,
      "costPerProblem": 0.012645734878004836,
      "qwenInputTokens": 6632620,
      "qwenOutputTokens": 3763246,
      "jevInputTokens": 10975698,
      "jevOutputTokens": 96287,
      "meanSteps": 14.655555555555555,
      "meanJevSeconds": 6.065001201155515,
      "missingAnswers": 24,
      "outputLimitRuns": 3,
      "stepLimitRuns": 26,
      "truncatedCalls": 1189,
      "retries": 0
    },
    {
      "method": "jev-reroll",
      "label": "Best-of-8 · Jev quality + rerolls",
      "count": 90,
      "correct": 11,
      "accuracy": 0.12222222222222222,
      "ci": [
        0.03333333333333333,
        0.23333333333333334
      ],
      "perSeed": [
        {
          "seed": 2025,
          "count": 30,
          "correct": 4
        },
        {
          "seed": 2026,
          "count": 30,
          "correct": 3
        },
        {
          "seed": 2027,
          "count": 30,
          "correct": 4
        }
      ],
      "meanLatencySeconds": 1145.9217828958888,
      "medianLatencySeconds": 1029.5642706504987,
      "p95LatencySeconds": 2593.7595970316497,
      "gpuSeconds": 27126.199111867,
      "activeGpuSeconds": 0,
      "costProvisional": false,
      "gpuCost": 3.0977449603057994,
      "jevCost": 2.161032384,
      "totalCost": 5.258777344305799,
      "costPerProblem": 0.058430859381175544,
      "actualRentalGpuCost": 3.756331879248405,
      "comparisonGpuUsdPerHour": 0.4111111111111111,
      "qwenInputTokens": 27946906,
      "qwenOutputTokens": 14498000,
      "jevInputTokens": 51453152,
      "jevOutputTokens": 656172,
      "meanSteps": 19.822222222222223,
      "meanJevSeconds": 32.5232538673006,
      "missingAnswers": 22,
      "outputLimitRuns": 2,
      "stepLimitRuns": 44,
      "truncatedCalls": 4574,
      "retries": 5,
      "rerolls": 3124,
      "fallbackSteps": 1501
    }
  ],
  "pairs": [
    {
      "other": "reasoning",
      "delta": -0.25555555555555554,
      "ci": [
        -0.35555555555555557,
        -0.15555555555555556
      ],
      "conclusion": "Jev lower in this experiment"
    },
    {
      "other": "self",
      "delta": 0.044444444444444446,
      "ci": [
        -0.02249999999999975,
        0.13333333333333333
      ],
      "conclusion": "Difference unresolved"
    },
    {
      "other": "nonreasoning",
      "delta": 0.011111111111111112,
      "ci": [
        -0.06666666666666667,
        0.08888888888888889
      ],
      "conclusion": "Difference unresolved"
    },
    {
      "method": "jev-reroll",
      "other": "jev",
      "delta": 0.011111111111111112,
      "ci": [
        -0.03333333333333333,
        0.05555555555555555
      ],
      "conclusion": "Difference unresolved"
    },
    {
      "method": "jev-reroll",
      "other": "reasoning",
      "delta": -0.24444444444444444,
      "ci": [
        -0.35555555555555557,
        -0.14444444444444443
      ],
      "conclusion": "Quality + rerolls lower in this follow-up"
    },
    {
      "method": "jev-reroll",
      "other": "self",
      "delta": 0.05555555555555555,
      "ci": [
        -0.044444444444444446,
        0.16666666666666666
      ],
      "conclusion": "Difference unresolved"
    },
    {
      "method": "jev-reroll",
      "other": "nonreasoning",
      "delta": 0.022222222222222223,
      "ci": [
        -0.06666666666666667,
        0.12222222222222222
      ],
      "conclusion": "Difference unresolved"
    }
  ],
  "costFrontier": [
    "nonreasoning",
    "reasoning"
  ],
  "runtimeFrontier": [
    "nonreasoning",
    "reasoning"
  ],
  "abstract": "We compared four methods: an LLM, a reasoning LLM, LLM-as-a-judge, and Jev-as-a-judge. All generation used Qwen3-1.7B. The first two methods disabled or enabled thinking; the latter two used non-thinking Qwen to generate eight candidate next steps and repeatedly selected one using Qwen or Jev 1.13.0, rather than choosing among complete solutions. Across all 30 AIME 2025 problems and three fixed seeds (90 trials per method), accuracy was 10.0%, 36.7%, 6.7%, and 11.1%, respectively. Jev underperformed the reasoning LLM in this setup. Its difference from Qwen self-judging remained unresolved. Jev's paired differences were -25.6 percentage points versus reasoning (95% problem-cluster interval -35.6 to -15.6) and +4.4 versus self-judging (-2.2 to +13.3). On one RTX 4090 at concurrency four, Jev averaged $0.0126 and 247.5 seconds per problem, versus $0.0064 and 211.6 seconds for reasoning. These results describe one small, public-benchmark experiment with unequal compute budgets, not an exhaustive comparison of search policies.",
  "totalMeasuredCost": 7.5147911256927005,
  "followup": {
    "complete": true,
    "completed": 90,
    "total": 90,
    "state": "complete",
    "config": {
      "version": "aime25-jev-reroll-v1",
      "created": "2026-09-21",
      "model": "Qwen/Qwen3-1.7B",
      "jevModel": "jev-1.13.0",
      "dataset": "math-ai/aime25",
      "problems": 30,
      "seeds": [
        2025,
        2026,
        2027
      ],
      "methods": [
        "jev-reroll"
      ],
      "concurrency": 4,
      "n": 8,
      "maxSteps": 32,
      "stepTokens": 384,
      "finalTokens": 2048,
      "baselineTokens": 38912,
      "reasoningSampling": {
        "temperature": 0.6,
        "top_p": 0.95,
        "top_k": 20
      },
      "nonreasoningSampling": {
        "temperature": 0.7,
        "top_p": 0.8,
        "top_k": 20
      },
      "candidateSampling": {
        "temperature": 1,
        "top_p": 0.95,
        "top_k": 20
      },
      "judgeSampling": {
        "temperature": 0,
        "top_p": 1,
        "top_k": -1
      },
      "jevUsdPerMillionInputTokens": 0.042,
      "gpu": "RTX 4090 24 GB",
      "gpuUsdPerHour": 0.49851417478456805,
      "instanceId": 51967955,
      "inference": "vLLM 0.10.2, bfloat16, max model length 40960, prefix caching disabled",
      "methodNotes": "Per-candidate Noul acceptance estimates replace the original relative Choice ranking. Accept the highest-scored candidate at >=0.8; if all fail, sample a fresh batch at the same prefix, up to two rerolls. If all three batches fail, choose the highest-scored candidate across all batches. Stable tie-break: earliest round, then shuffled candidate order. No judge feedback or discarded candidates are passed to Qwen. First-batch generation/shuffle seeds match the original; rerolls use distinct deterministic seed purposes. Accepted-step cap remains 32. Rerolls do not advance that cap. Grading/finalization/generation settings unchanged.",
      "measurement": "Three fixed seeds over all 30 problems (90 trials per method). Process one method and seed at a time with four concurrent problems. Rotate method order by seed. Time each trajectory end-to-end and each method batch separately. Cost = occupied GPU batch wall time at rental rate plus Jev input-token charges, allocated across completed trajectories. Includes GPU waiting during API calls, excludes installation, warmup, and analysis. Throughput cost and loaded latency are not isolated-request latency or token-list-price estimates.",
      "grading": "Last boxed integer in the final answer, normalized to 0..999; missing/invalid answers and output-budget exhaustion without a final answer count wrong. In thinking mode only the content after </think> is graded. Gold answers are loaded only by summarize.ts, not inference or judges.",
      "uncertainty": "Accuracy is mean of 90 binary trial outcomes. 95% percentile intervals and paired differences use 10000 bootstrap resamples of the 30 problem clusters, retaining all three seeds per problem. Report paired differences rather than a claimed win from small point differences. Only four measured settings, not an exhaustive frontier.",
      "limitations": "Follow-up chosen after inspecting the original four-method result. Changes both judging primitive/ranking and reroll policy, so not a pure reroll ablation. Noul values estimate step acceptability, not calibrated eventual-solve probability. More compute allowed. Same GPU model and serving configuration, separate rental/time; physical host and load may differ. Report paired descriptive intervals and actual plus common-rate cost accounting. Original 360 trials remain immutable.",
      "safety": "Six-hour inference limit, seven-hour rental watchdog, $10 estimated inference spend cap (excludes taxes/network/setup). Explicit instance-scoped cleanup and verification, durable results on trusted VM. No gold available to the inference process.",
      "modelRevision": "70d244cc86ccca08cf5af4e1e306ecf908b1ad5e",
      "datasetRevision": "563bb8404243c5f09de6ec262f2db674fe5bce9b",
      "infrastructureNote": "Initial rental 51966897 was destroyed and verified absent after a roughly 12-minute image-pull stall; no AIME inference ran on it. Its setup expense is kept separately. The active follow-up uses a separate Swedish RTX 4090 host from the original study, with otherwise matching serving settings. Source hashes are frozen before AIME evaluation.",
      "threshold": 0.8,
      "maxRerolls": 2,
      "maxInferenceHours": 6,
      "maxEstimatedInferenceUsd": 10,
      "comparisonGpuUsdPerHour": 0.4111111111111111,
      "followup": true,
      "preregisteredAt": "2026-09-21T21:44:34.625Z",
      "machineId": 48093,
      "gpuDriver": "590.48.01",
      "costComparison": "Use original $0.4111111111111111/hour GPU rate for the five-method plot; retain the new actual $0.49851417478456805/hour rental rate in expenses. Both include measured Jev token charges."
    },
    "method": {
      "method": "jev-reroll",
      "label": "Best-of-8 · Jev quality + rerolls",
      "count": 90,
      "correct": 11,
      "accuracy": 0.12222222222222222,
      "ci": [
        0.03333333333333333,
        0.23333333333333334
      ],
      "perSeed": [
        {
          "seed": 2025,
          "count": 30,
          "correct": 4
        },
        {
          "seed": 2026,
          "count": 30,
          "correct": 3
        },
        {
          "seed": 2027,
          "count": 30,
          "correct": 4
        }
      ],
      "meanLatencySeconds": 1145.9217828958888,
      "medianLatencySeconds": 1029.5642706504987,
      "p95LatencySeconds": 2593.7595970316497,
      "gpuSeconds": 27126.199111867,
      "activeGpuSeconds": 0,
      "costProvisional": false,
      "gpuCost": 3.0977449603057994,
      "jevCost": 2.161032384,
      "totalCost": 5.258777344305799,
      "costPerProblem": 0.058430859381175544,
      "actualRentalGpuCost": 3.756331879248405,
      "comparisonGpuUsdPerHour": 0.4111111111111111,
      "qwenInputTokens": 27946906,
      "qwenOutputTokens": 14498000,
      "jevInputTokens": 51453152,
      "jevOutputTokens": 656172,
      "meanSteps": 19.822222222222223,
      "meanJevSeconds": 32.5232538673006,
      "missingAnswers": 22,
      "outputLimitRuns": 2,
      "stepLimitRuns": 44,
      "truncatedCalls": 4574,
      "retries": 5,
      "rerolls": 3124,
      "fallbackSteps": 1501
    },
    "pairs": [
      {
        "method": "jev-reroll",
        "other": "jev",
        "delta": 0.011111111111111112,
        "ci": [
          -0.03333333333333333,
          0.05555555555555555
        ],
        "conclusion": "Difference unresolved"
      },
      {
        "method": "jev-reroll",
        "other": "reasoning",
        "delta": -0.24444444444444444,
        "ci": [
          -0.35555555555555557,
          -0.14444444444444443
        ],
        "conclusion": "Quality + rerolls lower in this follow-up"
      },
      {
        "method": "jev-reroll",
        "other": "self",
        "delta": 0.05555555555555555,
        "ci": [
          -0.044444444444444446,
          0.16666666666666666
        ],
        "conclusion": "Difference unresolved"
      },
      {
        "method": "jev-reroll",
        "other": "nonreasoning",
        "delta": 0.022222222222222223,
        "ci": [
          -0.06666666666666667,
          0.12222222222222222
        ],
        "conclusion": "Difference unresolved"
      }
    ],
    "cleanupVerified": true,
    "resumption": {
      "version": "aime25-reroll-resume-1",
      "created": "2026-09-21T22:32:02.124Z",
      "retained": [
        {
          "key": "jev-reroll-2025-24",
          "id": "24",
          "seed": 2025,
          "traceSha256": "ddfd4057f3784f78514fe99f9d449b07a659de1a5992c0e532deb3450c91b5fd"
        },
        {
          "key": "jev-reroll-2025-16",
          "id": "16",
          "seed": 2025,
          "traceSha256": "aa96d5f1a76de1e2008f5ac7610d6fbddcdb75f89521782c8eb090a3439fb145"
        },
        {
          "key": "jev-reroll-2025-7",
          "id": "7",
          "seed": 2025,
          "traceSha256": "ec44271c046e764df0e72186a362c0b32800ffc92950e0856d74eac017ac544b"
        },
        {
          "key": "jev-reroll-2025-9",
          "id": "9",
          "seed": 2025,
          "traceSha256": "9d554dc1e394d307b4c0a590a065e530a0bbd6bd4e1af37ed36d53b5fb7c6b41"
        },
        {
          "key": "jev-reroll-2025-19",
          "id": "19",
          "seed": 2025,
          "traceSha256": "75acfe11e8e17c8cbc4b4c225e68d2591a9ff8aeffba5a2949435068ee20fa2f"
        },
        {
          "key": "jev-reroll-2025-4",
          "id": "4",
          "seed": 2025,
          "traceSha256": "6e267c5cec33b91004a29ac9fcc4e7577de30d7a0d1d94444ffe08ea641e1d68"
        },
        {
          "key": "jev-reroll-2025-11",
          "id": "11",
          "seed": 2025,
          "traceSha256": "fe7b61b6fd6aefe9411ed738ba81ad4b3b77ea977b73d3be593498e7e2ca8d79"
        },
        {
          "key": "jev-reroll-2025-25",
          "id": "25",
          "seed": 2025,
          "traceSha256": "0acb7bce4aa546c501b4b4f096a159eff80976e2b210c22b1c20c6ce9ae5807c"
        },
        {
          "key": "jev-reroll-2025-26",
          "id": "26",
          "seed": 2025,
          "traceSha256": "5c5290c0e504c6f0e6385a95b0aaeb9f7f3f6c39caffadfcfa8ae61889190c79"
        },
        {
          "key": "jev-reroll-2025-13",
          "id": "13",
          "seed": 2025,
          "traceSha256": "9e696200a05a55f963ab66afa36b9836d338819baed72e0611ddc6e1817c51bd"
        }
      ],
      "runsSha256": "2c89ba6a379cf574211ab060fbac6e28a0dfb3a096c23894b8373a913c816140",
      "priorSettingsHash": "815a4504edbd7d576af2d400592361b95a88edd03296928de99bb245fe1c5d85",
      "remaining": 80,
      "transport": {
        "jevAttempts": 5,
        "qwenAttempts": 3,
        "jevTimeoutMs": 120000,
        "qwenTimeoutMs": 1200000,
        "backoffMs": 2000
      },
      "maxInferenceHours": 8,
      "maxEstimatedInferenceUsd": 10,
      "priorExpenseUsd": 0.6636699254719495,
      "note": "User requested completion of the remaining 80 trials after a Jev timeout. Retain all ten completed trials unchanged. Restart only interrupted/unstarted trials using their original seeds; preserve the interrupted trace. Sampling, scoring, threshold, reroll policy, accepted-step cap and finalizer are unchanged. Add bounded transport retries and attempt logging. Use a new RTX 4090 with the same pinned vLLM/model configuration. Charge both attempts GPU wall time and all known Jev usage, including interrupted work. Runtime adds interrupted active trajectory time to its restarted trial, excluding the gap between rentals. Eight-hour continuation deadline and nine-hour rental watchdog are operational safeguards, not accuracy-based stopping."
    },
    "continuationRental": {
      "id": 51973741,
      "gpuUsdPerHour": 0.49851417478456805,
      "started": "2026-09-21T22:38:09.489Z"
    },
    "note": "Follow-up chosen after the original results. Replaces relative Choice ranking with independent Noul step-acceptability scores and adds rerolls, so it is not a pure reroll ablation. Threshold 0.8, at most two extra batches of eight per step; after three failed batches, use the highest-scored candidate seen. Scores are not calibrated probabilities of eventually solving the problem. Same generator, seeds, accepted-step cap, finalizer and grader. Separate RTX 4090 rental. Plotted GPU cost uses the original hourly rate for comparability; actual rental expenses are reported separately. User requested completion of the remaining 80 trials after a Jev timeout. Retain all ten completed trials unchanged. Restart only interrupted/unstarted trials using their original seeds; preserve the interrupted trace. Sampling, scoring, threshold, reroll policy, accepted-step cap and finalizer are unchanged. Add bounded transport retries and attempt logging. Use a new RTX 4090 with the same pinned vLLM/model configuration. Charge both attempts GPU wall time and all known Jev usage, including interrupted work. Runtime adds interrupted active trajectory time to its restarted trial, excluding the gap between rentals. Eight-hour continuation deadline and nine-hour rental watchdog are operational safeguards, not accuracy-based stopping."
  }
}