GPT-6 Luna compaction · methods note
Designing a small experiment that can actually tell you something
Notes from rerunning the GPT-6 Luna compaction benchmark. Version 1 looked finished: frozen protocol, immutable request logs, audit, live page. It still could not answer its own question. This is what was wrong, what version 2 does about it, and the general rule behind each fix. Written for Rubric developers who want to run experiments, not just demos.
Version 1: https://benchmarks.rubric.sh/archive/luna-compaction-v1/ Version 2: https://benchmarks.rubric.sh/luna-compaction/
1. Count the information, not the trials
v1 ran 20 tasks once per arm. The headline read 16/20 vs 14/20. But in a paired design only discordant pairs carry information: 13 tasks passed in both arms, 3 failed in both, and the whole result rested on 4 tasks (3 encrypted-only, 1 textual-only). Exact sign test p = 0.63. The bootstrap interval (-10 to +30 pp) said the same thing, but the page led with the raw counts.
Rule: before running, estimate how many informative observations you will get, not how many trials. For paired pass/fail, that is the expected number of discordant pairs. If it is single digits, the study cannot distinguish a real effect from noise. Use a supported finding for the headline. Explain the discordant counts and interval in the results and discussion, rather than selling a noisy pass-rate gap.
v2: 4 repeats per task per arm (192 trials). Repeats do two things: they estimate within-task variance, which v1 assumed away, and they turn each task into a pass rate instead of a coin flip.
2. Make sure the treatment actually happened
Compaction only triggers when context passes 50k tokens. In v1, 8 of 20 tasks never compacted in either arm. On those tasks the two arms were byte-for-byte the same configuration; any difference was run-to-run noise. Yet torch-pipeline-parallelism, with zero compactions in either arm, was counted as an encrypted win. Half of the four informative pairs were placebo.
Rule: define the treatment exposure and check it per unit. Predeclare that the primary analysis uses only units where the treatment occurred. Report the untreated units separately as a placebo set; their estimate should be near zero, and if it is not, that tells you how noisy your measurement is.
v2: task compaction rate = fraction of scored trials (both arms pooled) with at least one qualified compaction. Primary set = rate at least 0.5. Placebo set = rate 0. The rule pools both arms so a task is in or out for both arms together.
3. Select units on exposure, never on outcome
v1 selected tasks from instructions and metadata. That was outcome-blind, good, but it produced tasks that were too short and too easy: 80% pass rate, 13/20 both-pass. A ceiling effect wastes trials the same way a floor effect does.
Rule: select on a covariate that predicts exposure and difficulty, measured before the experiment or under a neutral configuration. Never select or replace units based on how the arms scored. Write the rule down before looking at results.
v2: set A = the twelve v1 tasks that compacted in at least one arm (selection on observed context growth, symmetric across arms). Set B = unused Terminal-Bench 2.1 tasks with agent timeouts of 30 minutes or more in coding-adjacent categories (long timeouts are the dataset’s own proxy for long tasks). Excluded categories are listed in protocol.json. Terminal-Bench 4.0 was considered for a lower ceiling and rejected: its tasks carry 8-hour agent timeouts and 4 to 16 expert-hour estimates, which on a small efficient model means a floor effect and a blown wall clock.
4. Arms must differ in exactly one thing
v1’s arms differed in the thing under test (text summary vs encrypted state) and in several things that were not:
- The encrypted arm dropped every input item before the compaction item, including Pi’s developer message. After the first compaction, that arm ran without its system prompt. The textual arm kept it. Found by reading one request snapshot: input[0] was the compaction item, not the developer message.
- The textual arm needed a resume message after Pi’s idle-time compaction. v1 used “Continue solving the original terminal task from the compacted state. Finish the task and verify the result.” That is an extra instruction, including “verify”, the encrypted arm never received.
- Thresholds used different counters (Pi’s context estimate vs OpenAI’s server-side count) and different residuals (summary plus 20k recent tokens vs a small encrypted item), so the textual arm compacted 41 times to the encrypted arm’s 24.
Rule: list every difference between arms, then for each ask whether it is the treatment or a confound. Remove confounds where you can; measure and report the rest. A request snapshot from each arm at the interesting moment (here, right after compaction) is the cheapest audit there is.
v2: the encrypted arm keeps the developer message. The resume text is “Continue.” The audit checks both on every request. The threshold and residual differences are inherent to the two mechanisms, so they are measured (tokens at compaction, residual after compaction, compactions per trial) rather than hidden.
5. Measure the mechanism, not just the outcome
Pass/fail on 20 tasks is a blunt instrument. If compaction loses information, it should show up long before it flips a verifier: the agent re-reads files it already read, context after compaction is larger or smaller, more requests are needed.
Rule: for every hypothesis about why the treatment might work, add a cheap direct measurement. Mechanistic measures have far more statistical power than the final outcome and tell you whether a null result is “no effect” or “no exposure”.
v2: per arm, compactions per trial, median context tokens at compaction, median residual input tokens on the first request after a compaction, requests, cost, duration, timeouts, and file-read tool calls per request in the five requests after versus before each compaction. The gateway records tool calls per response so this needs no extra instrumentation in the agent.
6. Pair for shared conditions; report what pairing cannot fix
Three v1 trials hit the task’s wall-clock timeout. Timeouts penalise the slower arm, and the textual arm has extra summary round trips. Machine load varied with whichever other task was running.
Rule: run paired arms at the same time on the same machine so load is shared. Record duration and timeouts per arm and show them. State the scoring rule for timeouts explicitly (here, the verifier still runs on the container, so a timed-out trial can pass).
v2: both arms of a pair start together in one Harbor job. Rounds are round-major so an early stop leaves repeats evenly spread. Timeouts are counted per arm and marked per trial.
7. Check your cost model against the provider’s bill
v1 assumed compaction cost was in the usage numbers. Reading the ledger: responses that emit a compaction item report about 15 output tokens while carrying 5 kB of encrypted state. Server-side compaction is effectively free at billed rates. That is a real property of the product, but it has to be stated, otherwise readers assume the cost chart compares like with like.
Rule: whenever a comparison involves money, look at the raw usage of the unusual responses and say what the provider does and does not bill.
8. Publish the uncertainty where people will read it
v1 put “Encrypted 16/20, textual 14/20” in the headline and the interval in the footer. The no-JavaScript HTML said “Partial data, first results pending” forever, which is what crawlers and link previews saw.
Rule: lead with a punchy measured result, then the charts. Keep the interval and informative count in the visible discussion, not buried in a footer or loaded into the headline. Follow the charts with one short paragraph each for background, methodology, results, discussion, and conclusion. We write for developers, not a university. Static HTML carries the same current result as JavaScript. Budget caps and other internal operating details do not belong on the page.
9. Freeze, then keep hands off
Things v1 did right and v2 keeps: protocol frozen with hashes before launch; immutable request snapshots with SHA-256; the real API key never enters a task container; billed usage recomputed by the audit; one attempt per trial, no reruns; audit gates the COMPLETE state and the source archive; task files pinned by hash; actual model ID recorded on every response because the alias has no dated snapshot; synthetic smoke task exercises both arms end to end before spending on the real matrix.
Checklist for the next one
- Question in one sentence. Unit of analysis in one sentence.
- Expected number of informative observations. If small, add repeats or units before spending.
- Exposure defined and measured per unit. Primary analysis on exposed units. Placebo set reported.
- Unit selection rule written down, based on exposure and difficulty proxies, never on arm outcomes.
- Every arm difference listed and classified: treatment, removed confound, or measured confound.
- At least two mechanistic measures per hypothesis.
- Arms paired in time and machine. Duration and timeouts recorded and shown.
- Cost model checked against raw provider usage for the unusual responses.
- Headline: one supported result. Then 1–3 charts and five short paragraphs. Put intervals and informative counts in Discussion; keep static HTML current.
- Freeze hashes, smoke test both arms, then do not touch anything until the audit passes.