Does encrypted compaction help in coding tasks?

Encrypted compaction cut token use by 29%.

GPT-6 Luna · high reasoning · Terminal-Bench 2.1 · version 1, superseded by version 2

Pass rate, cost, and tokens

Background

When a coding agent runs out of context, should it keep a text summary or the provider’s encrypted state? This first pass compared both on GPT-6 Luna. It found a useful resource difference, but the quality result needed a better experiment.

Methodology

Pi 0.80.10 ran 20 selected Terminal-Bench 2.1 tasks once per method in fresh containers, with high reasoning and a 50,000-token compaction threshold. Tasks came from instructions and resource metadata, not model scores. Resource means use tasks with both methods scored; cost includes task and summary calls at OpenAI’s published rates.

Results

Encrypted compaction used 1.03 million tokens per task versus 1.45 million for text summaries, a 29% cut. Cost fell 15%, from $0.0449 to $0.0382 per task. Encrypted passed 80% of tasks versus textual’s 70%, but only four pairs had different outcomes: three encrypted-only wins and one textual-only win.

Discussion

Eight tasks never compacted, and two of the four differing outcomes came from those untreated tasks. The pass-rate gap had a 95% interval of −10 to +30 points. We also found setup differences: the encrypted method dropped Pi’s developer message, while the text method got an extra finish-and-verify instruction. Those confounds make this a poor test of quality.

Conclusion

Keep the resource saving as a reason to test encrypted compaction, not the 80% score as proof it is better at coding. Version 2 adds repeats, fixes the prompt differences, and separates tasks that actually compact. Read that run for the current comparison.

Data and run details

Partial data

Results by task

Only compactions followed by another model request count.

TaskPi textualOpenAI encryptedOutcome

What we measured

Pi 0.80.10 runs 20 tasks with GPT-6 Luna at high reasoning. At 50,000 tokens, one arm keeps a text summary; the other keeps OpenAI’s encrypted state. Each task runs once per arm in a fresh container.

The original seven tasks plus thirteen longer implementation, debugging, scientific-computing and build tasks. Selection uses task instructions and resource metadata, not GPT-6 results.

Method and cost

Topline scores and resource means use tasks with both arms scored. Each arm also shows all its completed results. Compaction counts distinguish observed events from events followed by another agent request.

Costs include task and summary calls at OpenAI’s published standard rates. Cached input counts once in token totals. Spend includes preflight calls; no GPU rental is needed. Errors are separate from verifier failures.

Original task timeouts and resource limits, four concurrent attempts, one run per arm/task. Task-paired uncertainty is reported after completion. This selected public subset is not an official leaderboard evaluation. Tasks: Terminal-Bench 2.1 (license).

Inspect the run

Results JSONProtocolSource and auditRun summarySummary JSON