Which compaction works on long coding tasks?
Encrypted compaction uses 22% fewer tokens than text summaries.
GPT-6 Luna · Terminal-Bench 2.1
Quality and token use
Background
Long coding sessions eventually fill the context window. Textual compaction asks Pi to write a summary. Encrypted compaction keeps OpenAI’s compacted state. Jev pruning scores older tool calls and results, then keeps, shortens, or drops them without rewriting the conversation. We wanted to know which saves work without losing the task.
Methodology
We ran 24 selected Terminal-Bench 2.1 tasks four times per method with GPT-6 Luna at high reasoning. All methods start compacting at 50,000 tokens. Textual and encrypted runs started in pairs; Jev ran later on the same tasks. Pass rates use the 12 tasks that compacted often in the original pair. Resource averages use all scored runs, including runs that never compacted.
Results
Encrypted compaction used 0.91 million GPT-6 tokens per run versus 1.18 million for text summaries, and cost 10% less. Pass rates were 53.5% encrypted, 56.9% textual, and 58.3% Jev. Jev used the most GPT-6 tokens and cost the most per scored run. The resource saving is the useful result, not a new quality winner.
Discussion
The pass-rate differences are too small to call: encrypted minus textual was −3.5 points, with a 95% interval of −15.3 to +6.9. Jev minus textual was +1.4 points (−12.5 to +13.9). Jev’s later run also leaves time and machine load as possible confounds. Errors are excluded, including one Jev run. A separate audit replay passed after excluding synthetic preflight records; the original flagged audit and raw logs remain available.
Conclusion
Use encrypted compaction as the first option to try when token use is the problem. It saved context and money here, but this run does not prove equal quality. Jev pruning needs a stronger quality or cost result to justify the extra moving parts.
Data and run details
The earlier paired study covered 12 compacting tasks. Jev ran on those tasks later, not alongside the original arms.
Complete · 192/192 earlier runs · about $9.84 spent
Three-way result
Jev runs finished · follow-up audit passed · 96/96 runs · 95 scored · 1 error · $4.81 spent (includes $0.06 Jev)
Follow-up audit passed. Jev vs textual: +1.4 pp (95% interval −12.5 to 13.9 pp); Jev vs encrypted: +4.9 pp (−6.3 to 17.4 pp). Historical comparison, not a concurrent pair.
More measures
The Jev leg ran later on the same 24 tasks × 4 repeats, not alongside the two earlier legs. Task-matched differences have a task-cluster 95% interval, but changes in time and machine load can confound them. One Jev run was unscored. The original audit flagged three synthetic preflight records; a separate replay excluding those records passed without changing trial data. Audit review.
Method, full table and task results
| Method | Pass rate on baseline-compacting tasks | Tokens / run | Cost / run | Median time |
|---|
On narrow screens, swipe the table to see tokens, cost and time.
Jev 1.13 scores each older tool call and result using fast-jev-compaction (pinned source). It sees the conversation with tool outputs omitted, keeps the six newest messages, and never rewrites assistant or user text. It either keeps a call and result, truncates its result to 300 characters, or drops the pair. Its key stays outside the task container. Failures do not silently switch to another method.
| Task | Group (from earlier run) | Textual | Encrypted | Jev | Jev compactions |
|---|
Jev results JSONJev protocolJev source and auditAudit review
Checks
Other task groups and paired outcomes
Mean pass rate across tasks. “Never compacted” checks ordinary run-to-run noise: both methods use the same setup there. Groups can change until all runs finish.
| Task group | Textual | Encrypted | Difference |
|---|
Earlier two-leg method
We ran 24 long Terminal-Bench 2.1 coding tasks with GPT-6 Luna at high reasoning, four times per method in fresh containers. Each pair starts together under similar machine load.
We chose twelve tasks that compacted in version 1 and twelve unused long tasks. We did not pick them by score. This is a selected subset, not an official leaderboard result.
Design, scoring and cost
Version 1 ran each task once and had only four discordant pairs, including two that never compacted. Version 2 repeats tasks, uses the frequently compacting group for the main estimate, and checks tasks that never compacted separately. We also kept Pi’s developer message in both methods after compaction and removed a finish-and-verify instruction from the text-summary arm.
We average each task’s pass-rate difference, then resample whole tasks for a 95% interval (20,000 resamples, fixed seed). Repeats of one task are not independent. The verifier scores the container after the agent stops or times out, so a timed-out run can still pass. Infrastructure errors are excluded and listed. A compaction counts only when another agent request follows it.
Costs use billed tokens at OpenAI’s published standard rates, including text-summary calls. One attempt per run, no score-dependent changes. Dataset license. See the methods note for the full design.
Limits
These long, selected tasks cannot establish a general coding benefit. Jev ran after the original paired study, so task matching does not remove time or machine-load effects. One Jev trial was unscored.
Explore the data
Results by task
Passes per scored run. Compaction rate pools both methods. A clock marks timed-out runs, which can still pass verification.
| Task | Set | Compaction rate | Text summary | Encrypted | Outcome |
|---|
Inspect the run
Results JSONProtocolSource and auditSummary JSONMethods noteVersion 1