Which compaction works on long coding tasks?

Encrypted compaction uses 22% fewer tokens than text summaries.

GPT-6 Luna · Terminal-Bench 2.1

Quality and token use

Background

Long coding sessions eventually fill the context window. Textual compaction asks Pi to write a summary. Encrypted compaction keeps OpenAI’s compacted state. Jev pruning scores older tool calls and results, then keeps, shortens, or drops them without rewriting the conversation. We wanted to know which saves work without losing the task.

Methodology

We ran 24 selected Terminal-Bench 2.1 tasks four times per method with GPT-6 Luna at high reasoning. All methods start compacting at 50,000 tokens. Textual and encrypted runs started in pairs; Jev ran later on the same tasks. Pass rates use the 12 tasks that compacted often in the original pair. Resource averages use all scored runs, including runs that never compacted.

Results

Encrypted compaction used 0.91 million GPT-6 tokens per run versus 1.18 million for text summaries, and cost 10% less. Pass rates were 53.5% encrypted, 56.9% textual, and 58.3% Jev. Jev used the most GPT-6 tokens and cost the most per scored run. The resource saving is the useful result, not a new quality winner.

Discussion

The pass-rate differences are too small to call: encrypted minus textual was −3.5 points, with a 95% interval of −15.3 to +6.9. Jev minus textual was +1.4 points (−12.5 to +13.9). Jev’s later run also leaves time and machine load as possible confounds. Errors are excluded, including one Jev run. A separate audit replay passed after excluding synthetic preflight records; the original flagged audit and raw logs remain available.

Conclusion

Use encrypted compaction as the first option to try when token use is the problem. It saved context and money here, but this run does not prove equal quality. Jev pruning needs a stronger quality or cost result to justify the extra moving parts.

Data and run details

The earlier paired study covered 12 compacting tasks. Jev ran on those tasks later, not alongside the original arms.

Complete · 192/192 earlier runs · about $9.84 spent

Three-way result

Jev runs finished · follow-up audit passed · 96/96 runs · 95 scored · 1 error · $4.81 spent (includes $0.06 Jev)

Follow-up audit passed. Jev vs textual: +1.4 pp (95% interval −12.5 to 13.9 pp); Jev vs encrypted: +4.9 pp (−6.3 to 17.4 pp). Historical comparison, not a concurrent pair.

More measures

The Jev leg ran later on the same 24 tasks × 4 repeats, not alongside the two earlier legs. Task-matched differences have a task-cluster 95% interval, but changes in time and machine load can confound them. One Jev run was unscored. The original audit flagged three synthetic preflight records; a separate replay excluding those records passed without changing trial data. Audit review.

Method, full table and task results
MethodPass rate on baseline-compacting tasksTokens / runCost / runMedian time

On narrow screens, swipe the table to see tokens, cost and time.

Jev 1.13 scores each older tool call and result using fast-jev-compaction (pinned source). It sees the conversation with tool outputs omitted, keeps the six newest messages, and never rewrites assistant or user text. It either keeps a call and result, truncates its result to 300 characters, or drops the pair. Its key stays outside the task container. Failures do not silently switch to another method.

TaskGroup (from earlier run)TextualEncryptedJevJev compactions

Jev results JSONJev protocolJev source and auditAudit review

Checks

Other task groups and paired outcomes

Mean pass rate across tasks. “Never compacted” checks ordinary run-to-run noise: both methods use the same setup there. Groups can change until all runs finish.

Task groupTextualEncryptedDifference

Earlier two-leg method

We ran 24 long Terminal-Bench 2.1 coding tasks with GPT-6 Luna at high reasoning, four times per method in fresh containers. Each pair starts together under similar machine load.

We chose twelve tasks that compacted in version 1 and twelve unused long tasks. We did not pick them by score. This is a selected subset, not an official leaderboard result.

Design, scoring and cost

Version 1 ran each task once and had only four discordant pairs, including two that never compacted. Version 2 repeats tasks, uses the frequently compacting group for the main estimate, and checks tasks that never compacted separately. We also kept Pi’s developer message in both methods after compaction and removed a finish-and-verify instruction from the text-summary arm.

We average each task’s pass-rate difference, then resample whole tasks for a 95% interval (20,000 resamples, fixed seed). Repeats of one task are not independent. The verifier scores the container after the agent stops or times out, so a timed-out run can still pass. Infrastructure errors are excluded and listed. A compaction counts only when another agent request follows it.

Costs use billed tokens at OpenAI’s published standard rates, including text-summary calls. One attempt per run, no score-dependent changes. Dataset license. See the methods note for the full design.

Limits

These long, selected tasks cannot establish a general coding benefit. Jev ran after the original paired study, so task matching does not remove time or machine-load effects. One Jev trial was unscored.

Explore the data

Results by task

Passes per scored run. Compaction rate pools both methods. A clock marks timed-out runs, which can still pass verification.

TaskSetCompaction rateText summaryEncryptedOutcome

Inspect the run

Results JSONProtocolSource and auditSummary JSONMethods noteVersion 1