Can Exo improve itself during a task?

Exo made zero self-edits and passed zero tasks.

GPT-5.6 Luna · xhigh reasoning · Terminal-Bench 4.0

What happened to the six tasks?

Loading task outcomes…

Three scored failures, one error, two blocked. No source edits or rebuilds.

Background

Exo can edit its own source, rebuild, and pick up the same conversation. That sounds useful when an agent gets stuck with the wrong tools. We tested whether giving it that freedom would actually lead to self-improvement during coding tasks.

Methodology

We gave GPT-5.6 Luna at xhigh reasoning six Terminal-Bench 4.0 development tasks, one attempt each. Every task got a fresh writable Exo and a prompt to read SELF.md and improve itself when useful. Changes did not carry between tasks. We tracked source edits, tool changes, rebuilds, and task verification separately, with fixed task resources and a one-hour limit.

Results

Exo made no source edits and no rebuilds during the task run. All three scored tasks failed. The fourth hit an error, which stopped the run and left two tasks blocked. The setup check had already shown that editing, rebuilding, and resuming worked, so the missing piece was using that capability on the tasks.

Discussion

This is a failed attempt to trigger useful self-editing, not proof that self-editing agents cannot work. The tasks were seen development cases, not a holdout, and the error cut the run short. Errors and blocked tasks are not scored failures. There was no fixed-runtime control, so we cannot measure a benefit from source access alone.

Conclusion

Do not count writable source as self-improvement. First get the agent to make and use a relevant change on a real task; then compare it with the same agent running a fixed runtime. This setup never cleared that first bar.

Data and run details

Loading run status…

Results by task

Errors and blocked tasks are not scored failures. Scroll for usage and rebuilds.

TaskStatusRewardCost¹TimeCallsTokens (in / cached / out)Source changesRebuilds (ok / failed / queued)Notes

What we measured

Each task gets a fresh, writable Exo. It can edit source, rebuild, and resume the same conversation. Changes do not carry between tasks.

Six seen development tasks, not a holdout. Source access alone is not improvement; tool edits, source edits, and rebuilds are counted separately.

Runtime and task limits

Exo uses its practical profile (shell, inspect_tools, manage_tool, rebuild_and_restart_exo). The guardian builds its writable clone; the driver resumes it after a rebuild.

Terminal-Bench images, tasks, resources, and verifiers stay fixed. Harbor attaches the task sandbox; the driver asks Exo to read SELF.md and improve when useful.

The driver enforces a one-hour task limit, an estimated $2 per-task cap, and two provider retries. An infrastructure or provider failure stops the remaining tasks. A scored task failure does not.

This run uses xhigh reasoning; earlier max-reasoning trials are not pooled with it.

Setup checks and cost

Setup fixed the restart path, guardian cleanup, and scheduler secret store. Keys stayed outside the sandbox.

The first setup check rebuilt but could not resume through macOS Keychain ($0.0039). The secret-store fix came before benchmark tasks.

¹ Costs use provider-reported usage where available, otherwise Luna pricing ($0.20 input, $0.02 cached input, $1.20 output per million tokens). In-flight calls are not counted until usage arrives, so cost caps can overshoot by a call. Setup spend is separate.

Inspect the run

Results JSONExo source