Browse docs

Ablation Protocol — does the kernel improve output, or only token count?

Purpose: Pre-register the experiment that answers the one question the token benchmarks cannot: whether an agent running coding-os produces better work, not just cheaper retrieval. Registering the arms, the metrics, and the scoring rule before any arm executes is what stops the result being chosen after the fact. Read when: running, changing, or citing the ablation. Skip when: the change touches only token accounting.

Nav: Section Index | Docs Index

Status

Registered, not yet run. No arm has executed, and therefore no result appears here or anywhere else. When a run happens, its numbers land in this file and nowhere earlier in the pipeline.

Revised 2026-08-15: the substrate moved from this repo's mined task history to SWE-bench Verified, after measuring that the mined acceptance commands cannot discriminate. Nothing had run under the old design, so nothing is retracted — this is a pre-registration replacing a pre-registration.

That order is the point. A published quality claim with no executed run is the same defect as a savings claim measured against a strawman — the thing this whole line of work exists to stop.

Why the existing benchmarks cannot answer it

third-party-token-bench.md measures the cost of one retrieval call against the cost of grepping. context-budget.md measures the toll the instruction layer charges. Neither observes whether the delivered change was correct. An agent can spend fewer tokens and produce worse code; the token ledger would call that a win.

The four arms

Each arm differs only in the instruction and retrieval layer. Model, model version, sampling settings, tool allow-list, container image, and starting commit are identical across arms.

Arm Instruction layer Graph tools What it isolates
raw none — stock mini-SWE-agent no the floor
graph none yes retrieval alone
rules full CLAUDE.md + .claude/rules/ no instruction discipline alone
full full yes the shipped product

raw vs graph isolates the graph. raw vs rules isolates the rules. If full does not beat both single-lever arms, the levers interfere and that is a finding worth publishing.

raw is mini-SWE-agent on its published bash-only configuration, unmodified — the same control the Verified leaderboard uses to compare language models, whose maintainers state they do not tune it for score. Adopting someone else's floor is the point: a floor this project defines is a floor this project can flatter.

What may and may not be published

Verified is saturated and contamination-prone; an absolute score on it ranks nothing. The only defensible statistic is the between-arm delta on one model, run in one window, with its spread. "coding-os scores X%" is not a claim this protocol can support and must not be made from its output. "With model M on date D, full resolved N more instances than raw, 95% CI [a, b]" is.

Cost probe before pilot

The pilot is not funded on an estimate. A guessed price here is the same defect as a guessed savings percentage — the thing this whole line of work exists to stop.

Pricing it takes one run, not ten. A single raw-arm run emits real input / cache-write / cache-read / output counts and real wall-clock in its trajectory; ×300 is the pilot figure. Two things make the extra nine runs waste:

  • Skip the grader. The tokens are in the trajectory whether or not the patch is correct, so evaluation buys container time and zero cost information. It answers completion rate, which is a pilot metric, not a probe metric.
  • Pick an instance with a native image for your architecture. On arm64, swebench/sweb.eval.arm64.* runs natively; an amd64-only instance is QEMU emulated and its wall-clock measures the emulator, not the agent.

Run src/scripts/ablation_probe.py --preflight first — it exits non-zero and names every missing prerequisite rather than half-starting.

Measured cost

Not yet measured — the probe has never run. Preflight on the maintainer's machine, 2026-08-16:

Prerequisite State
container runtime blocked — Docker 29.5.2 has 1.5 GiB free of 4.8 GiB
dataset 500 instances, reachable anonymously over the datasets-server API
control agent missinguv pip install mini-swe-agent clears it
model credential missing — the hard one

(Not uvx: it runs the package in a throwaway environment and installs nothing into the interpreter the check imports from, so the check would still fail.)

The credential is the blocker no local command fixes. mini-swe-agent routes every model call through litellm, which reads a provider key from the environment; a Claude Code session's own OAuth cannot be lent to it. So the probe needs a key exported deliberately, for a model both arms share.

The memory line is free, not total: the VM is 4.8 GiB but unrelated stacks hold ~3.3 GiB of it, and a SWE-bench image unpacks to about 1 GB and then runs a test suite inside itself. An earlier revision of the check read MemTotal and printed [OK] 4.8 GiB on this machine — a green on a box that cannot run the probe. It now subtracts what docker stats reports as committed and compares the remainder against a 2 GiB floor. That floor is an assumption; the first real run replaces it with an observation.

Pilot before fleet

Run 50 instances × 2 arms (raw, full) × 3 repetitions = 300 runs after the probe has priced it. The middle arms only earn their cost if the endpoints separate: if fullraw sits inside the noise at n=50, decomposing that null into graph and rules will not find a signal either, and the null is the result. Publish it either way — a pre-registration whose negative outcome goes unpublished is not a pre-registration.

Metrics — fixed before the first run

Primary:

  1. Completion rate — the instance's FAIL_TO_PASS tests pass and its PASS_TO_PASS tests still pass, evaluated by the SWE-bench harness in a container the agent never saw. Both halves are required: a patch that fixes the issue by breaking something else is a failure, not a partial credit.
  2. Weighted cost per taskinput + 1.25×cache_write + 0.1×cache_read + 5×output, the same weighting cos doctor --tokens uses. Read from the transcript, not estimated.

Secondary, reported always, never used to pick a winner:

  1. Tokens per task (raw, unweighted).
  2. Wall-clock per task.
  3. Turns per task.
  4. Quality per dollar — completion rate ÷ weighted cost.
  5. Completions per million tokens.

Rules that bind the run:

  • N = 3 runs per (task, arm). Report the median and the full spread; a single run of a stochastic agent is an anecdote.
  • No metric may be added after the first arm executes. Adding one is a new pre-registration and a new run.
  • The upstream test set is the scorer. No model-judged quality score, because a model judging an agent's output is a second uncontrolled variable — and no scorer written by this project, because a scorer written by the thing under test is the strawman baseline wearing a different hat.
  • Failures count. A task the agent abandons is a completion-rate 0, not an excluded outlier.

The task set — SWE-bench Verified, not this repo's history

Substrate: SWE-bench Verified — 500 human-filtered instances, each carrying FAIL_TO_PASS and PASS_TO_PASS test sets written by the upstream maintainers in the same PR that fixed the issue, and each running in its own container. Four properties this repo's own history cannot supply:

Requirement Why mining this repo fails it
The scorer must discriminate see the measurement below
Task author ≠ instruction author the same person wrote the rules and the tasks
Arms must not fight the git guards replaying a commit needs a checkout branch-guard blocks
A published floor mini-SWE-agent is the community's accepted control

Why the home-grown set was abandoned — the measurement

src/scripts/eval_taskset.py mined 148 candidates from 978 closed tasks on 2026-08-15. Those 148 carry only 58 distinct acceptance commands, and the head of the distribution is fatal:

Count Command
27 make ui-build
15 uv run pytest tests/test_cli.py -q
11 cos init
10 cos doctor
7 make regen-rules

70 of 148 (47%) are suite-level or idempotent tool invocations that pass at the starting commit, so they cannot separate a solved task from an untouched one. The remainder mostly name a test that the task's own commit added — so the scorer reduces to "did the agent guess the author's test-file name", which is not a measure of correctness either. No amount of validation rescues a set whose acceptance criteria are not independent of the solution.

The miner and its output stay in the repo as evidence for that conclusion, not as an input to a run.

The one finding worth keeping from it

A Then clause that names a command is strictly more useful than one describing a feeling of doneness — the "verify by executing" rule applied to the task template itself. It will not retroactively fix 978 closed tasks, and after this measurement it is a task-hygiene improvement, not an eval strategy.

Execution environment (non-negotiable)

Every arm runs against a throwaway clone, never this working tree. Replaying a historical commit requires git checkout <sha>, which branch-guard.sh BLOCKs in trunk mode, and Rule 21 forbids worktrees — a protocol that ignores that is a protocol that cannot run. SWE-bench's per-instance containers satisfy this by construction; any local arm must clone to a temp dir and be deleted after.

Threats to validity, stated up front

  • Verified is contaminated and saturated. Its instances predate most current models' training cutoffs. This is survivable only because every arm shares the contamination equally — it inflates the floor, not the delta — and it is exactly why no absolute score may be quoted.
  • Verified is Python, and mostly library code. A gain here does not transfer to the WordPress or Next.js consumers this kernel also ships to. Reporting it as a general result would repeat the strawman-baseline error in a new place.
  • The instruction layer was written for this repo, not for django or sympy. Stack rules that carry the most specific guidance will be absent or irrelevant, so the measured delta is a lower bound on a matched-stack project and says nothing about the upper one.
  • Cost of a full run is real — 4 arms × N instances × 3 repetitions, each a container plus a model session. Sample size will be reported, and an underpowered result will be labelled directional.
  • n=50 detects only a large effect. A null at pilot scale is "no effect this design can see", not "no effect".

See also