How to tune an agent harness

Follow a code change through implementation, review, repair and acceptance. Review the revised script · Read the article.

Get useful work finished

01 / 08

A guide to choosing settings for coding agents.

Scroll across the diagram to inspect it.

Get useful work finishedChoose agent settings by the work that reaches acceptance.Your coding agentsUseful work, finishedA change that passes its acceptance checksOur starting point0h1h2h3h4h5hAllowance exhausted at about 1hTwo weeks of evidence: work, review, verification and repair
Allowance incident: operator testimony. Settings and a cache-reporting issue prompted the audit; their quota impact was not measured.
Narration · scene 1

If you run coding agents, this video shows which settings you can control, and how to judge those choices by the work that actually gets finished. We started after using a five-hour subscription allowance in about an hour. Over two weeks, we examined where the tokens went, including review, verification and rework.

Follow the whole task

02 / 08

The first implementation is one step in the journey.

Scroll across the diagram to inspect it.

Follow the whole taskA task consumes capacity through implementation, review, repair and verification.ImplementReviewVerifyAcceptRepair, then review againJudge the complete path to acceptance
Illustrative workflow. The repair path is conceptual, not a reconstruction of the measured case.
Narration · scene 2

Follow a small code change. An agent implements it, another reviews it, and checks verify that it meets the requirements. If review finds a problem, the change goes back for repair, then through review again. That whole journey consumes capacity. A cheaper first attempt can create more work later.

Choose the pair per task

03 / 08

Model strength and effort are one task-level choice.

Scroll across the diagram to inspect it.

Choose the pair per taskModel strength and effort should be chosen together for the task.Task requirementsModel + effortBounded changeStarting hypothesis: capable mid-tier + medium effortDifficult change / substantive reviewConsider stronger model + higher effort
Starting hypotheses, where supported. Preserve independent acceptance checks when changing the pair.
Narration · scene 3

Your first control is the model and its effort setting: where supported, how much reasoning it spends on the task. For a bounded change, a capable mid-tier model at medium effort is a starting hypothesis. A difficult change or substantive review may need a stronger model and higher effort. Let the task and its risks decide.

Carry what the task needs

04 / 08

Context and recovery affect the work that comes back.

Scroll across the diagram to inspect it.

Carry what the task needsContext and recovery settings can save capacity or create rework.ContextInstructions + history + evidenceToo muchRepeated input uses capacityToo littleLost decisions can create rework~200K: a starting hypothesisPreserve decisions at handoff · Reserve capacity to finish
The ~200K starting window is a hypothesis for bounded maintenance, not a measured optimum.
Narration · scene 4

Next, control context: the instructions, history and evidence an agent carries into its next call. Too much can consume capacity; too little can force it to rediscover decisions or repeat mistakes. For bounded maintenance, try a window around two hundred thousand tokens, preserve decisions at handoff, and leave capacity for review and repair. Then measure the result.

The later steps add up

05 / 08

One measured case makes the accounting problem visible.

Scroll across the diagram to inspect it.

The later steps add upIn one accepted item, 53% of identifiable input came after implementation.ImplementationReviewVerification53% after implementationExcluded: unlinked rework + shared overheadMeasure through acceptance
One accepted item; identifiable child streams only. Input includes cache reads. Unlinked rework and shared overhead are excluded.
Narration · scene 5

Why include the later steps? In one accepted item from our Assay audit, review and verification accounted for fifty-three percent of the identifiable input. That count excluded unlinked rework and shared overhead, so it was not a complete delivery cost. The useful question is how much capacity it takes to reach an accepted result.

Keep the evidence connected

06 / 08

Every attempt belongs with the task and revision.

Scroll across the diagram to inspect it.

Keep the evidence connectedEvidence must join every attempt and its outcome to the task.Task + revisionActual model · Effort · Context · Acceptance checksAttempt 1FailedAttempt 2RepairAttempt 3AcceptedCash committedCapacity consumedKeep human correction, elapsed time and unfinished work visible
Illustrative record. Subscription cash allocation and usage are different measures; the article explains their limits.
Narration · scene 6

To answer that, give each task and revision a record. Capture its actual model, effort, context policy and acceptance checks. Attach every attempt, including failures, repairs, human corrections and elapsed time. Keep unfinished work visible too. On subscriptions, track the cash you commit separately from the capacity you consume. Fewer tokens alone do not establish a cheaper completed outcome.

Apply settings independently

07 / 08

The coordinator and task agent do different jobs.

Scroll across the diagram to inspect it.

Apply settings independentlyThe coordinator and task agent need independently resolved settings.cellctl: launch policyResolve provider, model and effortCoordinating loopIts own model + effortTask agentIts own model + effortTask brief → risks, requirements and acceptance checks
Policy principle and launcher example. Verify actual supported settings and installed behavior; learned routing remains proposed in the article.
Narration · scene 7

Once you have a policy, check what actually launches. Our launcher, cellctl, resolves provider, model and effort for role sessions. The coordinator assigns work; the task agent carries it out. They need separate settings. A low-effort coordinator should not force a difficult task to the same effort. The task brief supplies risks and acceptance criteria for that choice.

Start with one workflow

08 / 08

Change a control, then follow the outcome.

Scroll across the diagram to inspect it.

Start with one workflowCompare complete outcomes under fixed acceptance checks.Same taskSame starting statePolicy A / Policy BSame acceptance checksAccepted work · Repairs · Time · Human correctionCapture one complete taskStarting settings and comparison plan in the article
Controlled comparison proposed; no winning policy has been established by this observational audit.
Narration · scene 8

For your next trial, choose one representative workflow. Run two settings policies from the same starting state, with the same acceptance checks. Compare accepted work, repairs, elapsed time and human correction. Keep a change when the complete outcome improves within your available capacity. The article includes starting settings, the subscription report and a comparison plan. Start by capturing one complete task.