# How to tune an agent harness

An operator explainer for people using coding agents. Video: [watch the uploaded explainer](https://youtu.be/bws-W-DujxM). The narration below is unchanged by the article review corrections.

## Runtime

446 words at a planning rate of 140 words/minute: 191.2 seconds before pipeline gaps. A take 10% slower is 210.3 seconds. These are estimates, not measured narration.

## Scenes

| # | id | ~sec (estimate) | words | narration | on-screen text |
|---|---|---:|---:|---|---|
| 1 | orient | 22.29 | 52 | If you run coding agents, this video shows which settings you can control, and how to judge those choices by the work that actually gets finished. We started after using a five-hour subscription allowance in about an hour. Over two weeks, we examined where the tokens went, including review, verification and rework. | Get useful work finished |
| 2 | journey | 21.0 | 49 | Follow a small code change. An agent implements it, another reviews it, and checks verify that it meets the requirements. If review finds a problem, the change goes back for repair, then through review again. That whole journey consumes capacity. A cheaper first attempt can create more work later. | Follow the whole task |
| 3 | pair | 24.0 | 56 | Your first control is the model and its effort setting: where supported, how much reasoning it spends on the task. For a bounded change, a capable mid-tier model at medium effort is a starting hypothesis. A difficult change or substantive review may need a stronger model and higher effort. Let the task and its risks decide. | Choose the pair per task |
| 4 | context | 24.43 | 57 | Next, control context: the instructions, history and evidence an agent carries into its next call. Too much can consume capacity; too little can force it to rediscover decisions or repeat mistakes. For bounded maintenance, try a window around two hundred thousand tokens, preserve decisions at handoff, and leave capacity for review and repair. Then measure the result. | Carry what the task needs |
| 5 | evidence | 23.14 | 54 | Why include the later steps? In one accepted item from our Assay audit, review and verification accounted for fifty-three percent of the identifiable input. That count excluded unlinked rework and shared overhead, so it was not a complete delivery cost. The useful question is how much capacity it takes to reach an accepted result. | The later steps add up |
| 6 | record | 25.29 | 59 | To answer that, give each task and revision a record. Capture its actual model, effort, context policy and acceptance checks. Attach every attempt, including failures, repairs, human corrections and elapsed time. Keep unfinished work visible too. On subscriptions, track the cash you commit separately from the capacity you consume. Fewer tokens alone do not establish a cheaper completed outcome. | Keep the evidence connected |
| 7 | launch | 24.86 | 58 | Once you have a policy, check what actually launches. Our launcher, cellctl, resolves provider, model and effort for role sessions. The coordinator assigns work; the task agent carries it out. They need separate settings. A low-effort coordinator should not force a difficult task to the same effort. The task brief supplies risks and acceptance criteria for that choice. | Apply settings independently |
| 8 | compare | 26.14 | 61 | For your next trial, choose one representative workflow. Run two settings policies from the same starting state, with the same acceptance checks. Compare accepted work, repairs, elapsed time and human correction. Keep a change when the complete outcome improves within your available capacity. The article includes starting settings, the subscription report and a comparison plan. Start by capturing one complete task. | Start with one workflow |

## Description

How to choose model, effort and context settings for coding agents, and assess the complete path to acceptance. Includes review, repair, verification and subscription capacity, using a two-week Assay audit. Starting settings are hypotheses; the audit does not establish a winning policy.

https://assay.guide/blog/how-assay-works/tune-an-agent-harness/

## Recording notes

The first scene establishes the subject and audience before the incident. The workflow and attempt records are illustrative. Scene 5 is one measured case, with exclusions spoken and visible. The article carries the detailed defaults, subscription allocation and proposed learned routing.

There is intentionally no subtitle column: the engine uses the full narration for captions. Its current output is one cue per scene; split and align those cues to the reviewed speech before upload. On-screen headlines are separate from accessible captions. Review cellctl pronunciation and final voice timing. YouTube upload and publication remain separate human steps.
