This article explains the controls an operator can use to keep an agent harness productive within its usage limits: model choice, reasoning effort, context size, concurrency and recovery. It gives a starting configuration for routine software work, then shows how to adjust it using evidence from the whole workflow. Our example is Assay, the brief-driven process we use to coordinate implementation, review and verification, and two weeks of its session records.
The immediate reason for looking was less abstract. I exhausted the allowance for a five-hour subscription window within an hour. The investigation involved badly chosen context limits, the wrong model settings and a Claude cache-usage reporting bug. We needed to understand which controls were ours, how they interacted and what useful work the allowance actually bought. The logs let us examine those settings, although they cannot apportion the exhausted allowance among the causes.
start with a policy you can change
For a harness doing mostly bounded software maintenance, I would start with the settings below. These are proposed starting points, not configurations that our audit has proved optimal. “Mid” and “strong” mean capability tiers in your own provider catalog; effort names must map to settings that the selected model supports.
| Responsibility | Starting model and effort | Starting context policy |
|---|---|---|
| Main loop scheduling known work | Capable mid-tier model, low effort; send difficult judgments to a specialist | About 200K compaction target, durable queue and handoff state |
| Worker implementing a clear, bounded brief | Mid-tier model, medium effort where supported | One brief per child; about 200K target, retrieve relevant material |
| Reviewer judging a substantive code change | Strong model, high effort | Focused diff, requirements and relevant dependencies; about 200K target |
| Verifier checking acceptance evidence | Deterministic checks first; capable model, medium effort for interpretation | Fresh task-specific session; about 150–200K target |
An ambiguous brief, security-sensitive change or difficult systems problem should raise the model and effort before execution. A coordinator making those judgments needs the corresponding capability itself. Low effort is a starting point for routine dispatch, not permission to weaken every decision made by the main loop.
Start with two concurrent workers if you have no throughput baseline, and leave room for review and verification. Allow one targeted repair for a clear failure; if the same cause returns, stop and change the approach. These are deliberately simple policies to measure. A context target triggers summarization or handoff; it is neither a guarantee that all requests fit below that number nor a reason to discard necessary evidence.
Starting hypotheses for bounded software maintenance. High-risk work and different harnesses require different settings.
find out what “typical” means in your harness
A typical workload is a mix of tasks and responsibilities. In Assay, a standing loop dispatches work; children implement, review or verify specific briefs. Treating them as one population hides the distinction we most need to tune.
Across September 13–26, our four local Claude-profile trees contained 244,637 deduplicated model-message receipts across 5,278 streams. They reported 48.50 billion input tokens, about 98.3% of them cache reads. This is repeated context consumption, not 48.50 billion tokens of new material. [2]
The load looked different at the two levels:
Call-weighted medians, rounded. These describe the input presented to models, including cache reads; they do not establish how much context the tasks required. [2]
For implementation, the median main-loop call carried 441K input tokens; the median child call carried 133K. Review was 480K versus 93K; verification was 386K versus 85K. Long, active loops dominate a call-weighted measure. Giving each implementation main stream equal weight instead produces a median of its per-stream medians of 52K. Both are valid measurements; they answer different questions. Review and verification remain large when streams receive equal weight: 471K and 318K for main loops, versus 90K and 80K for their children.
The child tails matter too: implementation's 90th percentile was 271K, compared with 141K for review and 131K for verification. A 200K target is therefore a useful experiment for this workload, with an explicit test for lost constraints and extra repair. It is not a conclusion that every worker fits comfortably inside it.
Use your own last two weeks to group tasks by phase, risk, size and acceptance criteria. Separate parents from children; retain failures and unfinished work. Measure the median and tail for each group, then inspect examples from both. Your baseline should describe what the harness regularly does and where it struggles, rather than whatever its busiest session happens to consume.
include the work that comes back
The route to a completed outcome includes scoping, dispatch, implementation, review, repair, verification and sometimes reopening after a regression. Rework belongs to that same account. A cheaper first attempt can become an expensive completed change.
One accepted Assay item illustrates why the worker alone is an inadequate denominator. Its identifiable worker stream recorded 10.87 million input tokens. Review added 10.97 million and verification 1.45 million. Review and verification accounted for 53% of the 23.28 million observed direct input. Shared coordination and unlinked attempts were outside this reconstruction; it is one example, not a fleet average or a measured rework total. [1]
The repair path is conceptual. The three stage totals are measured; no numerical rework saving is claimed.
Every lever can alter that return path. Too little effort can miss an edge case. A smaller model can misunderstand a dependency. Tight compaction can lose a constraint. Excessive parallelism can create conflicting edits. Each apparent saving needs to survive review, repair and verification. Conversely, a clearer brief or a narrower task can make a smaller model sufficient and reduce the corrections required.
With subscriptions, keep cash and capacity separate. An included request may add nothing to the invoice while consuming allowance needed to finish the next job. Count subscription expense per accepted outcome, provider quota consumed, time to acceptance and human correction. Dividing a fixed fee by more tokens can reward pointless rereading.
are you capturing the evidence to decide?
You need to join the work, its attempts and its verdict. Record a stable work identifier, the revision being attempted, parent and child sessions, the resolved model and effort, context policy, usage, repair reason and independent acceptance. Keep rejected and abandoned attempts attached. Otherwise a weak route can look economical simply because its failures disappeared from the denominator.
Then ask a concrete question: did lowering this worker's effort reduce total consumption through acceptance without increasing serious findings or human correction? Did bounding the main loop reduce repeated context while preserving its queue and obligations? A lower token count alone cannot answer either question.
Our retrospective records support individual reconstructions and a workload profile. They do not yet support a complete fleet-wide cost per accepted outcome or a causal ranking of models. That gap tells us what to instrument next. Ironclad's conference discussion of trusted throughput makes the same broader move: generation must be assessed alongside review and downstream results. [3]
make the chosen policy repeatable
This is where cellctl enters. It launches our role sessions and resolves their provider, model and effort policy. The operator needs separate settings for the standing main loop and each task child. A dispatcher can run at low effort while a difficult child uses a stronger model at high effort; its context lifetime is a separate decision again. [6]
That separation must survive the launch path. The merged effort-inheritance correction removes an ambient Claude effort override so child-specific effort can take effect. The merged context change sets a 200K default for host Claude launches in the applicable cell types; it does not set that default for scrubbed or container cells, and a non-empty override takes precedence. Both the context-default change and the effort-inheritance correction merged on September 27, 2026. This establishes source status, not which version any operator has installed. A policy file expresses intent; inspecting the launched configuration and available runtime evidence checks whether it happened. [9]
The brief supplies the other half. Assay briefs already carry risk answers, an execution-tier floor, a rationale and an optional budget. They give the loop-runner context for choosing a subagent. The decision should be a model–effort pair for the task, with context, tools and budget alongside it. Our current tier field supplies part of that requirement; it does not by itself enforce a task-level effort floor. [5]
A proposed small decision model could take those brief features and past outcomes as input, then recommend a permitted model–effort pair or abstain. It would advise inside the risk and capability constraints, with independent verification still judging the work. That integration remains proposed.
Brief fields and the launcher exist. Complete enforcement, learned selection and outcome-linked feedback are separate pieces to establish.
A useful first tuning cycle is small: take a representative task class, set an explicit baseline, capture its complete path and change one control. Keep the change when the accepted work improves. The deeper guide below makes that cycle repeatable.
the operator guide: calibrate the controls on your workload
choose model and effort together, for the decision being made
A role is a convenient default. The task supplies the requirement. An easy change in a sensitive subsystem may still demand strong review; a long mechanical transformation may need little reasoning but precise tooling.
For an initial catalog, extend the starting table this way:
| Task class | Candidate starting pair | Reason to raise capability or effort |
|---|---|---|
| Idea generation and design alternatives | Strong model, high effort; bounded exploration | Incompatible constraints, unclear trade-offs or an irreversible decision |
| Issue triage and routine dispatch | Mid-tier model, low effort | Uncertain ownership, conflicting evidence or substantial prioritization judgment |
| Bounded implementation | Mid-tier model, medium effort where available | Ambiguous requirements, unfamiliar architecture, repeated substantive defects |
| Substantive review | Strong model, high effort | Widen context or add a specialist when the relevant system boundary is missing |
| Verification | Code for stable checks; capable model at medium effort for evidence interpretation | Incomplete evidence, subtle failure modes, high-risk acceptance decision |
| Operations changes | Strong model, high effort for diagnosis and change planning | Large impact or uncertainty requires stronger evidence and the applicable human gate |
These are suggested defaults to test, not a ranking obtained from our traces. In particular, a human approval gate does not make operations reasoning easy. Permission and competence solve different problems.
To name the initial models, use the model currently passing your independent checks as the baseline. Pick one permitted lower-cost candidate as the challenger for a bounded task class; keep the baseline for higher-risk work until that class has evidence. Record actual model IDs, not just tier labels.
Choose only supported effort levels and record the actual mapping. “High” across two models is not a standardized quantity of reasoning. Raising effort on a weaker model may not compensate for a missing capability. Compare permitted pairs on representative tasks, and check tool reliability, output quality and latency as well as tokens. Deterministic code belongs in the catalog wherever a stable rule can be checked directly.
bound context without losing the work
Separate persistent instructions, tool definitions, the brief, retrieved evidence, tool output and conversation history. Each needs a different treatment. Trim redundant boot instructions; return smaller tool results; keep long evidence in retrievable files; preserve decisions, unresolved obligations and acceptance criteria across handoff.
Begin around 200K for the standing loops in a workload like ours, provided the model and harness support it. For focused review and verification children, test 150–200K. If a child regularly needs more, inspect the task before changing the target: does it need a larger working set, a better retrieval strategy or a split into independently checkable parts? Large-context variants remain useful when the work actually needs them.
Record compaction and restart costs, then check the next decisions. A successful summary is one that preserves what the agent needs to finish. If forgotten constraints produce repair, the smaller context did not save all the tokens it appeared to save.
Keep stable instructions and tool schemas in a stable prefix where the harness permits. Avoid changing models or providers mid-task without a reason: cache rebuilding, different tool behavior and recovery can outweigh the next request's saving. The DigitalOcean routing talk treats cache stickiness as a routing constraint. [7]
Cache reads and writes require careful interpretation. Reads can refer to a cache created outside the local trace, and zero writes alone do not prove that a counter is false. Claude Code has also fixed a documented usage-reporting bug involving nested cache-creation fields. That establishes a reporting failure mode, not that this bug caused our subscription exhaustion. Check client versions, provider semantics and the provider's own usage meter before turning receipt totals into quota claims. [10]
treat rework as a diagnostic, not just a surcharge
Capture why work returned. “Attempt two” is less useful than “review found an omitted requirement,” “verification exposed an integration defect,” or “the environment prevented the check.” Distinguish necessary iteration, avoidable defects, infrastructure failures and new scope.
| Lever changed | Possible benefit | Rework signal to watch |
|---|---|---|
| Lower model tier or effort | Less expensive first attempt | Repeated missed requirements, invalid tool use, shallow diagnosis |
| Smaller context or earlier handoff | Less repeated input | Forgotten decisions, rediscovery, omitted acceptance checks |
| Smaller briefs | Easier independent completion | Excessive integration and repeated background loading |
| More parallel workers | Shorter implementation queue | Conflicts, stale changes, overloaded review |
| Early escalation | Avoid repeated weak attempts | Unnecessary strong-model use on recoverable routine failures |
| More retries | Recover transient failures | Same failure reproduced without new evidence |
A first repair can be appropriate when the correction is specific. After the same cause returns, change something substantive: model–effort pair, task boundary, missing tool, prerequisite or human decision. Track that intervention. Do not label every failed tool call as model weakness or every review cycle as waste.
Reserve completion capacity using what review and verification actually consume in your workload. Our one reconstructed item's 53% is a warning against ignoring later stages, not a universal reserve percentage. At the queue level, increase worker width only while accepted throughput improves without growing stale work or the correction burden.
build two views of “typical”
Use a work view for decisions: one versioned item, its attempts, rework, acceptance status and elapsed time. Use a load view for capacity: calls by model, effort, phase and parent/child, their input categories, output, context distributions and concurrency. Join them; neither replaces the other.
Two weeks is a useful starting sample if it includes the workload you intend to operate. It is not automatically representative. Identify routine maintenance, larger changes, investigation, operations and exceptional incidents. Record their proportions. Use a later period or held-out tasks to check whether a policy chosen on that sample still works.
Report per-work-item medians and tails, with enough examples to explain outliers. Keep call-weighted distributions for diagnosing repeated context. If output usage or effort is absent, mark it unknown. A dashboard that quietly fills missing values with zero makes a policy look better than the evidence warrants.
count the cost of a completed outcome
Under our assumption that Kimi, GLM, Claude and Codex use their respective maximum subscription offerings, the starting cash account is the actual subscription roster and fees attributable to the measured period. A local profile directory is not proof of a separate paid account. Discounts, account sharing and billing periods need explicit assumptions.
For a period, calculate attributable subscription expense ÷ independently accepted outcomes and show opening and closing work in progress beside it. For a task cohort, attach all attempts through its cutoff, including failed and abandoned items, and report both the accepted count and the unfinished count. Do not take all monthly expense and divide it by a conveniently small, differently timed completion sample.
Model-level allocation is a separate estimate: choose an allocation rule, state its coverage and keep provider-specific quota weights distinct from raw token totals. Allocated dollars are not marginal API charges. Under a fixed subscription, a reduction first buys headroom; it reduces cash spending when it permits a smaller subscription commitment or avoids another purchase.
The accompanying subscription analysis keeps measured usage separate from subscription scenarios. A complete dollar cost per outcome still needs reliable outcome joins and billing coverage. We have enough evidence to demonstrate the method and audit individual work, not enough to invent a reconciled fleet figure.
capture the route at every handoff
A minimal record links work ID and version, attempt ID, parent/child session, responsibility, brief risk, requested and resolved model–effort pair, context policy, tools, start/end times, usage categories, failure reason and acceptance evidence. Add provider reset and quota-delay observations where they affect completion. Preserve the client and policy version so a later audit can explain a change.
Check inheritance explicitly. Claude Code exposes model and effort settings for subagents, but the effective result also depends on overriding settings and the invocation path. [11] Test one inexpensive child before rolling out a changed launcher policy; inspect the available configuration evidence and keep unobservable fields marked unknown. The main loop should not accidentally force its low effort on a hard child, nor should every routine dispatch inherit the child's high effort.
Brief risk and capability floors should constrain the candidate set. The planned selector can then recommend model and effort using task features and historical outcomes. Before automatic use, shadow it to assess coverage, abstention and policy compliance. Shadow recommendations cannot prove that an unexecuted route would succeed. That requires executed comparisons.
Include selector context preparation, inference, latency and maintenance in its cost. A rule such as “bounded low-risk transformation with complete checks: use the baseline pair” may be sufficient. Learned advice earns its place when it improves decisions beyond that baseline. AI21's portfolio work provides useful hypotheses about sequences of models; its reported gains are not results from our harness. [4]
run a controlled comparison, then keep checking
A practical pilot would select 24–30 representative tasks and run two policies from the same starting revisions in isolated workspaces. Freeze prompts, tools, acceptance checks and budgets. Interleave or randomize the order, prevent the arms from seeing each other's outputs and standardize cache conditions. Keep reviewer and verifier policy fixed; blind them to the arm where practical.
Change one control if you want to identify its effect. Comparing two complete policies is legitimate too, but the result belongs to the bundle. Record every attempt, correction, unresolved outcome, quota delay and time to acceptance. Predeclare the quality loss that would disqualify a saving; serious regressions must not vanish inside an average.
Allow roughly three to five working days for instrumentation, a first paired pilot and adjudication, assuming available capacity. A replication plus a week of regression observation suggests ten to fourteen calendar days for a more defensible report. These are planning estimates, not a guarantee of statistical power. Pilot variance and task diversity determine the larger sample needed.
Review usage and queues daily during tuning; review completed cohorts weekly. Recheck after a model, harness, launcher or cache-policy change. Revisit the subscription portfolio when sustained headroom makes a smaller commitment plausible. Keep acceptance criteria anchored in domain expertise: the conference material from Langfuse is a useful reminder that an optimization loop can improve its score while missing what the operator actually values. [8]
The record to preserve is the completed work and what it took to get there: the initial route, the corrections, the evidence and the final verdict. That is what gives the next model–effort choice a sound basis.
the hedge
These are observational records from one operator, with incomplete output receipts and outcome joins. The settings are starting policies to test, not controlled benchmark winners. The later learned selector remains a proposed integration. The report distinguishes subscription assumptions from observed usage and includes failed or missing joins in its limits.
sources and measurement notes
[1] Local reconstruction of one accepted work item, September 23–24, 2026. The public methods appendix gives input totals, time boundaries and acceptance evidence without internal identifiers. Query: sum fresh input, cache creation and cache reads over each of the three identified child streams; do not include output in input. See the methods appendix.
[2] Local extraction, September 13 00:00 through September 27 00:00 UTC (two complete weeks), with the overnight verifier examined separately. Query: deduplicate reported message IDs, exclude synthetic messages, sum input + cache_creation + cache_read; group child records by parent session. Method, known receipt gaps and aggregation distinctions are in the methods appendix.
[3] Ironclad, “From Tokenmaxxing to Trusted Throughput,” AI Engineer World's Fair 2026. Talk. The local transcript-derived research notes informed the summary; no vendor productivity figure is adopted here.
[4] AI21, “Self-Optimizing Agents,” Agentic AI Summit 2026. Talk. Portfolio and execution-strategy discussion; reported benchmark gains are not reproduced in this article.
[5] Assay brief specification, inspected September 27, 2026: risk, gate, exec-tier, exec-tier-why and unit-bearing budget. The fields are defined in the public brief specification.
[6] Assay cellctl role-launch and model-policy resolution source, inspected September 27, 2026. This establishes the launch mechanism; it does not establish complete enforcement or an adaptive routing system across every execution path.
[7] DigitalOcean, “Preferences over Benchmarks: Model Routing for How Teams Actually Build,” Agentic AI Summit 2026. Talk.
[8] Langfuse, “Stop Burning Tokens: Why Self-Improvement Needs Domain Expertise First,” AI Engineer World's Fair 2026. Talk.
[9] Assay launcher status checked September 27, 2026: context defaults, PR 1762 merged; effort inheritance correction, PR 1760 merged. The earlier proposals were closed and replaced. Source status does not establish installation or a release version.
[10] Anthropic, Claude Code v2.1.152 release notes, documenting a fix to cache-write usage reporting. This is not evidence of how our provider charged the incident.
[11] Anthropic, subagent configuration and model configuration, inspected September 27, 2026. Supported settings vary by model and client version.



