Subscription capacity and cost allocation

Analysis snapshot · 27 September 2026

What one week of recorded agent work tells us about subscription utilization—and what we still need to measure the cost of a completed outcome.

Scope: September 20 00:00 to September 27 00:00 UTC. Usage combines the four inspected Claude-profile trees with the inspected Codex sessions. The article’s two-week context charts use a different window and the Claude-profile records only.

Assumed monthly commitment$759.00Four providers · one seat each
Allocated to the observed week$177.107/30 of the September monthly fee
Observed input with a subscription mapping24.719B98.15% cache reads · repeated input exposures

Dollar values are subscription scenarios and accounting allocations, not API charges or a reconciled invoice. More recorded tokens do not imply more useful work.

The subscription scenario

Assume the highest subscription tier used in the original audit for Claude, GLM, Kimi and Codex, with no direct API spend, extra credits or tax. The roster is one subscription per provider, not one subscription per local profile or model.

Monthly fees below come from the original audit: Claude $200 and ChatGPT Pro for Codex $200 were source-confirmed there; Kimi $199 and GLM $160 are unverified scenario inputs. These are not a fresh checkout quote.

Assumed subscriptionMonthly USD / seatSeatsSeven-day USDObserved input tokensCache-read share
ClaudeFee sourced in original audit$46.6723,326.07M98.18%
GLMFee scenario, unverified$37.33687.45M97.88%
KimiFee scenario, unverified$46.43378.55M98.75%
CodexFee sourced in original audit$46.67327.24M95.84%

Default scenario · values update locally; nothing is sent or saved.

Changing fees or seats changes the allocation. It does not create additional observed traffic or prove that the assumed accounts produced it. Other devices, web sessions and missing telemetry can consume the same subscriptions.

Show allocated dollars per million observed input tokens

This is a retrospective utilization ratio, not a provider price or a model-quality comparison. Redundant cached input increases the denominator and makes the ratio look cheaper. Use the whole-workflow evidence below to judge whether the capacity was useful.

SubscriptionAllocated USD / million observed input
Claude$0.0020
GLM$0.0543
Kimi$0.1227
Codex$0.1426

Where the observed input went

These volumes describe workload and usage coverage. They do not establish that one provider is more efficient: the task mix, models, session lifetimes and unobserved traffic differ.

Claude
23.326B
GLM
0.687B
Kimi
0.379B
Codex
0.327B
Horizontal scale: reported input tokens, billions · each bar starts at zero

Source: deduplicated Claude-profile message receipts plus Codex counter differences, September 20–26. Input includes cache reads. Another 0.590M DeepSeek input tokens have no subscription assumption and are excluded from the cash allocation.

Allocate each fee across its observed models

The rule is simple: provider’s seven-day fee × model’s share of that provider’s observed input. Every model shares its provider’s one fee. Cached input receives the same allocation weight as fresh input; this is an accounting convention, not a claim about quota rules.

Reported modelAssumed subscriptionInput, million tokensReported output, million tokensAllocated weekly USD
claude-opus-5-5Claude9,776.6597.614$19.5594
claude-sonnet-5Claude6,331.4199.579$12.6668
claude-opus-4-8Claude5,385.97514.460$10.7753
claude-fable-5-1Claude1,433.1651.903$2.8672
glm-5.3GLM577.5832.361$31.3669
k3Kimi378.5461.039$46.4333
claude-opus-5Claude343.3221.230$0.6869
gpt-6-astraCodex250.3061.024$35.6951
glm-5.3-flashGLM109.8650.537$5.9664
claude-opus-4-7Claude54.8101.226$0.1097
gpt-5.6-solCodex37.6230.074$5.3653
codex-auto-reviewCodex27.3600.045$3.9017
gpt-5.6-terraCodex10.0840.031$1.4380
gpt-6-lunaCodex1.8690.007$0.2666
claude-haiku-4-5-20251001Claude0.7160.001$0.0014
deepseek-v4-flashUnallocated0.5900.039—

Model identifiers are what the harness recorded, not independently verified serving identities. Output receipts have known completeness concerns and are excluded from the allocation. Values are rounded; the underlying calculation uses full input totals. Download the original default-scenario JSON.

What did a completed outcome cost?

One manually reconstructed accepted item links an implementation, a review and an independent verification. It shows why counting the worker alone misses much of the identifiable work. It does not yet supply a fully loaded delivery cost.

ONE ACCEPTED ITEM · IDENTIFIABLE CHILD STREAMS

23.28M input tokens
Worker · 10.87MReview · 10.97MVerify · 1.45M

53.3% of the observed direct input came from review and verification. Counting just the worker understates these three streams by a factor of 2.14.

ALLOCATION TO THESE CHILD STREAMS

$0.0466
Share of the assumed Claude weekly fee. This is not the price of delivering the outcome.

TIME TO INDEPENDENT ACCEPTANCE

About 13 hours
49.8 minutes of summed child-session spans; spans include tools and waiting, not measured compute or human time.

Source: one reconstructed accepted work item, with identifiers withheld. First worker activity September 23 at 23:50:49 UTC; acceptance September 24 at 12:50:47 UTC, with 4/4 verification rows passed. The repair path may include further work; shared controllers, intake, unlinked attempts, human correction and later regressions are outside this allocation.

Why there is no fleet-wide dollar figure yet

  • 56.2% of weekly child input is covered by candidate work-key joins.
  • 38 of 59 keys with a latest verified verdict also have attempts beginning after that verdict.
  • 707 candidate keys in the wider extraction are associations, not 707 accepted outcomes.

A reused brief key cannot let an old verdict approve a new revision. These gaps prevent a reliable fleet-wide accepted-outcome denominator.

The accounts to keep

  • Cash: attributable subscription expense ÷ accepted outcomes, with opening and closing work in progress.
  • Capacity: provider quota used through acceptance, including retries, failures and escalation.
  • Time: elapsed time to acceptance, quota waits and human correction minutes.
  • Quality: independent acceptance, serious findings and subsequent regressions.

What the operator can decide from this

A fixed subscription does not charge more for each included call. Better routing first buys headroom: less waiting, more accepted work, or room to finish review and verification. Cash falls when that headroom permits a smaller commitment or avoids another purchase.

The audit identifies where to investigate—long-running context, shared coordination and later-stage effort. It does not establish a winning model–effort pair or prove that a smaller context preserves quality. Capture those settings per task and follow the rework before deciding.

A controlled comparison

  1. Instrument and choose representative work. Link versioned work, attempts, resolved model and effort, usage and independent verdicts. Freeze task starting states and checks.
  2. Compare two policies on 24–30 tasks. Interleave runs, isolate outputs, keep reviewer/verifier policy fixed and record all repairs and unresolved outcomes.
  3. Adjudicate and follow regressions. Budget 3–5 working days for a first pilot, and roughly 10–14 calendar days for replication plus a week of follow-up, subject to quota and verification capacity.

These are planning estimates. Task diversity, runtime variance and the acceptable quality-loss bound determine the larger sample needed.

Methods and coverage

Usage extraction

Four Claude-profile trees were inventoried: 17,608 JSONL files, approximately 14.05 GB. Files modified since September 13 were parsed and records timestamp-filtered. This prefilter can miss imported files with older modification times. Logs were live during extraction.

For the primary week, Claude-profile records contain 118,456 unique reported model messages across 2,650 streams. Message IDs were deduplicated across blocks and copied histories, taking component maxima. Synthetic messages were excluded.

Cross-harness accounting

The Codex extraction covers 66 session IDs in the week, using cumulative-counter differences and excluding inherited initial counters. Cached input is already included in Codex total input. GLM and Kimi traffic inside Codex is assigned to those subscriptions, not counted as Codex subscription use.

Cache and quota semantics differ by provider. Missing or zero cache-write fields alone do not prove billing errors. The local receipt totals have not been reconciled with provider meters, account reset windows or invoices.

Methods and coverage · Two-week measurements and methods

Sources and assumptions

  1. Anthropic Max plan and OpenAI Pro tiers: fee sources cited by the original September 27 audit. Kimi and GLM fees remain scenario assumptions.
  2. Z.ai plan documentation: quota weights and time-dependent rules illustrate why raw tokens are not interchangeable quota units.
  3. Local model receipts, assignment joins, GitHub metadata and verification ledgers described in the audit. No new model execution or provider-billing reconciliation was performed for this report.