Subscription capacity and cost allocation
Analysis snapshot · 27 September 2026
What one week of recorded agent work tells us about subscription utilization—and what we still need to measure the cost of a completed outcome.
Scope: September 20 00:00 to September 27 00:00 UTC. Usage combines the four inspected Claude-profile trees with the inspected Codex sessions. The article’s two-week context charts use a different window and the Claude-profile records only.
Dollar values are subscription scenarios and accounting allocations, not API charges or a reconciled invoice. More recorded tokens do not imply more useful work.
The subscription scenario
Assume the highest subscription tier used in the original audit for Claude, GLM, Kimi and Codex, with no direct API spend, extra credits or tax. The roster is one subscription per provider, not one subscription per local profile or model.
Monthly fees below come from the original audit: Claude $200 and ChatGPT Pro for Codex $200 were source-confirmed there; Kimi $199 and GLM $160 are unverified scenario inputs. These are not a fresh checkout quote.
| Assumed subscription | Monthly USD / seat | Seats | Seven-day USD | Observed input tokens | Cache-read share |
|---|---|---|---|---|---|
| ClaudeFee sourced in original audit | $46.67 | 23,326.07M | 98.18% | ||
| GLMFee scenario, unverified | $37.33 | 687.45M | 97.88% | ||
| KimiFee scenario, unverified | $46.43 | 378.55M | 98.75% | ||
| CodexFee sourced in original audit | $46.67 | 327.24M | 95.84% |
Default scenario · values update locally; nothing is sent or saved.
Changing fees or seats changes the allocation. It does not create additional observed traffic or prove that the assumed accounts produced it. Other devices, web sessions and missing telemetry can consume the same subscriptions.
Show allocated dollars per million observed input tokens
This is a retrospective utilization ratio, not a provider price or a model-quality comparison. Redundant cached input increases the denominator and makes the ratio look cheaper. Use the whole-workflow evidence below to judge whether the capacity was useful.
| Subscription | Allocated USD / million observed input |
|---|---|
| Claude | $0.0020 |
| GLM | $0.0543 |
| Kimi | $0.1227 |
| Codex | $0.1426 |
Where the observed input went
These volumes describe workload and usage coverage. They do not establish that one provider is more efficient: the task mix, models, session lifetimes and unobserved traffic differ.
Source: deduplicated Claude-profile message receipts plus Codex counter differences, September 20–26. Input includes cache reads. Another 0.590M DeepSeek input tokens have no subscription assumption and are excluded from the cash allocation.
Allocate each fee across its observed models
The rule is simple: provider’s seven-day fee × model’s share of that provider’s observed input. Every model shares its provider’s one fee. Cached input receives the same allocation weight as fresh input; this is an accounting convention, not a claim about quota rules.
| Reported model | Assumed subscription | Input, million tokens | Reported output, million tokens | Allocated weekly USD |
|---|---|---|---|---|
claude-opus-5-5 | Claude | 9,776.659 | 7.614 | $19.5594 |
claude-sonnet-5 | Claude | 6,331.419 | 9.579 | $12.6668 |
claude-opus-4-8 | Claude | 5,385.975 | 14.460 | $10.7753 |
claude-fable-5-1 | Claude | 1,433.165 | 1.903 | $2.8672 |
glm-5.3 | GLM | 577.583 | 2.361 | $31.3669 |
k3 | Kimi | 378.546 | 1.039 | $46.4333 |
claude-opus-5 | Claude | 343.322 | 1.230 | $0.6869 |
gpt-6-astra | Codex | 250.306 | 1.024 | $35.6951 |
glm-5.3-flash | GLM | 109.865 | 0.537 | $5.9664 |
claude-opus-4-7 | Claude | 54.810 | 1.226 | $0.1097 |
gpt-5.6-sol | Codex | 37.623 | 0.074 | $5.3653 |
codex-auto-review | Codex | 27.360 | 0.045 | $3.9017 |
gpt-5.6-terra | Codex | 10.084 | 0.031 | $1.4380 |
gpt-6-luna | Codex | 1.869 | 0.007 | $0.2666 |
claude-haiku-4-5-20251001 | Claude | 0.716 | 0.001 | $0.0014 |
deepseek-v4-flash | Unallocated | 0.590 | 0.039 | — |
Model identifiers are what the harness recorded, not independently verified serving identities. Output receipts have known completeness concerns and are excluded from the allocation. Values are rounded; the underlying calculation uses full input totals. Download the original default-scenario JSON.
What did a completed outcome cost?
One manually reconstructed accepted item links an implementation, a review and an independent verification. It shows why counting the worker alone misses much of the identifiable work. It does not yet supply a fully loaded delivery cost.
ONE ACCEPTED ITEM · IDENTIFIABLE CHILD STREAMS
53.3% of the observed direct input came from review and verification. Counting just the worker understates these three streams by a factor of 2.14.
ALLOCATION TO THESE CHILD STREAMS
TIME TO INDEPENDENT ACCEPTANCE
Source: one reconstructed accepted work item, with identifiers withheld. First worker activity September 23 at 23:50:49 UTC; acceptance September 24 at 12:50:47 UTC, with 4/4 verification rows passed. The repair path may include further work; shared controllers, intake, unlinked attempts, human correction and later regressions are outside this allocation.
Why there is no fleet-wide dollar figure yet
- 56.2% of weekly child input is covered by candidate work-key joins.
- 38 of 59 keys with a latest verified verdict also have attempts beginning after that verdict.
- 707 candidate keys in the wider extraction are associations, not 707 accepted outcomes.
A reused brief key cannot let an old verdict approve a new revision. These gaps prevent a reliable fleet-wide accepted-outcome denominator.
The accounts to keep
- Cash: attributable subscription expense ÷ accepted outcomes, with opening and closing work in progress.
- Capacity: provider quota used through acceptance, including retries, failures and escalation.
- Time: elapsed time to acceptance, quota waits and human correction minutes.
- Quality: independent acceptance, serious findings and subsequent regressions.
What the operator can decide from this
A fixed subscription does not charge more for each included call. Better routing first buys headroom: less waiting, more accepted work, or room to finish review and verification. Cash falls when that headroom permits a smaller commitment or avoids another purchase.
The audit identifies where to investigate—long-running context, shared coordination and later-stage effort. It does not establish a winning model–effort pair or prove that a smaller context preserves quality. Capture those settings per task and follow the rework before deciding.
A controlled comparison
- Instrument and choose representative work. Link versioned work, attempts, resolved model and effort, usage and independent verdicts. Freeze task starting states and checks.
- Compare two policies on 24–30 tasks. Interleave runs, isolate outputs, keep reviewer/verifier policy fixed and record all repairs and unresolved outcomes.
- Adjudicate and follow regressions. Budget 3–5 working days for a first pilot, and roughly 10–14 calendar days for replication plus a week of follow-up, subject to quota and verification capacity.
These are planning estimates. Task diversity, runtime variance and the acceptable quality-loss bound determine the larger sample needed.
Methods and coverage
Usage extraction
Four Claude-profile trees were inventoried: 17,608 JSONL files, approximately 14.05 GB. Files modified since September 13 were parsed and records timestamp-filtered. This prefilter can miss imported files with older modification times. Logs were live during extraction.
For the primary week, Claude-profile records contain 118,456 unique reported model messages across 2,650 streams. Message IDs were deduplicated across blocks and copied histories, taking component maxima. Synthetic messages were excluded.
Cross-harness accounting
The Codex extraction covers 66 session IDs in the week, using cumulative-counter differences and excluding inherited initial counters. Cached input is already included in Codex total input. GLM and Kimi traffic inside Codex is assigned to those subscriptions, not counted as Codex subscription use.
Cache and quota semantics differ by provider. Missing or zero cache-write fields alone do not prove billing errors. The local receipt totals have not been reconciled with provider meters, account reset windows or invoices.
Sources and assumptions
- Anthropic Max plan and OpenAI Pro tiers: fee sources cited by the original September 27 audit. Kimi and GLM fees remain scenario assumptions.
- Z.ai plan documentation: quota weights and time-dependent rules illustrate why raw tokens are not interchangeable quota units.
- Local model receipts, assignment joins, GitHub metadata and verification ledgers described in the audit. No new model execution or provider-billing reconciliation was performed for this report.