# Methods and measurement boundaries

This appendix supports the article and subscription report. It contains aggregate measurements, not raw session transcripts or account records.

## Two-week context profile

Window: September 13 00:00 UTC inclusive to September 27 00:00 UTC exclusive. Four local Claude-profile trees; copied message IDs deduplicated, synthetic messages excluded. For repeated usage records, retain component maxima. Input means fresh input plus cache creation plus cache reads. Output has known completeness concerns.

244,637 model-message receipts across 5,278 streams reported 48,499,659,124 input tokens: 34,355,511 fresh, 787,576,100 cache writes and 47,677,727,513 cache reads. Cache reads are 98.3053% of recorded input.

The aggregation is: filter timestamps; group deduplicated receipts by recognized responsibility and parent/child; calculate median and 90th percentile of input per call. For equal-stream weighting, calculate each stream's median first, then take the median of those medians. These are consumption measures, not evidence of required context.

| Responsibility | Main call median | Child call median | Child call p90 | Main equal-stream median | Child equal-stream median |
|---|---:|---:|---:|---:|---:|
| Implementation |441090|132739|271016|52456|106690.5|
| Review |480240.5|93246.5|141164|471428|90183|
| Verification |386337.5|85157|130799|317586|80288.5|

## Seven-day subscription scenario

Window: September 20 00:00 UTC inclusive to September 27 00:00 UTC exclusive. The subscription report combines the Claude-profile receipts with inspected Codex sessions. Codex cumulative counter differences exclude inherited initial counters; its cached input is already part of total input. GLM/Kimi traffic in that harness maps to those subscriptions. The model-usage JSON retains exact aggregate inputs; all dollar figures follow the stated fee scenario.

Monthly fee × seats × 7/30 gives the period allocation. Each model receives its provider's period allocation multiplied by its share of observed input. This treats all input categories alike solely for allocation; provider quotas need separate documented weights. Output is displayed but not used in the allocation. The JSON is the original default scenario, not a saved copy of interactive edits.

## One reconstructed accepted item

Worker input: 10,870,518; review: 10,968,490; verification: 1,445,621. Sum: 23,284,629. Review plus verification: 53.3146%. First worker activity: September 23, 23:50:49 UTC; independent acceptance: September 24, 12:50:47 UTC. The verification ledger recorded four of four required checks passing. Summed child-session spans are 49.8 minutes and include tools and waiting. Neither those spans nor time to acceptance measure human labor.

The sum excludes shared coordination, intake, unlinked attempts and later correction. Repair arrows in the article illustrate a possible workflow, not measured savings or a full rework account.

## Why the cohort cannot supply a fleet-wide outcome cost

Candidate joins cover 56.2% of weekly child input. Of 59 work keys with a latest verified verdict, 38 also contain attempts beginning after that verdict. A work key can be reused; a versioned artifact and attempt identity are needed before assigning acceptance. The 707 candidate keys in the wider extraction are not 707 accepted outcomes.

The file inventory used a modification-time prefilter before record timestamp filtering; imported files with older modification times may be missed. Logs were live at extraction. Harness-reported model IDs do not independently establish upstream serving identities. Account counts, invoices, other-device use and provider quota meters were not reconciled.
