Design note, September 2026. The opening report example is illustrative. The worked design comes from our discussion about model cost and brief quality. Six-dimension assessment, independent acceptance review and authoring-quality analysis are proposed additions to Assay; they are not implemented controls or measured improvements.
Consider a request to build a report showing which coding model gives the best value. An agent produces a table, adds tests, and opens a pull request. The tests pass. But the report treats missing cost as zero and omits abandoned attempts. Its cheapest model may be the one with the least complete records.
The code can faithfully implement that interpretation. Reviewing it more carefully will not necessarily recover the question nobody settled: what counts as value, and which costs belong in the comparison?
This is the problem I want the brief-quality work to address. Preserve the decisions behind a brief, and preserve the brief that execution started from. Otherwise, a successful delivery can hide the gaps its implementer had to repair.
ambiguity can survive the whole development cycle
The report is an example, not an incident we measured. It shows how a missing requirement can travel through the software development lifecycle.
During requirements, “best value” leaves the denominator open. During design, someone chooses a treatment for missing data. During implementation, that choice becomes a default. During testing, fixtures reproduce the default. During review, the conversation concentrates on whether the code handles those fixtures correctly. By delivery, an unanswered question has acquired the appearance of an agreed requirement.
Each participant may have done competent work within the information they received. The handoffs lost the uncertainty.
An agent can also rescue the request: notice the missing definition, investigate it and improve the contract. That is a useful outcome. But if the final ticket replaces the original, the team cannot distinguish a well-authored brief from one that needed substantial repair. The next model-cost comparison rewards or penalizes the worker for work the brief never acknowledged.
This is why the records before implementation matter to evaluation afterwards. Faster code generation does not tell us whether the right problem reached the implementer.
an idea needs an outcome before it needs tasks
Our actual discussion began with a choice between a stronger model using less reasoning effort and a cheaper model using more. To compare them fairly, we needed to know what kind of task each received and what happened beyond its first response.
That exposed a more immediate question: were we measuring how well we authored the briefs themselves?
The proposed outcome became the ability to compare the brief as authored, the contract approved for dispatch, and the observed result. That gives the specification a job. It must define the assessment, the records and the acceptance process well enough that separate implementers do not invent incompatible meanings.
For the report example, cost would include authoring, review, implementation, repair and unsuccessful attempts in a declared cohort. The denominator would be independently verified, accepted briefs. Missing cost would remain missing, with coverage shown. A design that leaves these choices to the chart author has deferred requirements work into implementation. The draft specification states in its metrics section: “incomplete coverage yields partial-cost label, never complete-spend ranking.” That clause gives a reviewer grounds to reject the opening example’s cheapest-model conclusion.
Assay's existing brief contract records why, sources and dependencies. They are the beginning of a decision trail: a way to follow the request through the choices that made it executable. They do not make those choices correct by themselves.
six questions make different demands visible
The proposed assessment separates six dimensions because they influence different parts of delivery.
| Dimension | Question before dispatch | Why it matters later |
|---|---|---|
| Work size | How much work is expected? | Capacity and splitting depend on volume. |
| Specification completeness | Are facts missing or choices still open? | Discovery and authorized design need different treatment. |
| Reasoning difficulty | Is this an established pattern or a new reasoning problem? | The implementer may need judgment before persistence helps. |
| Coupling | Which shared contracts or component boundaries move? | Integration work and dependency order become visible. |
| Verification strength | How decisively can available checks establish success? | Review must cover judgments the tests cannot settle. |
| Failure consequence | What happens if the change is wrong? | Oversight and recovery planning should reflect the damage. |
A broad mechanical rename and a tiny retry change can occupy very different positions across these dimensions. Combining them into one complexity score would discard information the dispatcher and reviewer need.
Each assessment would carry evidence and anchored categories, including unknowns. The existing effort: S/M/L remains work size. Model reasoning effort belongs in the execution record. Neither a high-effort setting nor a capable model supplies missing requirements automatically.
a recommendation must survive becoming a ruling
The proposed independent acceptance review raises policy choices. Who writes the first definition of done? Does a reviewer block dispatch or observe in shadow mode? Is a separate reviewer sufficient, or must it use another model family? What evidence can the pilot retain?
The specification makes those alternatives inspectable. A ruling then records what the responsible authority chose, where that choice was made, and which specification revision it affects. Until then, a recommendation remains open.
This boundary matters during design: one agent's preferred option must not become another agent's instruction merely through confident wording. It matters during maintenance too. A later change should reveal which earlier decision it reverses and why.
The current recommendation keeps the author responsible for the initial contract. A non-author reviewer would first derive plausible failure cases from the intent and constraints, before reading the proposed verification. It would then challenge or amend those checks. The pilot and enforcement policy still require a ruling.
challenge success before the implementation teaches you what to test
For the illustrative report, an independent reviewer could ask what happens when cost is unavailable, when a worker abandons an attempt, or when a late defect changes an earlier outcome. These questions challenge the definition of success before an implementation supplies convenient answers.
The existing authoring skill already asks the author to imagine a shipped-but-wrong result and map each failure to a verification row or a review-only limitation. The proposal adds a separately attributable challenge and approval of the exact contract being dispatched.
This review has a different subject from code review. It asks whether the contract could accept the wrong result. Code review examines the implementation against that contract and the surrounding system. Post-merge verification checks the merged revision. None makes the others redundant; none proves every operational claim.
Using another model may diversify the analysis. It does not guarantee independence of assumptions or turn an assertion into evidence. That is why the proposed reviewer derives failure cases before seeing the author's test plan.
dependencies turn the design into an executable stream
The draft stream makes the prerequisite decisions visible in the work order; its README lists the critical path. Design and policy rulings precede schema changes. Procedures and event records can then develop on separate branches. Admission checks and outcome collection feed analysis; a bounded pilot comes afterwards.
The schema brief explicitly depends on the policy-rulings brief.
The schema implementer therefore cannot proceed as if those choices were settled. At the linked design revision, the stream is parked while its design remains unapproved.
This is where planning meets execution. The dependency prevents a schema implementer from choosing policy through a default value. Each brief then gives one implementer the purpose, relevant facts, scope, dependencies, and commands with expected results. Waves expose what can proceed together. They also expose what must wait.
In the later lifecycle, the implementation goes through review, then a fresh non-implementer verifies the merged work against the brief. The existing verify-desk contract distinguishes merged from verified. Offline checks remain limited to what they observe; a merged change is not evidence of production behavior.
keep enough history to improve the author
The proposed authored, approved and completed snapshots would let the report distinguish an omission caught before dispatch from one discovered during implementation. An amendment would preserve the old contract and state why it changed.
Outcome analysis would separate brief omissions, implementation errors, environment problems and changed requirements, while retaining unresolved attribution. A repair is evidence to investigate, not an automatic verdict against the author. More reviewer findings could mean better detection, noisier review or more difficult work.
The six dimensions help describe those differences. Execution records supply the actual model, reasoning effort, cost and repair history. Comparing like work with coverage reported is more informative than ranking authors or models by their raw completion counts.
This closes the development cycle back into requirements. A corrected brief guides today's worker. Its original version tells us what tomorrow's author needs to do differently.
the hedge
The new controls are a design, not evidence that extra review pays for itself. It can add delay, duplicate checks or reproduce the author's assumptions. Assessment categories may be applied inconsistently. A pilot needs comparable work, full cost accounting and an observation period for defects found later.
You can examine the underlying problem before adopting any new tooling. Take a completed change. Recover its original request, the decisions that settled its scope, and the acceptance criteria in force when work began. Compare those records with the final result and the repairs along the way. If only the final ticket survives, start preserving the next one before implementation begins.
The companion explainer follows the same illustrative report through the development cycle. Watch on YouTube or read the narration.