Concept register · Theme 09 of 14 4 concepts · 41 talks
assay · concepts · agent-evals
Agent evals
Evaluation is moving off the pre-launch checkpoint: scored continuously on live traffic, over the trajectory rather than only the artifact, by verifiers that are themselves audited, and capability scores are finally distinguished from deployment decisions.
assay: readiness exam shipped · live evals open
§1What it is
A standing process, not a gate
Production outcome evals run continuously against live traffic and are measured by what changed because of them, not by a headline pass rate. Three shapes recur: always-on outcome evals, shadow mode (the agent proposes, the human does, the diff feeds learning), and causal in-the-wild evaluation that defines its estimand before it counts.
Score the path, not just the artifact
Trajectory grading treats the reasoning path and the tool-call sequence as a first-class object alongside the artifact, because outcome-only grading passes an agent that arrived by a forbidden or wasteful route, the trajectory is where the bugs live.
QC the measuring stick
Benchmark numbers routinely fail to mean what they appear to: saturation and contamination, verifiers with their own error rates, agents attacking the eval harness, and run variance a single number hides. And a benchmark score measures a capability ceiling, deciding to deploy is a different exam, a per-use-case qualification with a risk envelope and an oversight policy.
§2The concepts in this theme
Each concept has its own page in the concept register, with sightings from every event we review, and where Assay stands on each.
- Production outcome evals: evaluation as a standing process, not a gate established 13
- Trajectory grading: scoring the path, not just the artifact established 12
- Benchmark integrity and run variance, QC for the measuring stick established 11
- Capability-ceiling evals vs deployment-readiness evals corroborated 9
§3How Assay implements this
The readiness exam is shipped
A brief that cannot state a falsifiable check is a brief that has not been specified, and the review gate before authoring is where that is caught, the readiness discipline written into a methodology rather than a benchmark suite. Verify rows are the task-plus-rubric slice; evidence-cited verdicts are the trajectory-plus-artifact slice. Abstention is already first-class: blocked-on-human and needs-decision are sanctioned exits, not failures.
Pre-merge, and once
Verify rows run before a merge and run once. That is the honest position: nothing always-on exists, no post-merge standing evaluation, no judge agent scoring live outcomes, no re-run of a brief’s verify rows after the world moves under them. And nothing ties a landed brief back to whether a higher-level outcome moved; that absence is precisely why the recurring reports count activity.
The named gaps
No trace-level metric of any kind exists, “tokens per verified brief” is the concrete upgrade over activity counting. The verifiers have no audit of their own: a verify row that passes on a sabotaged submission is a broken row, and sabotage runs as a routine audit of row quality is the single cheapest steal in this theme. Verify rows also stay closed-form and adversary-resistant rather than judged by a model, a verify row is a tiny verifier, and verifiers are the thing being attacked.
§4Talks that cover this theme
#043Panel: Agentic AI in Finance & LegalChandhok, Circle; Shafiq, Wells Fargo
#110Agent Arena: Causal Evaluations of Agents in the Real WorldAnastasios N Angelopoulos, Arena
#144From Training to Evaluation: Open Recipes for Agentic AIChenguang Wang, UCSC / Scale AI
#001Enterprise AIAdarsh Hiremath, Mercor
#108The Art & Science of Benchmarking AgentsVincent Sunn Chen, Snorkel
#142Data Benchmarks: Where Everything’s Made UpGrace Tang, Hex
#153The Exam Before Enterprise DeploymentYuan Emily Xue, Scale AI
#120Workshop: Future of Agent EvaluationBerkeley RDI and others
8 of 41 talks shown, the ones that reach the most concepts in this theme. Every sighting, per talk, is on the concept pages above.