Concept register · Theme 13 of 14 1 concepts · 14 talks
assay · concepts · autonomous-discovery
Autonomous discovery
Agent systems pointed at the whole research loop, ingest, hypothesize, implement, run real experiments, falsify most of it, write up what survived, with the human posing the question and accepting the result.
assay: shaped like the dreaming pass · unbuilt
§1What it is
Three infrastructural legs
The systems that work stand on auto-reset (a fresh environment per attempt), auto-improve (the loop drives its own upgrades) and auto-evaluate (outcomes judged by the environment, not by model confidence). Feedback must come from the environment, and selection matters as much as generation, the loop needs a mechanism for choosing among candidate ideas, not only for generating them.
Two dissents, recorded
The reported successes are mostly bounded, verifiable problems, exactly the class a register-and-gates methodology fits, and full autonomy over open-ended questions remains undemonstrated. The dissents are held, not dismissed.
§2The concepts in this theme
Each concept has its own page in the concept register, with sightings from every event we review, and where Assay stands on each.
- AI co-scientists and autonomous discovery loops established 11
§3How Assay implements this
The legs map onto what ships
Auto-reset is worktree discipline: every worker gets a fresh worktree off the mainline. Auto-improve is the worker desks. Auto-evaluate is the verify rows executed after merge by someone other than the implementer. The end-to-end research pipeline, out-of-band idea generation, agents arguing the candidates, harness-gated execution, human accept-or-reject at the end, is the shape the dreaming pass was designed to have. None of that runs today.
The strongest outside validation in the scan
The open-source “lab” results, branch per idea, sandboxed writes, findings written down, are Assay’s core bet arrived at independently, sandbox rationale included, with loops carrying all three legs saturating their task. Three transfers are recorded: an ELO-style tournament as a selection step among competing brief plans (none exists today), iterating with the environment as the argument for verify rows against real CI, and revisiting brief granularity on the task-horizon cadence rather than treating it as fixed.
§4Talks that cover this theme
#004A Lab Notebook for AgentsChuan Li, Lambda
#007From Models to Agents to DiscoverySaurabh Tiwary, Google
#017Opportunities and Challenges for Long Horizon AgentsJerry Tworek, OpenAI
#023Robotics: EndgameJim Fan, NVIDIA
#034The Eureka Machine: Recursive Superintelligence for ScienceRichard Socher, Recursive
#036Combining Experiments, LLMs, and Theory to Discover Quantum MaterialsEkin Dogus Cubuk, Periodic Labs
#079Unlocking Scientific Abundance by Learning from Superhuman AIEric Ho, Goodfire
#082Workshop: Open Source Agent InvestigationsLambda / Berkeley RDI
8 of 14 talks shown, the ones that reach the most concepts in this theme. Every sighting, per talk, is on the concept pages above.