Concept register · Theme 13 of 14 1 concepts · 14 talks

assay  ·  concepts  ·  autonomous-discovery

Autonomous discovery

Agent systems pointed at the whole research loop, ingest, hypothesize, implement, run real experiments, falsify most of it, write up what survived, with the human posing the question and accepting the result.

assay: shaped like the dreaming pass · unbuilt

AI co-scientists and autonomous discovery loops — 11 sources, established
1 concepts · 11 independent sources · 1 established

§1What it is

Three infrastructural legs

The systems that work stand on auto-reset (a fresh environment per attempt), auto-improve (the loop drives its own upgrades) and auto-evaluate (outcomes judged by the environment, not by model confidence). Feedback must come from the environment, and selection matters as much as generation, the loop needs a mechanism for choosing among candidate ideas, not only for generating them.

Two dissents, recorded

The reported successes are mostly bounded, verifiable problems, exactly the class a register-and-gates methodology fits, and full autonomy over open-ended questions remains undemonstrated. The dissents are held, not dismissed.


§2The concepts in this theme

Each concept has its own page in the concept register, with sightings from every event we review, and where Assay stands on each.


§3How Assay implements this

The legs map onto what ships

Auto-reset is worktree discipline: every worker gets a fresh worktree off the mainline. Auto-improve is the worker desks. Auto-evaluate is the verify rows executed after merge by someone other than the implementer. The end-to-end research pipeline, out-of-band idea generation, agents arguing the candidates, harness-gated execution, human accept-or-reject at the end, is the shape the dreaming pass was designed to have. None of that runs today.

The strongest outside validation in the scan

The open-source “lab” results, branch per idea, sandboxed writes, findings written down, are Assay’s core bet arrived at independently, sandbox rationale included, with loops carrying all three legs saturating their task. Three transfers are recorded: an ELO-style tournament as a selection step among competing brief plans (none exists today), iterating with the environment as the argument for verify rows against real CI, and revisiting brief granularity on the task-horizon cadence rather than treating it as fixed.


§4Talks that cover this theme