Concept register · Concept 62 of 64 · Theme: autonomous discovery Reviewed 2026-09-01

assay  ·  concepts  ·  autonomous-discovery

AI co-scientists and autonomous discovery loops

Agent systems are being pointed at the whole research loop rather than at one step of it: ingest the literature, generate hypotheses, implement, run real experiments, falsify most of them, and write up what survived, with the human moved up to posing the question and accepting or rejecting the result.

established · assay: designed, unbuilt

11 independent sources · sighted at the Agentic AI Summit 2026 · last reviewed 2026-09-01


§1What it is

Three infrastructural legs

The loop runs without a human in each turn only when three things exist: auto-reset of the environment, auto-improve of the artifact, and auto-evaluate of the result. Where all three are present, tasks have been driven from 0% to near-total success autonomously. Where one is missing, a human is back in every turn and the loop is a demo.

Feedback must come from the environment

A breakthrough is by definition outside the training set, so thinking harder about the literature does not substitute for running the experiment. The history cited against “thinkism” is pointed: the Nobel that mattered for superconductivity was for liquefying helium, a capability build, with the phenomenon stumbled into; a headline superconductor came from trying tens of thousands of materials rather than from theory. Mathematics and theoretical computer science are a different regime.

Selection matters as much as generation

The systems that work falsify aggressively, often by rating candidate ideas against each other in a tournament before committing compute. Progress arrives in jumps, not smoothly, one autonomous run had 240 experiments between its last two improvements, which makes cheap selection the difference between a loop that pays and one that burns.

Two dissents, recorded

A frontier lab reports seeing no particularly successful auto-research rollout with today’s models: progress appears early and then stalls, because research is serial and long-horizon. And the field’s own framing is that autonomy on grand challenges is gated by trust rather than by capability.


§2Sightings

Agentic AI Summit 2026 · 14 sightings

Also: AlphaEvolve; AlphaFold and AlphaGenome; the FrontierScience benchmark; PutnamBench; METR time-horizon measurements; Goodfire Silico; Prima Mente Pleiades; and “the lab” (open source, workspace, dashboard, one git branch per idea).


§3Where Assay stands

Designed, unbuilt: and the closest external match in the scan

The end-to-end research pipeline is the shape Assay’s dreaming pass was designed to have: out-of-band idea generation, agents arguing the candidates, harness-gated execution, and a human accept or reject at the end. None of that runs today.

The three legs map cleanly onto what is shipped

Auto-reset is worktree discipline: every worker gets a fresh worktree off the mainline. Auto-improve is the worker desks. Auto-evaluate is the Verify rows executed after merge by someone other than the implementer, per the lifecycle. The 0%-to-99% result is evidence that loops with all three legs saturate their task, which argues for closing the remaining gap rather than adding a fourth mechanism. The open-source “lab” is the strongest outside validation of Assay’s core bet, branch per idea, sandboxed writes, findings written down, arrived at independently, sandbox rationale included.

Three specific transfers

The ELO idea tournament is a candidate mechanism for choosing among competing brief plans or proposed memory updates, where Assay currently has no selection step at all. The iterate-with-the-environment stance is the external argument for why Verify rows run against real CI rather than against model confidence. And the time-horizon cadence is a standing review trigger: if reliably-executable task length doubles roughly every six months, the granularity ceiling on a brief should be revisited on that cadence rather than treated as fixed. The dissent is worth holding too, the successes reported here are mostly bounded, verifiable problems, which is exactly the class a brief is.


§4Watch

  • Whether any auto-research result is independently reproduced, the striking numbers here are each single-team single-run reports.
  • Whether the dissent softens: a long-horizon research rollout that does not stall would be the signal that the loop generalizes past bounded tasks.
  • Whether idea-selection mechanisms (ELO tournaments, verifier-fed replanning) show up outside research settings, which is what would make them portable to plan authoring.