The meta-loop that rewrites the optimization loop, live
How evo 0.5 runs the optimizer and a meta-controller as two threads on one event loop
the problem with a fixed loop
An autoresearch loop has a shape: look at what's been tried, gather signal, propose the next experiments, run them, keep what works, repeat. That shape is usually hard-coded. It runs the same way on round 1 and round 40 — same width, same phases, same prompts — no matter what the run is actually learning.
But a long run learns things about itself. Scan stops surfacing anything useful. One branch keeps producing the same failure. The whole search plateaus and needs to widen, or research a new direction, or just stop a doomed experiment before it burns an hour of GPU. A fixed loop can't act on any of that — it has no way to change its own shape mid-run.
evo 0.5 adds a second thread that can. The optimizer runs as before, and alongside it a meta-controller watches the run and reshapes the loop while it's running — retuning it, toggling phases, rewriting the prompts the loop uses, even stopping experiments. Here's how the two fit together.
two threads, one event loop
The whole thing is one workflow script. It launches two async loops and joins them with Promise.all: the optimize loop is the driver — its result is the run's result — and the meta loop is a concurrent, advisory observer. They never run in separate processes; they're two coroutines on a single-threaded event loop, sharing one mutable object between them.
harness object every round; the meta thread writes it. Same event loop, so writes land between the optimizer's awaits — no locks.The shape to hold onto: one loop does the work, the other can rewrite how the work is done, and the only thing passing between them is a plain object and a queue. To see where the meta gets its leverage, first look at what one turn of the optimize loop actually does.
anatomy of a round
Each round walks the same fixed pipeline. Scroll through it — the active step lights up, the rest fade.
one optimize round
Read the tree
An Explore agent reads the current experiment tree: the best score so far, the metric ceiling, and the open frontier — the branches still worth extending. If the ceiling's been hit or nothing's left to explore, the round stops here. Otherwise it takes the top width frontier nodes as this round's parents.
Gather signal
Scan agents comb the evaluated nodes in parallel for what's working and what's failing, while an aggregate agent looks for structural patterns across the whole tree. This phase is discretionary — the meta can switch it off when traces stop being informative, and the round briefs from prior signals instead.
Research a way out
When the search stalls — or every few commits — three research agents fire at once: one extrapolates the best branch, one dissects the failures, one reads the literature and the web. They surface directions local search alone wouldn't reach.
Write the experiments
A writer folds everything together — scan findings, structural patterns, the ideators' proposals, and any hints the meta queued since the last round — into concrete experiment briefs, then dedupes them so the round explores genuinely different ideas.
Run the lanes
One lane per brief, all in parallel. Each lane runs the same gauntlet: implement the change, pre-verify it (and revise up to three times if it's gaming the metric), run it, then post-audit the result with the verifier.
Prune, then loop
Dead lineages are pruned and notes recorded. Then the loop control: the stall counter resets only if a verified commit beat the prior best — a branch beating its own parent is progress, but not a new best. When stall hits the limit, the run ends. Otherwise the round repeats.
That pipeline runs the same way on round 1 and round 40 — which is exactly the limitation the meta thread exists to lift. So look at the meta loop itself.
the meta loop, tick by tick
The meta runs as its own loop alongside the optimizer — it observes, then acts through four channels. Scroll through one tick.
one meta tick
Self-paced, not per-round
The meta isn't called once per round — it runs on its own clock. Each tick opens with a ~5-minute wait, so it observes a handful of times during a run rather than reacting to every step. The wait is interruptible: when the optimizer finishes it drops a sentinel the meta polls every 15 seconds, so an in-flight tick exits promptly instead of stalling the run. Every tick is a fresh, stateless agent — no memory carried across ticks.
Read the run from outside
The tick takes in the experiment tree and best score, the live logs, GPU and host state, and cross-history patterns — including experiments that look stuck. It is strictly read-only: it never edits experiment code, runs training, or touches the benchmark, grader, or verifier.
Four channels out
From what it sees, the tick returns four lists — harness edits, brief hints, stops, and alerts. Most ticks are quiet; all-empty is the norm. It acts only on concrete evidence that the workflow's shape, not the experiments, is the bottleneck.
Reshape the workflow
Edits apply straight to the shared harness object — no allow-list, no caps — and take effect at the next round. This is the meta's distinctive power: retune a knob, toggle a phase, rewrite a prompt the loop uses, or inject an extra step. (The next section walks all four.)
Steer, don't reshape
Softer than an edit: hints are queued onto metaSignals and folded into the next round's brief — nudging what the loop tries next without changing the loop's structure.
Stop a doomed experiment
When the meta judges an experiment is heading nowhere, it doesn't kill it. It hands a recommendation to a separate gated enforcer that verifies the experiment is still active, then aborts it, annotates the diagnosis, and discards it with a failure class. Detect and act stay separate — never a silent kill. (Detailed two sections on.)
Surface what it can't fix
Runtime and host problems the meta can't solve itself — a dying GPU, a wedged process — go to the run log for a human. Then the tick ends, the wait restarts, and it repeats until the optimizer is done. A short streak of failed ticks disables the meta entirely: it's advisory, and must never take down a good run.
That's the outer loop. The most consequential of its four channels is the first — editing the harness — so that's next: the four edits, and exactly where each one lands in the round you just walked.