Skip to content
pioneerdesk.

Research · Research

EvoLab

A re-implementation of the architecture behind FunSearch and AlphaEvolve — a language model as mutation operator, a hard evaluator as selection. Local, without API keys.

The claim “nine AI agents optimise themselves” could not be substantiated. So we rebuilt the substantiated architecture: a language model proposes program variants, a deterministic evaluator decides what survives. The result is more honest than the claim.

The problem

Self-improving systems are surrounded by figures nobody can re-check. We wanted to know which component does the work — and found: not the agent swarm, but the evaluator that the candidate cannot trick.

How it works

  1. 01 Several sub-populations of program candidates with occasional exchange, so that the search does not freeze on one solution too early.
  2. 02 The local model sees two predecessors sorted by quality and is asked to write a better third version — it sees a direction, not just an example.
  3. 03 The evaluator is deterministic, runs in its own process with a time limit and cannot be influenced by the candidate.
  4. 04 Task: beat a known packing-problem heuristic. A positive value means: better than the reference.
  5. 05 Your own tasks by swapping only the evaluator.

Sovereignty & evidence

  • Runs entirely locally via Ollama on a single machine. No cloud service, no API key, no data leaving the machine.

The honest assessment

A local run on a desktop machine is not in DeepMind’s league. That was never the goal. The goal was to understand which component does the work. Answer: the evaluator. Everything else — the number of agents, the size of the model — is secondary as long as the selection is hard.

What comes of it

Nothing saleable. But a rule that applies in every customer project: whoever builds AI agents first builds the yardstick they are measured against.

Frequently asked questions

Has EvoLab discovered anything new?
No. What is substantiated is two short runs with one improvement found over a known heuristic. That is a working setup, not a research result. The well-known results of FunSearch and AlphaEvolve come from DeepMind, not from us.
What does this have to do with customer projects?
The lesson from it shapes every one of our systems: build the evaluator first, then the agents. An agent without a hard yardstick optimises into the void.

A comparable challenge on your desk?

Send us three sentences. We reply within one business day with an honest assessment — including when we are not the right partner.