Research · Research
EvoLab
A re-implementation of the architecture behind FunSearch and AlphaEvolve — a language model as mutation operator, a hard evaluator as selection. Local, without API keys.
The claim “nine AI agents optimise themselves” could not be substantiated. So we rebuilt the substantiated architecture: a language model proposes program variants, a deterministic evaluator decides what survives. The result is more honest than the claim.
The problem
Self-improving systems are surrounded by figures nobody can re-check. We wanted to know which component does the work — and found: not the agent swarm, but the evaluator that the candidate cannot trick.
How it works
- 01 Several sub-populations of program candidates with occasional exchange, so that the search does not freeze on one solution too early.
- 02 The local model sees two predecessors sorted by quality and is asked to write a better third version — it sees a direction, not just an example.
- 03 The evaluator is deterministic, runs in its own process with a time limit and cannot be influenced by the candidate.
- 04 Task: beat a known packing-problem heuristic. A positive value means: better than the reference.
- 05 Your own tasks by swapping only the evaluator.
Sovereignty & evidence
- Runs entirely locally via Ollama on a single machine. No cloud service, no API key, no data leaving the machine.
The honest assessment
A local run on a desktop machine is not in DeepMind’s league. That was never the goal. The goal was to understand which component does the work. Answer: the evaluator. Everything else — the number of agents, the size of the model — is secondary as long as the selection is hard.
What comes of it
Nothing saleable. But a rule that applies in every customer project: whoever builds AI agents first builds the yardstick they are measured against.
Frequently asked questions
- Has EvoLab discovered anything new?
- No. What is substantiated is two short runs with one improvement found over a known heuristic. That is a working setup, not a research result. The well-known results of FunSearch and AlphaEvolve come from DeepMind, not from us.
- What does this have to do with customer projects?
- The lesson from it shapes every one of our systems: build the evaluator first, then the agents. An agent without a hard yardstick optimises into the void.
A comparable challenge on your desk?
Send us three sentences. We reply within one business day with an honest assessment — including when we are not the right partner.