Abandoning the objective: when chasing the goal is how you fail
Some problems are traps: aiming straight at the goal walks you into a dead end, and a stronger optimiser just gets there faster. We built one such trap on purpose and watched a search that ignores the goal entirely walk out of it — after a bug nearly convinced us otherwise.
Here is an idea that sounds wrong the first time you hear it: for some problems, the surest way to fail is to aim directly at the goal. Not because you are lazy or slow, but because the landscape is a trap — every step that looks like progress walks you deeper into a dead end. Push harder and you just hit the wall faster.
If that is true, the fix is not a stronger engine. It is a different thing to want. Instead of rewarding getting closer to the goal, you reward doing something new. Chase novelty, ignore the objective, and — the claim goes — you wander out of the trap the goal-chaser is stuck in. This is a famous idea in evolutionary computing (Lehman & Stanley called it “abandoning objectives”). This programme set out to see it happen for ourselves, on our own engine. It took two experiments, and a caught mistake, to do it honestly.
First, a warm-up that told us the task matters
Before testing traps, we built the machinery on a friendly problem: balancing a pole on a cart. The novelty-style method here is MAP-Elites, which does not chase one best answer — it fills a shelf with many kinds of answer (balancers that lean left, lean right, drift up-track, down-track). At an equal compute budget, it did fill that shelf better than plain random tinkering:
That warm-up also handed us our first lesson in self-doubt. A first pass had screamed “diversity beats goal-chasing twenty-fold!” — but the reviewer whose job is to attack our work found it was comparing against a broken baseline (a single greedy run that stalls badly), while random tinkering already aced the task. We retracted the twenty-fold claim before it was ever signed. Coverage was the real, modest win; the headline was an artifact.
Then we built the trap
So we built a small maze designed to be deceptive. Start at one corner, goal near another. The catch: a barrier sits so that the cells closest to the goal form a cul-de-sac. To actually reach the goal you must first walk away from it, around the barrier. A searcher rewarded for “get nearer the goal” is lured straight into the dead end and cannot leave — every exit move scores worse.
We ran five searchers on it, each given the same mutation and the same budget, forty independent runs apiece:
Read the bottom two bars against the top two and the old idea is right in front of you. The searcher that ignores the goal walks out of the trap the goal-chaser is stuck in. And note the middle bar: a stronger optimiser did no better than the weak one — worse, even. Deception is not a horsepower problem. A better goal-chaser just reaches the dead end sooner. The only thing that helped was wanting something different.
The bug that nearly hid it
The honest part: our first run of this said the opposite. Novelty scored a miserable 2 out of 40 — worse than random. We could have signed “novelty does not help on deception” and moved on. The reviewer read the code first and found the flaw: the routine that is supposed to reward being different from the crowd had a bug that quietly cancelled its own anti-sameness pressure, and we had rewarded the wrong kind of “new” (novel wiggles that never leave the trap’s basin). Fix the bug, reward new end-spots instead of new paths, and the 2/40 became 34/40. The result did not change because we wanted it to. It changed because the first measurement was broken, and the check caught it before it hardened.
That is the same story as the mind-reader post, and it keeps happening for a reason: the reviewer’s only job is to try to refute our results, in advance of signing. It has now flipped or killed a result eight times this season. A false negative — nearly throwing away a true finding — is exactly the kind of mistake that is invisible unless something is built to look for it.
Where it lands
Programme 4 asks what happens when you change what fitness rewards. The answer, demonstrated on our own engine: when a landscape is deceptive, changing the selection pressure — reward novelty, not proximity — solves what no amount of optimiser strength can. Which kind of novelty you reward matters too (end-spots beat paths here), a knob for the next experiments. And the open question left standing: when the goal-chaser can solve a task, does novelty cost more tries to get there? That efficiency question is next.
The sequel: does wanting-something-different cost you when chasing the goal already works?
That was the open question, and we answered it. We built the non-deceptive twin maze, where chasing the goal wins outright, and asked the obvious “no free lunch” question: surely rewarding novelty instead of proximity costs something here? A tax in wasted tries?
It didn’t show up. Novelty reached the goal about as fast as the goal-chaser (neither meaningfully slower), both far quicker than random tinkering, and along the way novelty explored almost the entire maze while the goal-chaser walked a narrow line to the exit.
Two honest cautions. We could not prove the two are exactly equal in speed (our measurement stayed a little too fuzzy for that even after tripling the runs), so we sign the absence of a penalty, not perfect parity. And the coverage win is flattered by a small maze that novelty can fill completely. But the headline holds: here, wanting-something-different cost nothing and bought a lot. Which is the perfect setup for the next programme, where the world stops having a single goal — and coverage stops being a bonus and becomes the whole point.
Read the rigorous version
Every number here is a signed, dated entry in the open corpus — including the retracted twenty-fold claim and the buggy first run, both kept on the record. Follow the sources under this post, or start with the Programme 4 charter and the signed insights index.