Why we publish our failures
Most research quietly buries the experiments that didn't work. We publish every one, signed and dated, including the two theories we killed and the three experiments that found nothing. Here is why that is better science, not worse.
There is a quiet rot in most science, and it has a name: the file-drawer problem. An experiment that “doesn’t work” (a negative result, an inconclusive one, a refuted hunch) goes in a drawer and is never seen. Only the wins get written up. So everyone reads the successes, nobody reads the dead ends, the same mistakes get repeated across labs, and effects look far stronger and cleaner than they really are.
We do the opposite, on purpose. Every experiment we run gets a signed, numbered, dated entry in an open corpus, whether it confirmed our hunch, refuted it flatly, or landed in a shrug. This post is about why, and what it actually looks like when you commit to it.
What “publishing failures” actually looks like
Not a slogan. Concrete, from our own notebook, with the receipts:
- We contradicted ourselves, and left the trail. We reported that one memory mechanism was hopeless at holding two things at once. A week later we realised it was a testing mistake on our part, not a real limit. Insight 038 does not quietly overwrite the earlier claim; it corrects it in the open, so the whole reasoning stays visible.
- We killed our own favourite theory. We had an elegant explanation for a result and built an experiment specifically to confirm it. It refuted it flatly. Insight 040 is, in effect, titled “our nice theory was wrong.”
- We published three nothings in a row. On one question, three separate experiments (043, 044, 045) each found no difference between the things we were comparing. A conventional paper buries a single null; three would never see daylight. For us they were the finding: they revealed the thing we cared about belongs to a different regime entirely.
- We caught ourselves believing a fluke. An early result looked spectacular, from a single run. Our own rule (no claim about a rate below ten runs) forced a rerun. It did not survive. Insight 026 keeps both the exciting fluke and its quiet death.
This is not a handful of cherry-picked confessions. Here is the whole ledger:
The discipline that makes it cheap and safe
Publishing failures only works if the format removes the sting:
- Signed, numbered, never deleted. Entries are monotonic. You supersede an old one by writing a new one that references it. You never edit history.
- Evidence or it didn’t happen. Every claim cites a measurement, or a file and a line.
- Negatives are first-class. A refuted expectation gets the same care in the write-up as a confirmed one. Often it gets more.
- Report as replication, not discovery. When we reproduce something already known, we say so plainly. No inflating a re-run into a breakthrough.
Why a commons, specifically, does it this way
A company has an incentive to show only its wins: the story is the product. A commons has the exact opposite incentive. The whole point is that other people can build on ground that is actually solid. A buried negative is a trap left for the next person to step in. A published one is a gift to them.
So this is not modesty, and it is not confession. It is just better engineering, and it happens to be the only honest way to invite people in.
Come check our work
All of it is open. The corpus is signed and dated. The interactive demos let you reproduce the headline results with your own hands. The two claims we retracted are still there, labelled as such. If you go through it and find us wrong, that is not an embarrassment to us. That is the system doing exactly what it is for.
Start anywhere: the notebook, or the signed insights themselves.