How to pre-register a prediction so you can't move the goalposts
- Reading time
- 7 minutes
- Assumes
- You make decisions somebody later evaluates
- Updated
- Sep 6, 2026
The problem with explaining afterwards
Your team ships a change. Two months later the numbers are in. Somebody assembles a story about what happened, and the story fits.
It always fits. That is the problem. Given any outcome, a competent team can construct a plausible account of why it occurred, and having constructed it, will believe it. The account gets used to justify the next decision, which is how organizations accumulate confident folklore instead of knowledge.
You cannot fix this by being more rigorous after the fact, because the contamination happens at the moment you see the result. The only intervention that works is moving the claim earlier, to a point where you genuinely do not know.
What pre-registration actually means
Medicine hit this problem in the 1990s. Trials would measure fifteen outcomes, report the two that reached significance, and describe those as the goal all along. The fix was procedural rather than statistical: register what you are measuring and what counts as success, publicly, before you collect data.
The business version is smaller and needs no registry. Before the change ships, write down four things.
What you're changing. One or two sentences, specific enough that somebody could tell whether it happened.
What you expect to move. A named number, and the direction. "Support cost per ticket" rather than "efficiency."
By how much, and by when. A threshold and a date. "Down at least 8% by March 31." Without a threshold you will accept any movement as confirmation.
What would make you wrong. The single most useful line, and the one people resist writing.
That's it. It takes ten minutes and it is the difference between a team that learns and a team that narrates.
If you write one line
Write the threshold. A prediction without a number is not a prediction, it is a hope, and it will be satisfied by whatever happens.
Predict what might get worse, too
A prediction naming only the number you hope will rise is half a prediction. Every meaningful change has a cost somewhere, and the team that only wrote down the upside will find the downside later, usually from a customer.
Name the things you are willing to break and the tolerance you will accept. "Resolution time down 8%, and satisfaction holds within two points of where it is now." The second clause is what stops a win on the first from being reported as an unqualified success.
This is standard practice in online experimentation, where it goes by guardrail metrics, and it transfers cleanly to decisions that were never an A/B test.
Where teams cheat without meaning to
Nobody sets out to move a goalpost. It happens through four ordinary moves, each of which feels reasonable at the time.
Switching the outcome. The metric you named didn't move, but a different one did, and the writeup is about that one. This is the original sin pre-registration was invented to stop.
Extending the window. The deadline arrives, the number is flat, and somebody suggests it needs another month. Sometimes true. Write the extension down as an amendment with its reason, so the record shows the deadline moved.
Reinterpreting the threshold. You said 8%, you got 3%, and the writeup says "moved in the right direction." It did. It also missed.
Explaining away the guardrail. Satisfaction dropped four points, outside your two-point tolerance, and the explanation is seasonal. Maybe. Record it as a miss with the seasonal note attached, and see whether the same explanation is needed next quarter.
Common mistake
Editing the prediction after evidence lands, with good intentions, because the original was poorly worded. If the wording was genuinely wrong, supersede it with a new entry that points at the old one. An edit leaves no trace and makes every other prediction in the record less trustworthy.
Settling honestly
At the deadline, three outcomes are possible and all three are useful.
Met. The number crossed your threshold. Record what you now believe, in one line, and be specific about what it licenses you to do next.
Not met. Also a result. The most valuable entries in any decision record are the confident predictions that failed, because those are the ones that correct a model of the business.
Mixed. You made several claims and they disagree — the target moved, the guardrail slipped. Resist collapsing this into a win or a loss. Mixed is information, and forcing it into a binary throws away the part you'd most want to remember.
Write the lesson at settlement, never before. A lesson written while evidence is still arriving gets treated as settled by everyone who reads it later, including you, and the remaining evidence gets read against it.
What you get after a year
The individual entries are useful. The aggregate is what changes how a team operates.
Once you have thirty settled predictions you can ask a question almost no organization can answer: how often are we right? Not in general — for this kind of decision, made by this team, at this confidence. Most teams discover they are well calibrated on operational changes and badly overconfident on anything involving customer behavior, which is worth knowing before the next roadmap.
Keep this at the level of the workspace. Scoring individuals turns a learning instrument into a performance review, and the moment people believe their predictions are being graded, the predictions get safe and the record stops being worth keeping.
Before you ship the change
0 of 6 checked