We run post-mortems on failures and victory laps on wins. It is exactly backwards. A failure costs you once. An unrepeatable success costs you every time you scale it — the headcount you hire against it, the quarter you forecast on it, the channel you go all-in on, the training block you build around a session that happened to land on a good day.
The expensive mistake is not believing a fluke. It is announcing one. Once a number is in a plan, in a board deck, or in your own story about yourself, walking it back costs more than the number was worth.
So before you build on a good result, run it through this. It takes an afternoon to set up and one cycle to answer.
The one distinction that does all the work
Every result is a mix of three things:
- Skill — what you did. Repeats whenever you do it.
- Conditions — what was true around you. Repeats only when the conditions repeat.
- Variance — noise. Does not repeat, by definition.
A single good outcome tells you the total. It tells you nothing about the split. And the default human move is to silently assign the whole thing to the first bucket, because that is the flattering one and the one that comes with a story.
The Repeatability Test is just a structured way of forcing the split into the open before you spend money on it.
Our take: The number you should plan with is never the first result. It is the second one. First results are estimates with a sample size of one and a strong upward bias — you noticed this outcome precisely because it was unusual. The second result, run under conditions you deliberately changed, is the first honest number you have. Everything in this playbook is machinery for getting to that second number cheaply, before someone builds a budget on the first.
Step 1 — Write the claim down before you look again
Ten minutes, one sentence, in this shape:
“X caused Y. If I do X again under conditions C, I should get at least Y′.”
Three rules for the sentence:
- Y′ is a floor, not a range. “Somewhere between 5,000 and 40,000” is not a prediction, it is an alibi. Pick the number below which you would admit you were wrong.
- C is written out explicitly. Same audience? Same season? Same person doing the work? Same one enthusiastic customer? If you cannot list the conditions, you do not yet know what you are claiming.
- It goes in writing, with a date, before the next run. Not because you are dishonest — because hindsight will explain any outcome you get, and a written prediction is the only thing hindsight cannot edit.
If you cannot write a falsifiable sentence, that is the finding. You do not have a result yet, you have an anecdote you enjoyed.
Step 2 — Count your real n, and find the base rate
The felt sample size is almost always bigger than the real one. Three wins with the same client in the same quarter is not n = 3; it is roughly n = 1, because the thing you are worried about — that the client, not the method, was doing the work — is held constant across all three.
Rough ladder:
- n = 1: anecdote. Useful as a hypothesis, worthless as a plan.
- n = 2–3, deliberately varied: weak signal. Enough to invest your own time.
- n = 5+, varied: a real effect. Enough to invest someone else’s money.
Then find the base rate, which is the step almost everybody skips. What fraction of attempts like this succeed for anybody? If roughly one post in twenty goes 10x for everyone in your category, and one of your twenty posts went 10x, you have learned something about the category and nothing about yourself. Base rates turn “this worked” into “this worked more often than chance,” which is the only version worth acting on.
Step 3 — Change exactly one condition, then rerun
This is the whole test. Everything else is bookkeeping.
Pick the single condition you most suspect was doing the work — not the one that is easiest to change. Be honest here; the suspicious condition is usually the one you have been carefully not thinking about. Change that one. Hold everything else fixed.
| If you suspect… | Vary this | Hold this |
|---|---|---|
| One motivated buyer | The customer | Offer, price, pitch |
| One unusually good week | The timing | Method, effort, team |
| One talented operator | The person | Process, tooling, brief |
| One big account amplified it | Distribution | The content itself |
| One easy set of inputs | The test cases | The system, the prompt |
Two practical notes. First, the cheapest replication is usually smaller, not bigger. You do not need to rerun the whole campaign; you need one honest instance under a changed condition. Second, if you change three things at once and the result holds, you have learned nothing you can name — and if it collapses, you will not know which change killed it.
Step 4 — Score it against three verdicts
Compare the second result to the Y′ you wrote down in step 1.
- Repeats (80%+ of Y′). Skill, or a condition durable enough to count on. Scale it — and write down which conditions you are now dependent on.
- Partially repeats (roughly 30–80%). The most common and most useful verdict. The effect is real and smaller than you thought, or it is condition-dependent. Do not scale yet. Go back to step 3, vary a different condition, and find out which one carries the weight.
- Does not repeat (under 30%). Variance, or a condition you have now lost. Bank the win, take the money, and do not build on it. Then ask the genuinely interesting question: what was unusual about the first run? That answer is often worth more than the win was.
Write the verdict next to the original claim. Two lines. This file becomes the most useful document you own within a year, because it is a record of how calibrated you actually are — which is a skill that does compound.
Step 5 — Size the bet to the verdict
The rule: never commit more than the evidence would let you lose gracefully.
- One result: spend your own time. Nothing else. Do not announce it.
- Two results, varied: spend money you have already written off in your head. Tell your team it is a test, in those words.
- Three or more, varied: now you can hire against it, forecast it, put it in the plan, say it out loud to people who will remember.
Most damage happens at rung one, where the cost of the bet is small but the cost of the announcement is not. Reputational commitment is a bet too, and it is the one people forget to size.
Five worked examples
1. The post that went 20x. 41,000 impressions against a 2,000 median. Claim: “The teardown format caused this; another teardown should clear 8,000.” Suspect condition: a large account reshared it in hour two. Vary distribution — publish the next teardown with no outreach. Result 3,100. Verdict: partially repeats. The format is worth roughly 1.5x; the reshare was worth the other 13x. Correct action is to invest in relationships with resharers, not in more teardowns.
2. The 140% quarter. Claim: “The new discovery script drove it; next quarter clears 110%.” Real n: one quarter, one rep, one deal worth 60% of the overage. Strip the whale — the rest of the quarter came in at 88%. Verdict: does not repeat as stated. The honest plan number is 88–95%, and the script is untested because one deal drowned out the signal. Rerun on the other three reps.
3. The new hire’s first project. Shipped in three weeks against a six-week estimate. Suspect condition: it was a project she had built twice before elsewhere. Vary the domain, hold the process. If the second project lands at four weeks against six, that is speed. If it lands at seven, you hired for prior art, which is fine — you just need to staff her accordingly.
4. The training PR. A lift or a time that jumps well beyond trend. Suspect conditions: nine hours of sleep, a deload week before, a training partner, ideal temperature. Vary one — repeat it on a normal Tuesday. A PR you can hit on a normal Tuesday is fitness. A PR that needs four aligned conditions is a peak, which is a real and useful thing, but you cannot program from it. See the Peak Day Playbook for building those conditions deliberately.
5. The AI workflow that got 9 out of 10. Suspect condition: you chose the ten test cases, after building the thing. Vary the inputs — have someone else pull ten real cases at random, including the ugly ones you avoided. A drop from 9/10 to 6/10 is the normal, healthy result and tells you exactly where the work is. Build the metric properly before you run this one.
Three scripts
To your boss, before the number becomes a plan: “July came in at 140%, and I want to flag how much of that is repeatable before we build Q4 on it. One deal is 60% of the overage. Stripping it, the underlying run rate is about 90%. I’d plan on 95 and treat anything above as upside. I’m rerunning the new script with the other three reps this month to see if it’s real.”
To a client, asking for the second pilot: “The pilot hit the target, and I don’t want to oversell one run. The conditions were unusually favourable in two ways — your team was fully available and we picked the cleanest data set. I’d like to run it once more on a messier slice before we scope the rollout. Same fee structure, half the duration. If it holds, you buy with real numbers. If it doesn’t, you’ve saved a rollout.” This costs you a fortnight and wins you the deals that would otherwise have died in month three.
To yourself: “What would have to be true for this to be luck?” Then go looking for those things with the same energy you would use to defend the win. If you cannot find any, you have earned some confidence. If you find three in five minutes, you already knew.
Seven ways this goes wrong
- Writing the prediction after the second run. The single most common failure. It converts the test into a story.
- Varying everything. A rerun with a new person, new market and new offer answers no question at all.
- Rerunning too big. If the replication costs more than the original, you will not run it, and you will scale on n = 1 instead.
- Treating “partially repeats” as a failure. It is the most informative verdict available. It means there is a real effect and you have not yet found its boundary.
- Only testing wins. Bad results are just as often variance. A method binned after one poor run is a real cost, and nobody audits it.
- Ignoring the base rate. If everyone gets this result one time in twenty, so did you.
- Running the test after the announcement. By then you are not measuring; you are hoping. Decide first, then commit — see the Decision Playbook for the order of operations.
The one-page version
- Write the claim: X caused Y; repeating X under C should give at least Y′. Date it.
- Count your real n. Find the base rate.
- Change the one condition you most suspect. Hold the rest. Rerun small.
- Score: repeats / partially repeats / does not repeat. Plan with the second number.
- Size the bet: own time → written-off money → headcount and forecasts.
An eight-second win gets argued about. A 1:18 win does not. That is not a metaphor about margins — it is the actual mechanism. A result that repeats under conditions nobody controls ends the conversation about whether it was real, and no amount of arguing about the first result ever does. Build the second data point before you need it.
