There is a specific, repeatable moment that decides the fate of most work: someone with budget authority looks at what you have been doing for six months and decides, in about ninety seconds, whether it is working. They are not hostile. They are busy. They have four other things in front of them and no way to independently verify any of it.
In that moment, effort is invisible. Activity is invisible. Your deck is invisible. What survives is a number — and only if it is a number the person cannot argue with, cannot restate to mean something else, and did not have to take your word for.
Almost nobody builds that number in advance. They build it retroactively, under pressure, from whatever data happens to exist, which is why it always looks like it was built retroactively under pressure. This playbook is the other approach: pick the number on day one, wire it up before there is anything to report, and let it accumulate credibility while you work.
Part 1 — The mental model: claim, proof, noise
Every number you could report falls into one of three buckets, and knowing which is which is most of the skill.
- Claim. A number that requires believing you. “This saves the team about ten hours a week.” The estimate may be right. It is still an assertion with a decimal point on it.
- Noise. A number that is real, measurable, and does not distinguish a working project from a dead one. Signups. Seats provisioned. Meetings held. Pages shipped.
- Proof. A number that would be different if the work were not working, that someone other than you produced or could reproduce, and that moves on a timescale short enough to see.
The test for proof is a single counterfactual question: if this project were quietly cancelled tomorrow, would this number change within a month? If the answer is no, you have noise. Most dashboards are entirely noise, which is why nobody reads them.
Part 2 — The five metrics that get you laughed out of the room
These are not bad numbers. They are numbers that a skeptical person can dismiss in one sentence, and knowing the dismissal in advance saves you the meeting.
- Registrations or seats. Dismissal: “How many of them came back?” Access granted is not value received. Every enterprise tool ever bought has 100% seat provisioning and single-digit weekly use.
- Total volume of anything you produced. Reports generated, tickets touched, drafts written. Dismissal: “Would we have missed any of them?” Output is your cost, not your return.
- Satisfaction scores collected by you. Dismissal: “Who answered?” The people who fill out your survey are the people who like you. Self-collected sentiment is the weakest evidence in business.
- Cumulative anything. A line that can only go up cannot go down, so it carries no information. Dismissal: “What did it do last month?”
- Estimated savings. Dismissal: “Whose budget went down?” If no line item shrank and no headcount was redeployed, the savings are hypothetical and everyone in the room knows it.
Part 3 — The C.L.E.A.R. test
Run every candidate metric through five gates. A real proof metric passes all five. Four out of five is a supporting metric — useful, not load-bearing.
- C — Counterfactual. It changes if the work stops. (The gate from Part 1.)
- L — Lean. One number, one unit, no composite index. The moment you weight three things together, you own an argument about the weights.
- E — External. Someone other than you touches the data: a finance system, a vendor invoice, a customer action, a third-party benchmark, a log you do not control.
- A — Actual. It describes something that happened, not something projected. Past tense only.
- R — Repeated. It refreshes on a fixed cadence — weekly or monthly — so a trend exists before anyone asks for one.
The E gate does the heaviest lifting and is the one people skip. A number your own tool reports about your own tool is an opinion in a nice font. Route it through the accounting system, the CRM, the payment processor, the support queue, the vendor bill — anything with an owner who does not report to you.
Part 4 — The 20-minute selection drill
Set a timer. On one sheet:
- Minutes 0–5. Write the sentence: “This work exists so that ______ happens more, or ______ happens less.” Fill both blanks with things that happen to other people, not to you.
- Minutes 5–12. List every number that already exists in a system you can query today which touches either blank. Existing beats ideal — a mediocre metric with nine months of history outranks a perfect one starting from zero.
- Minutes 12–17. Run each through C.L.E.A.R. Cross out anything that fails two gates.
- Minutes 17–20. Pick one primary and one guardrail. The guardrail is the number that would move in the wrong direction if you gamed the primary. You will report both, always, in the same breath.
Worked examples
- You run an internal AI tool. Weak: seats activated. Strong primary: percentage of weekly active users who used it in three consecutive weeks, pulled from the auth log. Guardrail: support tickets tagged to the tool per 100 active users. The pair says “people keep choosing it and it isn’t creating a mess.”
- You are a freelancer or consultant. Weak: hours billed. Strong primary: share of this quarter’s revenue from clients who bought before. Guardrail: days from proposal sent to signature. The pair says “they come back and it isn’t a fight.”
- You lead a support or ops function. Weak: tickets closed. Strong primary: percentage of issues resolved without a second contact, from the ticketing system. Guardrail: median time to first response. The pair blocks the obvious game of closing fast to look good.
- You are trying to justify your own role in a reorg. Weak: projects delivered. Strong primary: number of recurring commitments that pass through you and their on-time rate, taken from someone else’s calendar or tracker. Guardrail: how many of them have a named backup. The second number is the one that gets you promoted rather than trapped — irreplaceable and unpromotable are the same status.
Our take: The guardrail is not a nicety. Any single metric, held up long enough, gets optimized until it stops meaning anything — that is not cynicism, it is arithmetic. Publishing the guardrail yourself, before anyone asks for it, is also the fastest credibility purchase available. It tells the room you already thought of the objection they were forming.
Part 5 — Instrument it before there is anything to show
Do this in the first week of any project, when the number is embarrassing. An embarrassing baseline is an asset; you cannot demonstrate a change from a starting point you never recorded.
- Write down today’s value and today’s date. In a file. With the query or the steps used to get it, verbatim.
- Pull the last 12 periods if they exist. Historical variance tells you what “normal noise” looks like, so that later you can distinguish a real move from a Tuesday.
- Automate the pull or calendar it. A recurring 15-minute block beats an unbuilt dashboard. The failure mode of instrumentation is a beautiful pipeline that ships in month four.
- Send it to one person immediately. A short note to your manager or client: “Baseline for this project is X as of today; I’ll send this same number monthly.” That email is your timestamp, and it is unforgeable in a way a spreadsheet is not.
Part 6 — The one-page proof memo
Six lines. Send it on the same day every month. Do not attach a deck.
- Line 1 — The number. “[Metric] is now [value], from [baseline] on [date].”
- Line 2 — The source. “Pulled from [system], same query as the baseline. Anyone can rerun it here: [link].”
- Line 3 — The guardrail. “[Guardrail] over the same window: [value], from [baseline].”
- Line 4 — The honest caveat. One sentence naming the strongest alternative explanation. “Some of this is seasonal; last year the same window rose 4%.” Volunteering the counter-argument is what separates a proof memo from a marketing memo.
- Line 5 — The decision it supports. “On this basis I’d keep funding at current level / expand to team B / stop.” Name the option you would take, including stopping.
- Line 6 — The next checkpoint. “Next reading [date]. If it is below [threshold], I will recommend we stop.”
Line 6 is the whole memo. Pre-committing to a kill threshold converts you from an advocate into an evaluator, and evaluators get believed. It also means that when the number is good, nobody has to wonder whether you would have told them if it weren’t.
Our take: The counter-intuitive move in this entire playbook is Line 4. Every instinct says to hide the alternative explanation. But the person you are trying to convince is going to think of it either way — the only variable is whether they think of it while trusting you or while auditing you.
Part 7 — Six ways a good proof metric goes bad
- Definition drift. The query quietly changes and the trend line becomes fiction. Fix: store the exact query with the baseline and diff it before every send.
- Population drift. The denominator changes — a team is added, a segment is excluded — and the ratio moves for reasons that have nothing to do with you. Fix: report the denominator alongside the ratio, every time.
- The victory lap. One great month, so you stop sending the memo. Fix: send it on the schedule regardless. A missing month reads as a bad month.
- Metric inflation. Someone asks for more context, so you add a second number, then a fourth, and within two quarters you have a dashboard nobody opens. Fix: one primary, one guardrail, in the body of an email, forever.
- Silent gaming. The team starts optimizing the metric without deciding to. Fix: the guardrail, plus one qualitative check per quarter — talk to three actual users and ask what changed.
- Orphaning. You leave, or get reassigned, and the number dies with you. Fix: the pull instructions live in a shared doc, not your head, and one other named person has run it at least once.
The objection clinic
- “My work genuinely can’t be measured.” Then measure its absence. What breaks, and how often, when you are on vacation? That is a countable event in someone else’s system.
- “The number won’t move for a year.” Then your primary is the leading indicator and the year-long outcome is a footnote. Pick the earliest observable behavior on the causal chain.
- “It might make me look bad.” It might. But a metric that can only flatter you is one nobody believes, which means it cannot help you either. Downside risk is what makes upside readable.
- “Nobody has asked me for this.” Correct. They ask during budget season, and by then the answer takes six months of history you did not collect.
Your first week, hour by hour
- Monday, 20 minutes. The selection drill. Pick a primary and a guardrail.
- Tuesday, 45 minutes. Pull today’s value and whatever history exists. Save the query verbatim.
- Wednesday, 10 minutes. Send the baseline email to one person. Put the monthly reminder on your calendar in the same sitting.
- Thursday, 30 minutes. Write the pull instructions in a shared doc and ask one colleague to run them once, so you learn what is ambiguous.
- Friday, 15 minutes. Draft the six-line memo with placeholder values. Next month it takes eight minutes.
Two hours, once. Against the alternative — assembling evidence retroactively while somebody decides your project’s future in a meeting you are not in — it is the highest-leverage two hours on your calendar this quarter.
