Something in your business will break in front of customers. The site will go down, an integration will silently stop syncing, a billing run will double-charge, a vendor you depend on will have a bad day and hand you their outage. This is not a risk to manage away. It is a scheduled event with an unknown date.
What varies is the cost. And the cost is set almost entirely in the first sixty minutes — not by how fast you fix the thing, but by how you behave while it is still broken. Customers forgive outages. They do not forgive being kept in the dark, told a comforting lie, or discovering the problem themselves and finding your last public word was “all systems operational.”
This is the hour, broken into moves you can run without thinking. Which is the point: nobody thinks clearly at minute four.
The three jobs, in this order
An incident has exactly three jobs, and almost every bad response comes from doing them in the wrong order or letting one person do all three.
- Contain. Stop it getting worse. Not fix — contain.
- Communicate. Tell affected people what is true, before they ask.
- Record. Capture what happened while it is happening, because you will not reconstruct it later.
The classic failure is doing job one until it’s finished and then doing job two. That feels responsible. It is how a 40-minute outage becomes a trust problem that lasts a quarter, because the gap between “customers noticed” and “you said something” is the only number they will remember.
Minute 0–5: Declare
The first decision is binary and should take under a minute: is this an incident? The test is not severity. It is this:
Is a customer, right now, having an experience we did not intend and would not want them to have?
If yes, declare. Say the word out loud in whatever channel your team lives in: “Declaring an incident. I’m Incident Lead.” That sentence does three things — it stops the ambient “is anyone looking at this?” drift, it names a decision-maker, and it starts the clock everyone will later measure you against.
Small teams under-declare because declaring feels dramatic. Invert it. Declaring costs you nothing and can be stood down in four minutes. Not declaring costs you the first twenty.
Minute 5–15: Four roles, named out loud
Even if you are a team of three, name the roles. One person can hold two. Nobody holds all four.
- Incident Lead. Makes calls, does not debug. Their only outputs are decisions and the next checkpoint time. If your best engineer is Lead, you have lost your best debugger.
- Operator. Hands on keyboard. Fixes things. Speaks only to the Lead.
- Comms. Owns everything customer-facing. Writes the updates. Never asks the Operator for an ETA more than once every fifteen minutes.
- Scribe. Timestamps everything in one thread: what was observed, what was changed, when. This is the least glamorous role and the one that pays for itself two days later.
The Lead’s first act is to set a checkpoint: “We regroup in 15 minutes regardless of progress.” Fixed intervals stop the two states that kill incidents — the silent hero going dark for an hour, and the pile-on where nine people ask for status and nobody works.
Minute 15–20: The first public message
Send it before you understand the problem. That is not a typo. The first message is not an explanation; it is proof that a human is awake.
Four elements, nothing else:
- What is affected, in customer language.
- What it means for them right now.
- What you are doing.
- When you will next speak — a specific time.
Template:
We’re aware that [thing customers actually touch] is [not working / behaving incorrectly] as of [time, with timezone]. If you’re affected, you’ll see [specific symptom]. Our team is investigating now. We don’t have a cause yet. Next update by [time], whether or not we have news.
Worked example:
Since about 09:40 ET, invoices sent from the platform are not reaching recipients. If you sent an invoice this morning, assume it did not arrive. Nothing has been double-charged and no data has been lost. We’re investigating and don’t yet know the cause. Next update by 10:15 ET.
Notice what it does not contain: an apology paragraph, a cause, a fix ETA, or the phrase “a small number of users.” Every one of those is a trap.
Four sentences to delete on sight
- “A small number of users…” — you don’t know the number yet, and the affected customer is at 100%.
- “We expect this to be resolved shortly.” — an ETA you invented becomes a promise you broke.
- “Due to an issue with our upstream provider…” — true, irrelevant, and reads as blame-shifting. Name the vendor in the post-mortem, not the first update.
- “We apologise for any inconvenience this may have caused.” — the “any” and “may” are hedges that tell the reader you don’t think it’s serious. Apologise later, specifically, once.
The cadence rule
Once you have promised a next update time, that promise outranks the fix. Comms sends at the stated time even if the entire content is: “10:15 ET. Still investigating. No cause identified yet. Next update by 10:45.”
A team that posts “no news” on schedule reads as in control. A team that goes quiet because there was nothing worth saying reads as absent. The information content is identical; the trust content is not.
Practical cadence: every 15 minutes for the first hour, every 30 after, hourly past three hours. Shorten it when the news is bad, never lengthen it silently.
Minute 20–60: Two tracks, running in parallel
Operator track. Contain first. Roll back before you diagnose — a revert you can explain later beats a fix you understand now. If a rollback is not available, degrade deliberately: turn the broken feature off, put up a maintenance state, queue the writes. A visibly disabled feature is a far cheaper outcome than one that silently corrupts data for another forty minutes.
Comms track. While the Operator works, Comms builds the blast radius list: who was affected, what did it do to them, and which ones need an individual message rather than a status post. Someone whose payroll run failed does not get told by status page. That list is the single most valuable artefact of the whole hour, and it can only be built while the evidence is fresh.
The Lead’s job across both tracks is to keep asking one unpopular question: “What’s the worst thing that’s still happening that we haven’t stopped?”
The resolution message
Do not send this the moment it works. Send it once you have watched it work for ten minutes. “Resolved” followed by “actually, still broken” costs more credibility than the original outage.
This is resolved as of [time]. [Symptom] is working normally and we’ve confirmed it over the last [N] minutes. The cause was [one plain sentence]. If you [specific action during window], here’s what you should do: [concrete instruction]. We’ll publish a full write-up by [date]. I’m sorry — this one was on us.
One apology, unqualified, at the end. Not five hedged ones throughout.
Five failure modes
- The hero. One person disappears into the problem and stops reporting. Fix: fixed checkpoints, enforced by the Lead, not by the hero’s judgement.
- Debugging in the comms channel. Technical speculation leaks into a customer-facing update. Fix: Comms owns the send button. Nobody else has it.
- The premature all-clear. Covered above. Ten-minute soak, no exceptions.
- Cause-hunting before containment. Understanding why is satisfying. It is also job three of one. Contain, then be curious.
- No scribe. Two days later you are reconstructing a timeline from memory and Slack scrollback, and the post-mortem produces a vibe instead of a rule. This is the failure that guarantees the same incident happens twice.
The 20 minutes of prep that does the work
None of the above is hard in the moment if four things exist beforehand. Do these this week, not during your next incident.
- Pre-write the first message. Put the template above in a document with the blanks visible. At minute fifteen you want a fill-in-the-blank, not a writing task.
- Decide where you speak. One canonical place — status page, a pinned post, an email list — and make sure customers know it exists before they need it. A status page nobody has bookmarked is a diary.
- Build the escalation list. Names and phone numbers for: your top ten customers by dependency, your critical vendors’ support paths, and whoever can approve a refund without a meeting. Phone numbers, because the thing that’s broken may be how you normally message people.
- Run one 20-minute dry run. Pick a plausible failure, declare it, assign roles, write the first message, stop. You will find the gap — usually that nobody knows who can post to the status page — and you will find it for free.
The teams that come out of a public failure with their reputation intact are almost never the ones who fixed it fastest. They are the ones who, at minute fifteen, said something true to the people it happened to — and then kept saying true things on a schedule until it was over.
That is a process, not a temperament. Write it down before you need it.
