Writing runbooks people open during an incident

At 3:12 in the morning, the on-call engineer's phone goes off: authorization approvals on the card program have dropped by a third in ten minutes. There is a runbook for this. It is fourteen pages long, lives in a wiki space she has never opened, and begins with two paragraphs on the history of the processor integration.

She doesn't open it. She messages the colleague who fixed this last time and waits to see if he is awake. By the time he answers, the incident is forty minutes old and support has a queue.

The runbook was written by the person who built the integration and reviewed by someone checking that a runbook existed. Nobody had read it in the situation it was written for. Almost everything that makes a runbook useful follows from writing it for that third reader.

Write for the first five minutes

The first screen, meaning whatever appears on a laptop without scrolling, carries most of the weight. Someone who has just woken up decides within a few seconds whether a document is going to help. If the opening answers the questions they have right now, they keep reading. If it opens with background, they close it and start typing in the incident channel.

We put four things on that screen, in this order.

  1. What this looks like when it starts

    The alert name and the symptoms a person would notice, so the reader can confirm within seconds that they are in the right document.

  2. What is safe to do immediately

    One or two actions that cannot make anything worse, such as checking the processor's status page, pulling the decline reasons by issuer, or opening the incident channel.

  3. Who to call, and when

    The escalation path with rota links, plus a time limit. If the cause is not identified within fifteen minutes, page second line. Without the time limit, people wait far too long out of politeness.

  4. How to tell it is getting worse

    A threshold that changes the response, for example declines spreading from one BIN range to all of them, or the queue of pending settlements passing a size you have agreed in advance.

Everything else, including the architecture, the integration history and the full table of error codes, goes further down or behind a link. A long runbook with a good first screen works fine, because the reader only goes deeper when the simple checks have failed.

Decision points, written as questions

Most runbooks are written as a sequence of steps, and the hard part of an incident is usually choosing between them. Write the choices as questions with a short answer and a destination. Are the declines concentrated in one BIN range? If yes, go to section C. Is the processor's status page showing an incident? If yes, raise a ticket with them and move to the customer communication steps.

Each decision point should also say who is allowed to make the call. Switching traffic to a backup processor, pausing merchant payouts and telling the sponsor bank are decisions an on-call engineer may or may not be authorized to take at 3am. The runbook is the right place to settle that in advance. If the answer is "the head of payments", put the phone number in the document and add a rule for what happens if they don't pick up within ten minutes.

Pre-authorizing a small number of actions saves more time than any amount of formatting. An engineer who knows she may pause payouts for up to an hour on her own authority does it at minute five, with a note in the channel, instead of waiting for someone to wake up and agree.

An owner and a date at the top

Every runbook needs a named owner, a last-reviewed date and a last-tested date, shown where the reader sees them. A reader who knows the document was last tested eighteen months ago treats its dashboard links with suspicion from the start, which beats discovering the problem halfway through an incident.

Runbooks go stale in predictable ways. A dashboard gets renamed, the processor changes its support number, the person listed as escalation moves to another team. Tie the review to those events where you can. When someone leaves the on-call rota, their name comes out of every runbook that week. When you change a vendor, the runbooks that mention it go on the change ticket as a task with an owner.

Put them where the alert already is

Nobody searches a wiki at 3am. The link to the runbook belongs in the alert itself, so the message that wakes someone up also tells them where to start. If your alerting tool supports it, add the link to every alert definition and treat an alert without one as unfinished work.

Keep a copy somewhere that does not depend on the systems the runbook is about. The outage in July was a reminder that laptops, single sign-on and the wiki can all be unavailable at the same moment. An exported set of runbooks for your most serious scenarios, stored where the on-call team can reach it from a phone, costs very little to keep current.

Drills, and what they find

Run a drill each quarter for the scenarios you would least like to face. Start with a tabletop: someone reads out the alert, and the newest person on the rota talks through what they would do using only the runbook. Every time they hesitate, ask what they were looking for and write it down. That list is the edit.

Later, run the same scenario in a test environment with a timer. What drills usually find is small: a broken link, a missing permission, a step that assumes access the on-call engineer does not have, a threshold that made sense before last year's volume growth. Make the fixes the same week, while the reasons are still obvious, and update the last-tested date. If a drill turns up nothing at all, you probably ran it with someone who already knew the answer.

Discuss your operations

More insights

All articles