Post-incident reviews that change something

Almost every operations team we meet runs post-incident reviews. Far fewer can point at a specific change that happened because of one. The difference is usually in how the actions were written and who was expected to check them.

Imagine the review for an incident that starts on a Monday morning. The processor's daily settlement file arrives four hours late, reconciliation starts late as a result, and a batch of merchant payouts misses the sponsor bank's cut-off. The review document is neat. It has a timeline, a root cause recorded as "vendor delay", and three actions: improve monitoring, follow up with the vendor, update the runbook. No owners, no dates. Six months later the file is late again, the payouts miss the cut-off again, and the new review carries the same three actions.

Nothing in that document was wrong. It was too general to change anything, and nobody was responsible for checking that it had.

Blameless and specific

Blameless reviews exist so that people describe what they actually saw and did. The failure mode is that blameless slides into vague, because naming the precise step feels uncomfortably close to naming the person who took it. You can be exact about the system, the time and the decision without making anyone the cause.

Compare two lines from a timeline. "The on-call analyst missed the alert" names a person and explains nothing. "The late-file alert went to a chat channel that has been muted overnight since the October rota change" names nobody and points straight at a fix. The second one also tells you to go and look at what else is routed to that channel.

  1. What did we expect to happen, and when

    Write the normal path with its times: the file lands by 05:00, reconciliation finishes by 08:00, payouts go out before the bank's cut-off. The distance between that and what happened is the incident.

  2. How did we find out

    Record whether monitoring caught it or a customer did. If the first report came from outside the team, detection becomes an action item in its own right, separate from the cause.

  3. What slowed the response

    Look for waiting: for access, for a decision nobody felt entitled to make, for the one person who knows the settlement tool. Note whether anyone opened the runbook, and if they did not, why.

  4. Has this happened before

    Search earlier reviews for the same component or the same cause before writing any actions. If there is a match, start from the actions that were closed last time.

  5. What changes, who owns it, by when

    Each action gets a named person, a date and a definition of done that a third party could verify.

Actions someone can finish

"Improve monitoring" cannot be finished, so it never is. Rewrite it until it can be: alert the payments on-call phone if the settlement file has not arrived by 06:00 Gulf time, owned by the payments operations lead, due 31 January, done when a simulated late file triggers the alert in the test environment. That version is small, dull and checkable by someone who was not in the room.

Keep investigations separate from fixes. "Find out why the processor's file was late" is an investigation. Its output is more actions, and it still needs an owner and a date, otherwise it drifts until the next incident makes it urgent again.

Keep the list short. A review with fifteen actions usually completes none of them, because nobody can hold fifteen items open alongside their normal work. Four actions that close are worth more than twelve that age in a backlog.

When the cause sits with a vendor, the action on your side is still yours. A contract conversation with a date belongs to you, and so does a monitoring change that warns you earlier while there is still time to act. "Follow up with the vendor" belongs to nobody, which is why it survives unchallenged in so many reviews. The rehearsed manual fallback we described after the July 2024 outage is the same discipline applied to a smaller failure.

Track them where the rest of the work lives

Actions that stay inside the review document are forgotten by the following week. Put them in the same tracker as everything else the team does, tagged to the incident, and review the open list at a standing monthly meeting. The operational risk meeting is a natural home, since the people who can unblock an action are usually in the room.

Report two numbers there: actions overdue, and actions closed since the last meeting. Close an action only against evidence: the test result, or the link to the updated runbook. An action closed without evidence gets reopened, and after that happens once, the team stops closing things hopefully.

Where the action updates a runbook, reference the incident in the runbook's change log so the next person on call can see why the step exists. Steps with no visible reason are the ones people skip at 3am, which is part of why we argued for runbooks people open during an incident.

Look for the repeats

Tag every incident with the component involved and the kind of cause. Once a quarter, sort by tag and read down the list. Repeats surface quickly, and they are the most useful output a review program produces.

If an incident comes back after its actions were closed, review the actions before you review the incident.

A repeat means the actions addressed the wrong thing, or they were closed without being done. Either answer changes what the new review should contain, and it is a more honest starting point than writing the timeline again from scratch.

One exercise for this month: pull every review from last year, list the actions still open, and check how many have an owner who is still in the team. Reassign or close each one before the first review of 2026 lands on top of them.

Discuss your operations

More insights

All articles