What the July outage showed about manual fallback

On Friday, July 19, a faulty content update to CrowdStrike's Falcon software crashed Windows machines around the world, and Microsoft estimated that about 8.5 million devices were affected. Banks, airlines and broadcasters were hit, along with some payment terminals.

For a payments or fintech operation, a day like that tests something most continuity plans cover in a single line. Your core platform may run on servers that never noticed. The people who operate it still log in from company laptops, and the vendors they call during an incident are in the same position. When that layer fails, you find out how much of the operation your team can run by hand.

What going manual means for a payments team

Business continuity plans often say something like "revert to manual processing" and stop there. That phrase covers a lot of different work. In a payments operation it might mean releasing high-value payouts through the sponsor bank's portal, answering customers without the CRM, or deciding what to do with outbound payments that can't be screened because the sanctions tool is unreachable. The answer to that last one is usually to hold them, and the plan should say so in writing so nobody has to make the call under pressure.

For each process, someone needs to write down four things: the manual version of the step, who does it, which device and credentials they use, and how much volume it can handle.

Volume is where most plans are thinnest. Say a PSP normally releases 600 merchant payouts on a Friday, and the manual route is the bank portal with maker-checker approval, taking about five minutes per payment for a pair of people. Two pairs clear roughly 24 an hour, or fewer than 200 over a working day. (The figures are illustrative; your own will differ, but the gap is usually of this kind.) So someone has to decide in advance which payouts go first, perhaps those with contractual settlement dates, then the largest merchants. Making that decision at 11am on the day, with merchants already calling, goes badly.

Has anyone done it recently

The test of a fallback procedure is whether someone has run it in the last twelve months, with the people who would be on shift, against a clock. Running it tends to surface problems nobody would find by reading the document.

The hardware tokens for the bank portal are assigned to two people, and one of them left in March. The break-glass admin account exists, but its password is in a password manager that only opens on a corporate laptop. The processor's "incident line" turns out to be a general support inbox with a two-day response target. The printed copy of the procedure, if anyone printed it, is in a drawer in an office nobody visited that Friday.

Use the people who would actually be on shift for the rehearsal. Managers who wrote the plan know what each line means, which makes them poor testers of it. Pick an ordinary weekday, set a time limit, and write down every point where someone had to stop and ask a question.

Who tells customers, and how

Communication tends to fail in the same outage, because the tools you use for it share the dependency. If everyone reaches Slack or Teams from the laptops that just crashed, you need another channel ready: a call tree, or a group on a messaging app on personal phones, with the numbers kept current.

Customer and merchant messaging needs the same preparation. Host the status page outside your main infrastructure and make sure at least two people can update it from a phone. Write the first two or three messages now ("we are aware, payouts are delayed, next update at 14:00") and get compliance to approve them now, so support is not waiting on sign-off while the queue grows. Your sponsor bank or program manager will want to hear from you early, and your agreement with them may say how early. Find that clause before you need it.

If the manual procedure needs a laptop, a password manager and a messaging app, check that none of them depend on the thing that just failed.

Every action taken by hand also needs a record. Payouts released through a portal, refunds approved over the phone and balance adjustments made to calm a customer will all have to be matched later. If the procedure doesn't say where those actions get logged, the outage turns into weeks of reconciliation breaks that nobody can explain.

Look for the shared dependency

The July failure came from one piece of software installed on a very large number of machines. Your own version is probably smaller and easier to miss: one identity provider in front of every system, one laptop image across the whole operations team, one vendor whose support desk is the only route to a fix. Go through the processes you listed above and ask, for each one, what single thing would take out both the normal route and the fallback. Where the answer is the same for both, the fallback exists in name only.

The same question applies to your vendors. Ask your processor, your KYC provider and your core platform what happened to them on July 19, how long they took to recover and what their staff did in the meantime. Their answers will tell you more about their resilience than the questionnaire they filled in last year.

While the day is still fresh, pick the processes you would least like to lose for a full business day, probably payouts and customer contact to start with. Run each one by hand for an hour with the people who would be on shift, and fix the questions they raise before you rewrite anything else in the plan.

Discuss your operations

More insights

All articles