Handing a project to the team that will run it

Project plans tend to stop at go-live, and the trouble tends to start a few weeks after it, once the project team has moved on and the people running the new process find out what nobody told them.

Most organizations treat the handover as an event: a meeting near the end, and a slide with a RACI on it. We think it works better as a set of conditions that operations defines early in the project and signs off at the end, with a hypercare period in between that closes when the conditions are met. Whether that happens on the planned date is a separate matter.

Readiness as acceptance criteria

Projects already have acceptance criteria for the build: the system meets the specification and UAT is signed off. Operations readiness deserves the same treatment, written down early enough that the project can plan and budget for it. The details depend on the project, but a payments go-live usually needs at least these:

  • Every new exception type has a queue and a named owner, and the first response is written down. If the new processor can return a status the old one never did, someone should know where it lands and what to do with it.
  • The operations team has run the daily process without the project team in the room, for at least a week of parallel run or UAT, including the end-of-day reconciliation.
  • The reports operations needs on day one run on schedule and have been checked against source data at least once.
  • Access and approval roles sit with named operations staff, and the project team's admin accounts have a removal date.
  • Alerts go to the operations on-call rota, and someone has triggered a test alert and watched it arrive.

The head of operations, or whoever will be accountable for the service, signs these off, and that signature is a condition of go-live. If the project can go live without it, the criteria get traded away in the last fortnight, when the deadline is close and every item starts to look negotiable.

Training on the work the team will do

Training delivered as a sandbox demo two weeks before go-live has mostly faded by the time anyone needs it. People remember what they did with their own hands, on cases that looked like their real ones. So run the training inside UAT or the parallel run where you can, with operations staff executing the test scripts themselves, and weight it toward the exception paths. A clean payment takes five minutes to learn. A partial refund against a transaction that has already been disputed takes longer, and it is the case that will arrive late on a Friday.

Train the people who will train the next hire, too. Six months after go-live, new joiners will learn the process from whoever sits next to them. If that person was trained properly and has a set of worked examples to hand, the knowledge survives the project team's departure. Keep a simple record of who has been trained on what, so gaps show up before someone is alone on a shift with a case they have never seen.

Documentation that outlives the project

Most project documentation is written for the project: design documents, decision logs, configuration workbooks and wiki pages that get archived when the budget code closes. Operations needs a shorter set, kept where the team already works. We would include a procedure for each recurring task, a configuration register, a known-issues list with the agreed workaround for each, and a contact sheet for vendor support and escalation.

The configuration register is the one most often missing. It records what is set up where, why it was set that way and who may change it. A year after go-live someone will ask why a certain merchant category routes to the second acquirer, or why a matching tolerance is set at two cents, and without a register the answer is in a closed ticket if it exists at all.

Every document in the set needs an owner in operations and a last-reviewed date. The procedures should follow the same rules as runbooks people open during an incident: short, with the decisions visible on the first screen.

Hypercare, and who fixes what afterwards

Hypercare is the period after go-live when the project team stays close, triages issues daily with operations and fixes defects quickly. It should end on exit criteria agreed in advance, with a planned date that is allowed to move. Two criteria we would expect in almost any version are no severity 1 incidents in the last two weeks, and open defects below an agreed number, each with a fix date. We would add one that often gets left out: operations has handled every exception type at least once without calling the project team.

The harder question is what happens after the exit meeting. Software defects go to the vendor or the development team, under whatever support terms apply. Configuration changes need a named owner, often a systems or platform role inside operations, working under change control. Process changes belong to operations. Reports and data feeds often belong to a data or finance team who may not have been involved in the project at all. Write this down as a short matrix before hypercare ends, and get each owner to accept their part in writing.

Pay attention to anything that sits between two teams. After you move payment providers, for example, the mapping of the new provider's settlement file into your ledger can end up between the finance systems team and the departing engineers, with each assuming the other has it. Items like that tend to break in month four, when nobody remembers who built them.

If a project of yours goes live this quarter, draft the hypercare exit criteria this week and put them in front of the project sponsor and the head of operations together. Agreeing them brings the other handover questions forward to a point where they are still cheap to answer.

Discuss your operations

More insights

All articles