Skip to content
Back to blog
Major incidentCrisis managementCommunication

A playbook for major incidents: staying calm when everything breaks at once

Ruben van der Graaf4 min read

A major incident playbook prevents chaos during big outages. Learn how to arrange roles, communication and evaluation before the outage, not during it.

During a major outage there is no time to figure out who does what. The phone rings, several teams jump in at once, and everyone looks at everyone else. Without a playbook agreed in advance, chaos sets in at exactly the moment calm and clarity are needed most. A good playbook arranges in advance who solves the problem, who communicates and how you evaluate afterward, so a big outage becomes a controlled process instead of headless scrambling. This article shows how to build that.

What makes a major incident different

A regular request follows a known path: accept, prioritize, resolve. A major incident is different because the impact is large, several parties are involved and outside pressure grows as the outage lasts longer. Managers call, users email, and the team working on the fix gets distracted by "when will this be done" while still figuring out what is wrong. A playbook makes sure solving and communicating do not get in each other's way.

The roles you need to arrange in advance

The incident commander: who is in charge

For every major incident, appoint one person to run the response. This person does not necessarily solve the problem themselves, but makes sure the right people are at the table, keeps an overview and decides next steps. Without this role, specialists get involved later than needed.

The solvers: focus on the technical problem

The specialists working on the outage need to focus on the problem. They should not be answering phone calls or writing updates. Shield them from that, so their attention stays on the fix.

The communicator: one voice to the organization

Appoint someone to handle communication to the organization, separate from solving the problem. This person gathers updates from the incident commander and translates them into plain language. One fixed source prevents conflicting messages that undermine trust.

Communicating with the organization

During a major incident, silence is the worst thing you can do. Even without a solution yet, the organization wants to know work is underway and when the next update comes. Agree a fixed rhythm in advance, for example every half hour, regardless of news. Use a fixed channel everyone knows, so nobody has to search for the latest information. Be honest about what you do not know yet; "we are still investigating the cause" beats no update at all.

The playbook in practice

A playbook only works if it is written down and practiced in advance, not invented during the first real outage.

  • Define when something is a major incident. Agree clear criteria, for example based on impact and the number of affected users, so escalating is not up for debate.
  • Assign roles in advance. Record who can be incident commander, who communicates and how to reach those people quickly.
  • Set a fixed communication rhythm. Agree how often and through which channel you share updates, even without news.
  • Document during the incident. Keep a timeline of what happens and when, as the basis for the evaluation.
  • Evaluate after every major incident. Look at what went well and what did not, and update the playbook.
That way, nobody has to figure out who does what while the outage is happening.

What it delivers

A good playbook does not necessarily shorten the time it takes to find the technical cause, but it prevents the chaos around it. Roles are clear, communication is predictable, and the organization feels informed instead of ignored. Afterward, the team learns from every outage instead of repeating the same confusion. That is achieving more with the same people: not by running faster during a crisis, but by arranging in advance what makes it manageable.

Frequently asked questions

Who should be incident commander? Not necessarily the most technical person, but someone who can keep an overview and is willing to make decisions. Often that is a team lead or an experienced coordinator.

How often should we practice the playbook? At least once a year with a simulated outage. That shows whether roles and contact details are still correct.

What if a major incident lasts longer than expected? Rotate staffing where needed so people do not burn out, but keep the fixed roles and communication rhythm in place.

Want a major incident playbook that fits your organization and that your team actually knows? book a call and we will build it together, from roles to communication.

Further reading

Want to apply this in your own organization?

Schedule a no-obligation conversation. Together we look at where you stand and what the first step is.

Get in touch