Skip to main content
Start a Conversation

An Incident Runbook a Small Team Will Use

The best runbook is short, practiced, and tied to actual service symptoms.

Shawn Iuliucci
4 min read
Reliability & Platform Engineering
On this page

A small team does not need a large incident bureaucracy. It needs to know who declares an incident, how customer impact is checked, where the current version and dashboards are, which changes are safe to reverse, and how decisions are recorded.

Begin with symptoms

Write entries for failed login, missing data, slow checkout, or broken form delivery rather than abstract infrastructure categories. Include one reliable first check for each symptom.

Protect the recovery path

Record rollback and restore steps with prerequisites and expected outcomes. Do not improvise a database change during an incident because the command looks familiar.

Learn after service returns

Capture the timeline, customer impact, detection gap, and one or two system improvements with owners. Avoid turning the review into a search for blame.

Build the first page around decisions

The first page should name the incident lead and backup, the customer journey being affected, where to see current health, and how to contact the people who can act. Add safe first checks for the most likely symptoms, with links to current dashboards and deployment history. Label any command that changes data or service state with prerequisites and expected result. A runbook that depends on one person's memory will fail during a holiday or a second simultaneous issue. Keep it short enough to use under pressure and store it where responders can reach it during an outage.

An incident decision card can fit on one screen. It should say how to declare impact, identify the current release, find the health view, pause a bad change, verify recovery, and notify the owner of any stranded work. Link detailed commands elsewhere and label their prerequisites; keep the first page focused on sequence and authority. Add one small log for time, action, result, and decision maker during the event. Rehearsing the card reveals missing permissions and dead links while the stakes are low. After an incident, update the card from observed confusion rather than adding pages of generic process.

Rehearse a bounded failure

Pick a safe exercise, such as a broken nonproduction integration or a deliberately failed synthetic form delivery. Walk through detection, declaration, communication, rollback or repair, verification, and closure. Record the time to each step and the point at which responders had to guess. Afterward, update the runbook and assign one or two system improvements with owners. The review should distinguish the trigger from the conditions that made recovery slow. A practiced path is more useful than a long document nobody has tried.

Decision checklist

  • Name an incident lead and backup.
  • Link to current dashboards and deployment history.
  • Practice one rollback or restore step.
  • Review the runbook after each material incident.

A small test before committing

Walk through a recent or plausible failure with the actual people who would be paged. Without coaching, ask them to find the current owner, customer impact, first diagnostic view, safe rollback, and communication channel. Time the search and note missing access or stale instructions. Keep the runbook short enough to use under pressure, then repeat the exercise after changes. A long document that no one can find is not incident readiness.

Worked scenario

A hypothetical release breaks a quote form. The runbook should lead an operator to test the form, inspect the deployed version, pause the bad release, confirm accepted submissions, and reconcile any stored requests before saying the incident is resolved.

For a scoped application of this decision, see Reliability & Platform Engineering.

Apply this decision to your own system.

Share your current workflow and constraints so the next step can be scoped around real work.

Discuss Your Project