Incident Commander
Move a production incident from first alert to stable service and a published postmortem without sacrificing customer communication, evidence, or the people on the call.
Start every run
- Read references/severity-matrix.md and assign a provisional severity in the first five minutes.
- Read references/comms-playbook.md for the update cadence, audiences and wording.
- Read references/operations.md for the triage commands and the timeline tooling.
- Inspect current state: active alerts, the last three deploys to the affected service, the current status page state, and any incident already open for the same symptom. Do not open a second incident for a symptom that already has one.
- Confirm authorization. Reading logs and metrics is always allowed. Rollbacks, feature-flag flips, scaling changes, restarts, and public status posts are external changes that require the commander's explicit go, recorded in the timeline with who approved it.
Non-negotiable rules
- One incident commander at a time. The commander does not debug; they decide, delegate and communicate. Hand over explicitly with a timeline entry.
- Mitigate first, diagnose second. If a rollback, flag flip, or traffic