incident-runbook-templates
wshobson/agents
Create structured incident response runbooks with detection, triage, mitigation, and escalation procedures.
What is incident-runbook-templates?
This skill provides production-ready templates for building incident response runbooks that guide on-call engineers through detection, triage, mitigation, resolution, and communication during outages. Use it when standardizing incident procedures across teams, onboarding new responders, or documenting recovery steps for specific services.
- Generate runbooks with severity levels (SEV1-SEV4) and corresponding response times
- Structure incident procedures covering detection, triage, mitigation, root cause investigation, and resolution
- Create escalation matrices and communication templates for stakeholder updates
- Include rollback procedures and verification steps to prevent cascading failures
- Add prerequisite checks and failure handling notes to prevent runbook steps from failing under stress
- Establish metadata tracking (last verified date, owner, review cadence) to prevent runbook rot
How to install incident-runbook-templates
npx skills add https://github.com/wshobson/agents --skill incident-runbook-templatesHow to use incident-runbook-templates
- 1.Define incident severity levels (SEV1-SEV4) and response time targets for your service
- 2.Structure your runbook using the eight-section template: Overview, Detection, Triage, Mitigation, Root Cause Investigation, Resolution, Verification, and Communication
- 3.Write each step assuming a stressed 3 AM responder—include prerequisites, expected outputs, and failure handling for every command
- 4.Add a numbered quick checklist at the top that mirrors section numbers for progress tracking under stress
- 5.Include rollback steps and verification procedures after each mitigation action
- 6.Add metadata (last verified date, owner, review cadence) and set up CI checks to validate commands and endpoints remain current
- 7.Test runbooks regularly via game days or chaos engineering exercises and update after every incident
Use cases
- Building a payment processing system outage runbook with detection alerts, mitigation steps, and rollback procedures
- Creating database incident procedures for connection pool exhaustion, replication lag, and disk space alerts
- Onboarding new on-call engineers with step-by-step recovery guides designed for high-stress situations
- Standardizing escalation matrices and communication protocols across multiple engineering teams
- Documenting service-specific runbooks with dashboard links and prerequisite checks for production incidents
- On-call engineers and incident responders
- Platform and infrastructure teams
- Engineering managers building incident response processes
- DevOps and SRE practitioners
- Teams standardizing incident procedures across services
incident-runbook-templates FAQ
Add prerequisite checks and a 'what to do if this fails' note for each command. Document assumptions about cluster state, configuration, and environment. Include dry-run outputs for destructive operations before execution.
Add a 'Last Verified' date and owner at the top. Set up CI checks that validate curl endpoints and kubectl context names. Review and update after every SEV1/SEV2 incident and on a regular cadence.
Add a numbered quick checklist at the very top of the runbook that mirrors section numbers. This lets responders track progress under stress without reading the full document.
Assign a dedicated incident communicator role separate from the incident commander. Use a standing communication template that posts updates every 15 minutes with current status, impact, what you're doing, and next update time.
Add explicit warnings before destructive SQL commands. Require a dry-run output check and verification of reasonable counts before executing any termination or modification commands.
Full instructions (SKILL.md)
Source of truth, from wshobson/agents.
name: incident-runbook-templates description: Create structured incident response runbooks with step-by-step procedures, escalation paths, and recovery actions. Use this skill when building a service outage runbook for a payment processing system; creating database incident procedures covering connection pool exhaustion, replication lag, and disk space alerts; onboarding new on-call engineers who need step-by-step recovery guides written for a 3 AM brain; or standardizing escalation matrices across multiple engineering teams.
Incident Runbook Templates
Production-ready templates for incident response runbooks covering detection, triage, mitigation, resolution, and communication.
When to Use This Skill
- Creating incident response procedures
- Building service-specific runbooks
- Establishing escalation paths
- Documenting recovery procedures
- Responding to active incidents
- Onboarding on-call engineers
Core Concepts
1. Incident Severity Levels
| Severity | Impact | Response Time | Example |
|---|---|---|---|
| SEV1 | Complete outage, data loss | 15 min | Production down |
| SEV2 | Major degradation | 30 min | Critical feature broken |
| SEV3 | Minor impact | 2 hours | Non-critical bug |
| SEV4 | Minimal impact | Next business day | Cosmetic issue |
2. Runbook Structure
1. Overview & Impact
2. Detection & Alerts
3. Initial Triage
4. Mitigation Steps
5. Root Cause Investigation
6. Resolution Procedures
7. Verification & Rollback
8. Communication Templates
9. Escalation Matrix
Detailed patterns and worked examples
Detailed pattern documentation lives in references/details.md. Read that file when the navigation tier above is insufficient.
Best Practices
Do's
- Keep runbooks updated - Review after every incident
- Test runbooks regularly - Game days, chaos engineering
- Include rollback steps - Always have an escape hatch
- Document assumptions - What must be true for steps to work
- Link to dashboards - Quick access during stress
Don'ts
- Don't assume knowledge - Write for 3 AM brain
- Don't skip verification - Confirm each step worked
- Don't forget communication - Keep stakeholders informed
- Don't work alone - Escalate early
- Don't skip postmortems - Learn from every incident
Troubleshooting
Runbook steps work in staging but fail during a real incident
Steps often assume preconditions that are true in a healthy environment but not during an outage. For each command in your runbook, add a prerequisite check and a "what to do if this command fails" note:
# Step: Check pod status
kubectl get pods -n payments
# Prerequisites: kubectl configured, kubeconfig points to correct cluster
# If this fails: run `aws eks update-kubeconfig --name prod-cluster --region us-east-1`
# Expected output: pods in Running state
On-call engineer panics and skips steps out of order
Add a numbered checklist at the top of the runbook that mirrors the section numbers, so responders can track progress under stress without reading the full document:
## Quick Checklist
- [ ] 1. Declare incident severity and open war room
- [ ] 2. Check service health (Section 4.1)
- [ ] 3. Check recent deployments (Section 4.1)
- [ ] 4. Roll back if deploy is suspect (Section 4.1)
- [ ] 5. Post initial notification to #payments-incidents
- [ ] 6. Escalate if > 15 min unresolved
Runbook is outdated — commands reference old cluster names or endpoints
Runbooks rot because they're updated manually. Include a "Last Verified" date and owner at the top, and add a CI check that validates all curl endpoints and kubectl context names are still valid:
## Runbook Metadata
| Field | Value |
|---|---|
| Last verified | 2024-11-15 |
| Owner | @platform-team |
| Review cadence | After every SEV1/SEV2 |
Stakeholder communication is delayed while engineers are heads-down
Assign a dedicated incident communicator role (separate from the incident commander) whose only job is to post status updates. Add a standing agenda in the communication template:
Update every 15 minutes (even if no new information):
- Current status (Investigating / Mitigating / Monitoring)
- Impact (what is broken, who is affected, % of traffic)
- What we are doing right now
- Next update in: 15 minutes
Database runbook commands cause additional downtime when run incorrectly
Add explicit warnings before destructive SQL commands and require a dry-run output check before executing:
-- WARNING: This terminates active connections. Verify count first.
-- DRY RUN (check count before terminating):
SELECT count(*) FROM pg_stat_activity WHERE state = 'idle' AND query_start < now() - interval '10 minutes';
-- EXECUTE only after verifying count is reasonable (< 50):
SELECT pg_terminate_backend(pid) FROM pg_stat_activity
WHERE state = 'idle' AND query_start < now() - interval '10 minutes';
Related Skills
postmortem-writing- After resolving an incident, use postmortem templates to capture root cause and preventive actionson-call-handoff-patterns- Structure shift handoffs so the incoming responder has full context on active incidents
Related skills
More from wshobson/agents and the wider catalog.

interaction-design
Design microinteractions, motion, and transitions that enhance UI polish and user delight.

istio-traffic-management
Configure Istio service mesh traffic routing, load balancing, circuit breakers, and canary deployments.

javascript-testing-patterns
Implement comprehensive testing strategies with Jest, Vitest, and Testing Library for JavaScript/TypeScript.

k8s-manifest-generator
Generate production-ready Kubernetes manifests with built-in best practices and security standards.

k8s-security-policies
Implement NetworkPolicy, PodSecurityPolicy, RBAC, and Pod Security Standards for production Kubernetes security.

kpi-dashboard-design
Design executive KPI dashboards with metrics selection, visualization patterns, and real-time monitoring best practices.