writing-postmortems
riekelt/technical-writer
Write blameless, evidence-based postmortems and incident reports that outlive the crisis.
What is writing-postmortems?
A structured approach to documenting incidents, outages, and near-misses after resolution. Use this to transform incident channels and war-room threads into durable, factual records that guide future prevention—grounded in evidence, mechanisms, and owned action items, never blame.
- Enforces blameless framing by focusing on structural causes and process gaps, not individuals
- Builds timestamped timelines traceable to logs, alerts, commits, and messages—facts only, no interpretation
- Documents root causes and contributing factors grounded in code and system mechanisms
- Captures wrong turns and misleading signals followed during response for learning
- Generates owned action items with acceptance criteria, filed as tracker issues and linked
How to install writing-postmortems
npx skills add https://github.com/riekelt/technical-writer --skill writing-postmortems- The `technical-writing` skill (required background for hard rules, truth rules, and style)
- Access to incident evidence: logs, alerts, commits, messages, and timestamps
- An organizational severity taxonomy (or explicit acknowledgment if none exists)
- A tracker system for filing and linking action items
How to use writing-postmortems
- 1.Gather all incident evidence: logs, alerts, commits, Slack threads, and timestamps before starting
- 2.Fill the metadata table: incident date, duration, severity (from org taxonomy), and status
- 3.Write a one-sentence summary of root cause and quantified user impact
- 4.Build the timeline as timestamped, evidence-traceable entries—no interpretation or narrative
- 5.Document root cause and contributing factors grounded in mechanisms and code, naming which gates missed or caught the issue
- 6.List wrong turns and misleading signals that extended response time
- 7.Create action items with owners and acceptance checks; file as tracker issues and link them
- 8.Add sign-off line(s) with reviewer name, date, and scope of approval; keep the postmortem Draft until links exist
Use cases
- Writing a postmortem after a production outage affecting users
- Documenting a defect that reached production to prevent recurrence
- Recording a near-miss where a guard caught a problem before user impact
- Converting an incident Slack channel or war-room thread into an immutable historical record
- Creating a root-cause analysis for a data issue or system failure
- On-call engineers and incident commanders
- Technical leads and engineering managers reviewing incidents
- Site reliability engineers (SREs) building incident response culture
- Teams building runbooks and prevention systems
writing-postmortems FAQ
After an incident is resolved or at a stable milestone of a long one. Invoke for any defect reaching users or near-miss caught by a guard. Do not invoke during the live incident (the runbook governs that) or for stakeholder status updates.
Name the role, gate, dependency, or missing guard—never the person. Blaming a person stops at 'be more careful'; blaming a structure produces an action item. Ownership of action items is assignment, not blame.
Every entry must trace to something checkable: a log line, alert, commit, or message with a timestamp from systems, not memory. This ensures the document is a factual record that survives the incident and guides future prevention.
Leave the severity field explicitly unassigned rather than inventing one. Severity should come from a defined org taxonomy, not assumed.
No. Once reviewed, the postmortem is immutable. New findings are added as dated addenda. A rewritten postmortem falsifies the record.
Full instructions (SKILL.md)
Source of truth, from riekelt/technical-writer.
name: writing-postmortems description: Use when writing a postmortem, incident report, or root-cause analysis after an outage, a defect that reached users, a data issue, or a near miss - or when turning an incident channel, alert log, or war-room thread into a durable document. Encodes the blameless framing, the evidence-only timeline, contributing factors, and owned action items. Use whenever something broke and the write-up must outlive the incident.
Writing postmortems
REQUIRED BACKGROUND: the technical-writing skill (hard rules, truth rules, style).
Overview
Audience: the engineer who hits something similar in a year, not the people in the room. Core principle: facts from evidence, causes from mechanisms, lessons from both, and no names. Classification: historical; once reviewed it is immutable, and corrections are dated addenda.
When to invoke, and not
Invoke after an incident is resolved, or at a stable milestone of a long one. A defect that reached users and a near miss (a guard caught what review missed) earn the same write-up. Do NOT invoke during the live incident: the runbook governs that. Not for assigning accountability: a postmortem that needs a person's name to make sense is describing a process hole. Not for a status update to stakeholders, which is a report.
Skeleton
# [System]: [the failure, as its symptom, one line]
| | |
|---|---|
| **Incident date** | YYYY-MM-DD |
| **Duration** | detection to resolution |
| **Severity** | [per the org taxonomy, defined where used] |
| **Status** | Draft / Reviewed |
## Summary
[User-visible impact with numbers, the root cause in one sentence, current state.
The reader who stops here still knows what happened.]
## Impact
[Who and what, quantified: requests failed, records affected, money or time lost.
Estimates labeled as estimates; "no evidence of data loss" only if actually checked,
and say what was checked.]
## Timeline
[Timestamped entries, facts only, each traceable to evidence: an alert, a log line,
a commit, a message. Interpretation lives in the sections below, never here.]
## Root cause and contributing factors
[The mechanism, grounded in code and commits. Contributing factors as a list;
"every component was correct and the defect lived between them" is a valid root
cause. Name which layer of checking missed it and which caught it.]
## Wrong turns during the response
[The wrong first fix, the misleading signal followed, the theory that cost an hour.]
## Action items
[Each with an owner and an acceptance check, filed as tracker issues and linked.
"Investigate X" without an owner is banned here as everywhere.]
## Sign-off
[- YYYY-MM-DD <reviewer>: approved <scope>, one line per review event.]
Rules
- Blameless means structural. Name the role, the gate, the dependency, the missing guard; never the person. Blaming a person stops at "be more careful"; blaming a structure produces an action item. This governs the narrative (timeline, causes, response); action-item ownership is assignment, not blame: a role in this document, an individual in the linked tracker issue.
- The timeline is evidence, not narrative. Every entry traces to something checkable, timestamps from the systems rather than memory. Where memory is the only source, say so.
- Wrong fixes are content. The fix that made sense and did not work belongs in the document with the reasoning that made it plausible; embarrassment is not a retention policy.
- Severity comes from the org taxonomy, defined where used, not assumed. When no taxonomy exists, leave the field explicitly unassigned rather than inventing one.
- Action items follow
writing-issues: outcome, owner, acceptance check, filed and linked, never left as prose intentions in the postmortem. When filing is not possible from where you sit, mark each item[to file: <who files it>]; the postmortem stays Draft until the links exist. - Immutable once reviewed. New findings are dated addenda; a rewritten postmortem is a falsified record. The review itself goes in a sign-off line (see
references/truth.mdin the core skill). - Near misses use the same skeleton, with Impact describing what would have happened, labeled as the counterfactual it is.
The decision that often follows a postmortem (a new invariant, a policy change) is recorded via recording-decisions and linked, not embedded. Runbook updates the incident exposed go through writing-runbooks in the same change as the fix; documentation is part of done, not a follow-up ticket.
Related skills
More from riekelt/technical-writer and the wider catalog.

writing-runbooks
Write operational documentation—runbooks, setup guides, release procedures, migration guides—that people execute under time pressure.

diagramming-processes
Diagram business processes, workflows, and system interactions as maintainable source code.

documenting-contracts
Document HTTP APIs, message contracts, and file formats with exhaustive wire-level detail at the right abstraction level.

documenting-legacy-codebases
Reconstruct documentation for legacy codebases by grounding it in code, not memory or stale docs.

multiplayer-game
Pragmatic patterns for building multiplayer games with matchmaking, tick loops, realtime state, and validation.

rivet-actors
Actors: The primitive for agent orchestration.