{ "schema_version": 2, "kind": "shortening-method", "format": "agent-skill", "id": "incident-postmortem", "name": "Incident postmortem", "category": "Technical", "summary": "Compress a completed incident into impact, causes, response, and follow-up.", "use_cases": [ "Completed incident reviews", "Reliability reports" ], "word_count": 144, "url": "https://sho.rten.it/methods/incident-postmortem/", "instructions_url": "https://sho.rten.it/methods/incident-postmortem/SKILL.md", "skill_url": "https://sho.rten.it/methods/incident-postmortem/SKILL.md", "json_url": "https://sho.rten.it/methods/incident-postmortem/llms.txt", "plain_text_url": "https://sho.rten.it/methods/incident-postmortem/prompt.txt", "license": "MIT", "sources_url": "https://sho.rten.it/sources/#incident-postmortem", "skill_name": "incident-postmortem", "skill_description": "Compress a completed incident into impact, causes, response, and follow-up. Use for Completed incident reviews, Reliability reports.", "agents_md_url": "https://sho.rten.it/methods/incident-postmortem/AGENTS.md", "sources": [], "instructions": "Condense the supplied account of a completed incident into a review that supports learning and follow-up. Do not turn it into a live status announcement.\n\nKeep the impact, affected scope, start and recovery times, and the evidence for the causal chain. Distinguish the triggering change from contributing conditions and detection gaps. Compress the timeline to events that explain impact, detection, decisions, or recovery. Retain a failed mitigation if it explains the duration or a lesson.\n\nKeep confirmed facts separate from unresolved questions. Preserve the difference between no data loss and no data loss confirmed. Carry action owners, deadlines, and completion states exactly as supplied. Remove repeated chronology, blame, and meeting narration.\n\nUse brief sections: Impact, Cause, Response, and Follow-up. Include only supported sections. Do not assign responsibility, declare an action complete, or claim a prevention measure will eliminate recurrence without evidence in the source.", "example": { "context": "Database connection-pool incident", "before": "On June 11, from 08:10 to 08:24 UTC, the EU preferences service rejected 6% of update requests. These failed writes were not retried automatically, so affected users must retry them. No data loss has been confirmed; the audit is still open. The 08:10 deployment reduced the database connection pool from 40 to 4, which exhausted available connections. The slow-query alert did not fire because queries remained fast. Support detected the issue at 08:14. We rolled back at 08:22 and confirmed recovery at 08:24. Iris will add a connection-wait alert by June 18. A proposed pool-size load test has no owner or date yet.", "after": "Impact: EU preferences rejected 6% of updates on June 11, 08:10–08:24 UTC. Affected users must retry; no automatic retry. Data-loss audit remains open; none confirmed.\nCause: The 08:10 deployment cut the connection pool from 40 to 4, exhausting connections. The slow-query alert missed it because queries remained fast.\nResponse: Support detected the issue at 08:14; rollback at 08:22; recovery confirmed at 08:24 UTC.\nFollow-up: Iris owns a connection-wait alert due June 18. Pool-size load test proposed; owner and date unset.", "must_preserve": [ "EU preferences; 6%; June 11; 08:10–08:24 UTC", "Failed writes need user retry; no automatic retry", "No data loss confirmed; audit still open", "Connection pool fell from 40 to 4; slow-query alert missed connection exhaustion", "Support detection 08:14; rollback 08:22; recovery confirmed 08:24 UTC", "Iris alert due June 18; load test proposed without owner or date" ], "kind": "illustrative", "omitted": [ "First-person narration and repeated references to failed writes and the issue." ] } }