Back to cookbook

ChatGPT Prompt to Write a Clear Incident Postmortem Report

2 views Updated
Share

This incident postmortem prompt helps engineers and engineering managers turn a messy timeline of logs, Slack messages, and on-call notes into a structured postmortem report after a production outage or bug. It's built for anyone who needs to write up what broke, why it broke, and what changes will stop it from happening again, without burying the root cause under vague summaries or blame-focused language.

The prompt pushes the model to separate the timeline of events from the root cause analysis and the action items, which is where most postmortems fall apart. A model given raw logs and a short description will often default to a generic "we had downtime and fixed it" summary. Structuring the request around specific sections forces a more honest accounting of detection time, impact, and contributing factors, which is what most postmortem templates actually require for a review.

Because the output is only as precise as the detail you feed in, pairing this with a tool that trims a long incident timeline down to the lines that actually matter can make the draft tighter and more accurate. If your incident thread is thousands of lines of logs, the Token Optimizer can compress that raw material into a cost-effective prompt before you run it through this template.

Prompt template

ROLE: You are a senior engineer writing a blameless incident postmortem for [TEAM OR COMPANY NAME].

CONTEXT:

  • Incident summary: [ONE-SENTENCE DESCRIPTION OF WHAT WENT WRONG]
  • Start time: [TIMESTAMP INCIDENT BEGAN]
  • Detection time: [TIMESTAMP TEAM NOTICED OR WAS ALERTED]
  • Resolution time: [TIMESTAMP SERVICE WAS RESTORED]
  • Systems affected: [LIST OF SERVICES, APIS, OR FEATURES IMPACTED]
  • User impact: [WHO WAS AFFECTED AND HOW, E.G. ERROR RATE, DOWNTIME DURATION]
  • Raw timeline or logs: [PASTE RELEVANT LOG LINES, ALERTS, OR CHAT MESSAGES WITH TIMESTAMPS]
  • Known root cause (if any): [WHAT YOU'VE CONFIRMED SO FAR, OR "UNDER INVESTIGATION"]

CONSTRAINTS:

  • Use blameless language: describe system and process failures, never name individuals
  • Clearly separate confirmed facts from open questions still being investigated
  • Keep the timeline in chronological order with timestamps for each major event
  • Action items must each have a specific owner role and a target timeframe

OUTPUT FORMAT:

  1. Summary (2-3 sentences: what happened, impact, current status)
  2. Timeline (chronological, timestamped events from detection through resolution)
  3. Root Cause Analysis (what caused it, and any contributing factors)
  4. Impact (who/what was affected, quantified where possible)
  5. Action Items (numbered list, each with owner role and target date)
  6. Open Questions (anything still unresolved or under investigation)

Example input

ROLE: You are a senior engineer writing a blameless incident postmortem for Fintrail, a payments platform.

CONTEXT:

  • Incident summary: Checkout API returned 500 errors for roughly 40 minutes, blocking new payments
  • Start time: 2026-09-14 14:02 UTC
  • Detection time: 2026-09-14 14:11 UTC (PagerDuty alert on 5xx rate)
  • Resolution time: 2026-09-14 14:43 UTC
  • Systems affected: checkout-api, payment-gateway-adapter
  • User impact: ~12% of checkout attempts failed during the window, roughly 340 failed transactions
  • Raw timeline or logs: 14:02 UTC - deploy of checkout-api v2.14.0 completed 14:03 UTC - error rate on /checkout/submit begins climbing 14:11 UTC - PagerDuty fires "High 5xx rate: checkout-api" 14:14 UTC - on-call engineer acknowledges, starts investigating logs 14:22 UTC - identifies null pointer exception in new discount-code validation path 14:30 UTC - rollback to v2.13.8 initiated 14:43 UTC - error rate returns to baseline, incident closed
  • Known root cause: v2.14.0 shipped a discount-code validation change that didn't handle null promo codes, causing a crash on any checkout without a promo code applied

CONSTRAINTS:

  • Use blameless language: describe system and process failures, never name individuals
  • Clearly separate confirmed facts from open questions still being investigated
  • Keep the timeline in chronological order with timestamps for each major event
  • Action items must each have a specific owner role and a target timeframe

OUTPUT FORMAT:

  1. Summary (2-3 sentences: what happened, impact, current status)
  2. Timeline (chronological, timestamped events from detection through resolution)
  3. Root Cause Analysis (what caused it, and any contributing factors)
  4. Impact (who/what was affected, quantified where possible)
  5. Action Items (numbered list, each with owner role and target date)
  6. Open Questions (anything still unresolved or under investigation)

Example output

  1. Summary A deploy of checkout-api v2.14.0 introduced a null pointer exception in the discount-code validation path, causing 500 errors on checkout submissions for 40 minutes. The incident was resolved by rolling back to the previous version. All checkout functionality has been restored.

  2. Timeline

  • 14:02 UTC: checkout-api v2.14.0 deployed to production
  • 14:03 UTC: error rate on /checkout/submit begins climbing
  • 14:11 UTC: PagerDuty alert fires for high 5xx rate on checkout-api
  • 14:14 UTC: On-call engineering acknowledges alert and begins log investigation
  • 14:22 UTC: Null pointer exception identified in new discount-code validation logic
  • 14:30 UTC: Rollback to v2.13.8 initiated
  • 14:43 UTC: Error rate returns to baseline; incident closed
  1. Root Cause Analysis The v2.14.0 release introduced validation logic for discount codes that assumed a promo code value was always present. Checkouts submitted without a promo code triggered a null pointer exception, causing the request to fail with a 500 error. A contributing factor was that the pre-release test suite did not include a case for checkout without a promo code.

  2. Impact Approximately 12% of checkout attempts failed during the 40-minute window, affecting an estimated 340 transactions. No payment data was lost or duplicated; affected users were able to retry successfully after the rollback.

  3. Action Items

  4. Add a null-check regression test for checkout without a promo code — Owner: Checkout team engineer, Target: within 3 business days

  5. Add a required test case for "no promo code" to the checkout pre-release test suite — Owner: QA lead, Target: before next release cycle

  6. Add a canary deployment stage for checkout-api that monitors 5xx rate before full rollout — Owner: Platform engineering, Target: within 2 sprints

  7. Open Questions

  • Whether the same null-check gap exists in the mobile checkout client is still under investigation.
  • Confirming whether any of the 340 failed transactions resulted in abandoned carts versus successful retries is still being analyzed with the analytics team.

When to use it

  • You need a first draft of a postmortem within hours of an incident being resolved, while the timeline is still fresh
  • You have raw inputs like PagerDuty alerts, Slack threads, or deployment logs and need them turned into a coherent narrative
  • Your team requires a standard postmortem format (timeline, impact, root cause, action items) for every production incident
  • You're writing a postmortem for a non-technical stakeholder audience and need the technical detail translated into plain impact language

Best practices

  • Paste in the actual timestamps from alerts and deploy logs rather than paraphrasing them, so the model can reconstruct an accurate sequence of events
  • Separate what you know for certain (confirmed root cause) from what's still being investigated, and tell the model to flag open questions instead of guessing
  • Ask for blameless language explicitly; without that instruction, drafts can drift toward naming individuals instead of describing system or process gaps
  • Request specific, owned action items with due dates rather than generic "improve monitoring" bullets, since vague follow-ups rarely get tracked to completion

Common mistakes

  • Feeding in only a verbal summary of the incident instead of the actual logs or timestamps, which leads to a generic-sounding report with made-up specifics
  • Skipping the "detection" and "response" timeline stages and jumping straight to root cause, which hides how long the team took to notice and react
  • Letting the model write action items as vague intentions instead of concrete, assignable tasks with owners and deadlines
  • Not specifying the audience, so the draft mixes deep technical detail with executive-level impact summary in a way that serves neither reader well

FAQs

How do I write a blameless postmortem with ChatGPT?

Give the model a role instruction that explicitly says to describe system and process failures rather than naming individuals, then feed it the actual timeline of events. The blameless framing works best when it's stated as a constraint up front, not left implicit, since models will otherwise default to whatever tone the raw input suggests.

What should I include in an incident postmortem besides the root cause?

A complete postmortem needs a timestamped timeline from detection to resolution, a quantified impact statement, the root cause and any contributing factors, and a list of specific action items with owners and deadlines. Skipping the timeline or impact sections is the most common reason postmortems read as incomplete during review.

Can AI write an accurate postmortem from just logs, without a human writing the summary first?

AI can draft a strong first pass from raw logs and alert timestamps, but it should not be the final word on root cause if the cause is still under investigation. The model will produce a plausible-sounding explanation even from incomplete data, so any output should be reviewed by someone who was involved in the incident before it's shared widely.

How long should an incident postmortem be?

Most production postmortems run one to two pages: a short summary, a timeline, root cause, impact, and action items. Length should scale with incident severity rather than being fixed, since a minor five-minute blip doesn't need the same depth as a multi-hour outage.

Which Cuelara tool can help me tighten this prompt before using it?

Token Optimizer — if your incident timeline includes long log dumps or chat exports, this tool compresses that raw material into a leaner prompt so the postmortem draft stays accurate without the extra token cost.

Found this prompt useful? Share it.

Share

More in Engineering

Engineering

Secure API Request Handler

A prompt that makes the model write an API route the way a security conscious reviewer would want it: validated input, safe data access, uni…

Write a Node.js Express route handler for a [METHOD] request to '[PATH]'.

Requirements:
- Validate the body with Zod: [FIELDS AND CONSTRAINTS].
- Use parameterized queries / the ORM only. No string-concatenated SQL.
- Wrap the logic in try/catch.
- Return standardized errors: { "error": { "code": string, "message": string } } with 400 for validation failures and 500 for server errors. Never leak stack traces.
- Add a short comment above any security-relevant line.

Return only the code.
Engineering

Senior React Developer Persona

This is a system prompt that turns a general purpose model into a disciplined senior frontend engineer. Instead of asking for "a React compo…

You are a Senior Frontend Engineer specializing in React, Next.js (App Router) and Tailwind CSS.

Rules:
1. Use functional components with TypeScript. Export a typed props interface.
2. Prefer Tailwind utility classes over custom CSS.
3. Use Framer Motion for micro-interactions only; respect prefers-reduced-motion.
4. Make every interactive element keyboard accessible with correct ARIA attributes.
5. Respond with ONLY the code block. No explanations unless I ask.

Task: [DESCRIBE THE COMPONENT]
Engineering

AI Prompt to Write a Rollback Plan for a Risky Deployment

This AI prompt for a deployment rollback plan helps engineers document exactly how to reverse a risky release before it ships, not after som…

ROLE: You are a senior site reliability engineer writing a rollback plan for a production deployment.

CONTEXT:
- Service or system being deployed: [SERVICE NAME]
- What the deployment changes: [CODE CHANGES, CONFIG CHANGES, DATABASE MIGRATIONS, ETC.]
- Deployment method: [CI/CD PIPELINE, MANUAL DEPLOY, FEATURE FLAG ROLLOUT]
- Dependent services or consumers affected: [LIST DEPENDENT SERVICES]

TASK:
Write a rollback plan for this deployment that includes:
1. The specific monitoring signals or alerts that indicate a rollback is needed
2. Step-by-step rollback instructions in execution order, each with an owner
3. Any step that is irreversible or partially irreversible, flagged separately
4. A verification step to confirm the rollback succeeded
5. An estimated time to complete the rollback

CONSTRAINTS:
- Assume the person executing the rollback may not be the person who wrote the plan
- Do not assume manual database fixes are safe without naming the exact commands or scripts
- Keep each step to one action

OUTPUT FORMAT:
A numbered list of rollback steps, followed by a short "Irreversible Actions" section and a "Verification" section.