Incidents and post-mortems
Confirmed incidents
Section titled “Confirmed incidents”A quorum-confirmed failure opens an incident with its start at the confirmation second. Every event lands on the timeline: region votes, notifications, acknowledgment, recovery. Linked status-page components and the curated availability record use this timeline.
Manual incidents
Section titled “Manual incidents”Not everything a customer notices is a probe failure. Incidents can be opened by hand, with or without a bound monitor, curated, and explicitly published to the status pages you choose. Nothing publishes implicitly. Manual incidents start with Sentinel; automatic incidents and acknowledgement are available in every plan.
Post-mortems and action items
Section titled “Post-mortems and action items”Writing and publishing post-mortems is part of incident command, so it needs Sentinel or above - the same gate as opening incidents by hand. Automatic, quorum-confirmed incidents keep opening on every plan; what the gate covers is the manual work on top of them.
A post-mortem attaches to the incident: impact, root cause, lessons, plus action items with owners. It moves from draft to published deliberately, internal or public. Each post-mortem has its own address in the app, so it can be linked from a ticket or a report.
Saving and publishing are separate steps. Writing never puts anything in front of your customers; publishing is its own action, and so is choosing whether the result is internal or public.
Every change appears in version history with the actor, time, and action such as save, publish, or restore. Restoring an earlier state creates another history entry, and the activity log records the affected fields. This is attributable and reversible version history, not an immutable or tamper-proof store. The object-specific retention period for these artifacts is part of the current product and contract review; confirm it before relying on a duration.
Acknowledgment and ownership
Section titled “Acknowledgment and ownership”Acknowledging takes the incident: escalation stops, other devices quiet down, reminders address the owner. The acknowledgment timestamp is part of the record, which is why the escalation ladder ends in an acknowledged incident, not in silence.