Alerting and on-call
Channels
Section titled “Channels”Push (native iOS app and Apple Watch) and email are available on every plan. Personal SMS starts with Pulse. On-call SMS and phone calls start with Sentinel as escalation steps to verified numbers. Successful personal and on-call SMS and voice deliveries share the monthly and rolling 24-hour quotas on plans and limits. The full channel table is on alerting and integrations.
The live display
Section titled “The live display”An open incident puts a live display on the phone it woke: on the lock screen, in the Dynamic Island, on an Apple Watch in the Smart Stack, and on the CarPlay dashboard. Lay the phone on its side on a charger and it fills the screen, readable from across the room.
It follows the situation rather than the moment it appeared: how many incidents are open, how bad the worst of them is, which service, and how long it has been running. Once somebody takes a case, the display on every other phone stops offering to take it.
With exactly one incident open it carries Acknowledge, and the case is taken straight from the lock screen without unlocking. With several open there is no button, deliberately: nothing would say which one you meant. There is none on the watch or in CarPlay either, where a button in a live display cannot perform an action at all, and a button that does nothing is worst exactly where somebody is relying on it.
Apple ends a live display a few hours after it began, whatever we do. Nothing in the escalation depends on it: the display is the quiet view beside the alarm, never a replacement for it.
Acknowledging by voice
Section titled “Acknowledging by voice”Somebody woken at three in the morning has the earbuds in and the phone somewhere across the room. Acknowledging is the one act that counts in that minute, because it stops the ladder before it wakes the next colleague. It should not depend on finding a device and unlocking it.
Say “Acknowledge the incident in Perstat” or “Take over the incident in Perstat”. With exactly one case open it is taken, no question asked. With several open, Siri asks which one and offers the service names: taking the wrong case is worse than a question, because it leaves the right one unattended. With nothing open, Siri says that too.
The outcome is spoken, the failure as well. Somebody who says “acknowledge” at night, hears “taken over” and turns over is relying on that word.
This works without unlocking, and that is deliberate: unlocking is exactly what it must not depend on. Anyone with physical access to an on-call phone can stop a running escalation this way. If your policy rules that out, turn Siri off on the locked device (Settings → Face ID & Passcode → Siri when locked).
The built-in phrases are English and German. In the Shortcuts app you can give the action a wording of your own, in whichever language your Siri speaks.
The escalation ladder
Section titled “The escalation ladder”An incident pages the on-call person by push first; without acknowledgment the ladder walks to SMS and then to a phone call, with stage timings configured per schedule. Calls are capped at 2 per incident: escalation should wake a human, not drain a battery.
The chain itself is yours to define, under On-call in the Escalation chain section. Every step has a moment (immediately, or after X minutes), its recipients and its channels. Steps can be added and removed; whoever never changes anything keeps the chain the two numbers used to describe.
A step reaches several recipients at the same time. “After 15 minutes the owners and the team lead and the Slack channel” is one step with three targets, not three steps. A target is whoever is on duty, a specific person, a role, owners and admins, or an endpoint: any connector you set up under Integrations (Slack, Microsoft Teams, Discord, Google Chat, PagerDuty, Opsgenie, or a signed webhook of your own).
Channels can be set per recipient. A step names the channels its people are reached on, and that is the setting most chains ever need. Where one recipient should be reached differently, the step’s card carries a small as the step control next to that recipient: switch it, and this one recipient is paged the way you chose while everybody else on the same rung keeps the step’s setting. So “at minute 15 call the person on duty, but only push the second level” is one step, not three steps at three artificial minutes. An endpoint has exactly one way in and ignores channels entirely, on the step and on itself.
The grid above the steps shows the channels that actually go out, not the step’s default: if every recipient of a rung overrides it, the rung shows what those recipients get.
Because a step already holds several recipients, two steps cannot share the same minute. Saving that is refused, and rightly so: only one of them would ever have gone out.
Above the steps the chain is drawn as a grid: the trigger on top, the steps below it in the order of time, and beside each step its fan of recipients. It shows at a glance what a list cannot, namely how many people one step wakes at once. Editing happens in the cards underneath; the grid is the map, not the workbench.
An endpoint that is switched off or failed on its last delivery is marked right there at the step, with the reason. A chain rung that quietly reaches nothing is worth more attention than a tidy display.
An endpoint that a step alerted also receives the all-clear, even when it is not subscribed to that event under Integrations. This matters most with PagerDuty and Opsgenie: naming an endpoint in a step opens an alert there, and an alert that is never closed keeps ringing after the incident is over. What the chain opens, the chain closes.
Two rules apply regardless of your chain: an incident that has been open for a long time lands on the latest due step rather than walking the chain from the start, and if nobody is on duty, the end of the chain applies immediately. Waiting for a timer that reaches nobody is the one outcome an escalation must not have.
Phone number verification
Section titled “Phone number verification”A number must be verified before it can ever be paged: a 6-digit one-time code, valid for 10 minutes, with a 60-second resend cooldown and at most 5 attempts. Paging only ever goes to verified numbers.
Acknowledgment semantics
Section titled “Acknowledgment semantics”Acknowledging an incident takes ownership: the escalation stops, every other device quiets down, and reminders go to the owner rather than the whole team. The acknowledgment, like every event, lands on the incident timeline with its timestamp.
Snooze
Section titled “Snooze”Snoozing is the smaller gesture, and it is deliberately not the same thing. It silences an incident for 15 minutes and for you alone. Nobody else goes quiet, the escalation is not stopped, the incident is not owned by anyone, and nothing lands on the timeline as a decision.
Use it when you already know and are on your way to a laptop, and acknowledge when you are the person handling it. Choosing snooze where acknowledgment belongs leaves the team believing an incident is unowned; choosing acknowledgment where snooze belongs stops the escalation for everyone.
Because a snooze only ever affects the person who asks for it, it needs no particular role: whoever can be woken by an alert can quiet their own phone.
Dependent alarms
Section titled “Dependent alarms”When one failure causes ten, you want one alarm and not ten. A monitor can name the monitors it depends on, and while one of those origins is provably down, this monitor holds its alarm chain rather than adding to the noise.
Set it on the monitor that would be woken for, not on the origin. The rule is a list of 1 to 20 origin monitors plus a deadline between one minute and 24 hours.
What is held, and what is not, is the part worth reading twice:
| Held | The alarm chain: push, SMS, phone call, the escalation ladder |
| Not held | The incident itself, its start time, downtime, the SLA record, the status page, and status-page subscribers |
The incident opens exactly as it otherwise would. A red light on your public page without the accompanying email is the one inconsistency your customers’ customers would notice, so outward notification continues; only the waking up inside your team waits.
Three properties keep this from becoming the place where alarms disappear:
- Loud in every case of doubt. The alarm is held only when the failure of an origin can be positively established. No rule, an unreachable origin, a failed lookup, or an origin that is up or merely unknown all mean loud.
- A deadline that always exists, counted from the start of the incident rather than from now. Once it passes, the alarm goes loud regardless of what the origin is doing. Counting from the incident start is what stops a flapping origin from extending the silence indefinitely.
- Visible while it is held. A held alarm shows in the notification bell with who is holding it and until when. A held alarm you cannot see would be indistinguishable from a lost one.
A monitor that names itself as an origin holds nothing, otherwise a service would suppress its own alarm.
On-call schedules
Section titled “On-call schedules”Rotations with escalation stages and a coverage-gap warning when nobody is on duty. On-call, like incident command, is included from Sentinel; the exact gates are on plans and limits.
What alerting is not
Section titled “What alerting is not”No alert AI, no anomaly scores deciding whether you get woken up. For the eleven regional probe types, the configured policy confirms the failure; its default quorum is two regions. Agent and heartbeat monitors use neither regions nor quorum and follow their own incoming-signal rules. The quorum rule explains the regional mechanism.