Five things, done properly, instead of thirty half-finished ones.
Weekly, daily or follow-the-sun. Overrides for holidays take ten seconds and everyone sees who is actually on call right now.
Per severity, per service. Thirty-second granularity, and a policy can page a person, a rota or a whole team.
Fifty alerts from one root cause become one incident. Configurable by service, host or a key in the payload.
Every event with a timestamp, exportable as markdown so the post-mortem lives with the code.
Written from the incident, held for a human. Nothing is ever published to customers automatically.
Weekly: which alerts fired most and were never actionable. Most teams delete a third of their rules in the first month.
Paging runs in two regions with independent SMS and voice providers. If both fail, configured fallback numbers are called directly by the provider from a list we keep synced daily. We publish every incident of our own on the status page, including the embarrassing ones.
No. Per user, per month, and nothing else. Charging per alert punishes teams for monitoring more, which is the opposite of what should happen.
Not currently, and we would rather say that than string you along. Data residency in the UAE or EU is available on the Scale plan.
A team of ten with a dozen services is usually running in an afternoon. We import rotas and escalation policies from the common tools; alert routing rules have to be re-pointed by hand because they never map cleanly.