Intermediate
Open
Pro
Redesign a Paging Policy That's Causing Alert Fatigue
An on-call rotation for a data pipeline service reports being paged 12-18 times per week, and admits to muting the pager on their phone "most nights" because almost every page turns out to be a transient blip that resolves on its own within a few minutes. Current alerting rules: page immediately whenever CPU utilization exceeds 80% for any single 1-minute measurement, whenever the error rate exceeds 0.1% for any single 1-minute measurement, and whenever any individual downstream dependency's latency exceeds its p99 baseline by any amount.
- Diagnose specifically what is wrong with each of the three existing alert rules, using the principles of good paging policy.
- Redesign the alerting: propose what should page immediately, what should file a ticket instead, and what should be dashboard-only with no alert at all. Justify each choice.
- The team is worried that loosening thresholds will mean a real incident gets missed. What design element addresses that concern directly, without just making every threshold looser?
Share this question