Match a job Paths Subjects Questions Quizzes Pricing
Intermediate Open Pro

Redesign a Paging Policy That's Causing Alert Fatigue

An on-call rotation for a data pipeline service reports being paged 12-18 times per week, and admits to muting the pager on their phone "most nights" because almost every page turns out to be a transient blip that resolves on its own within a few minutes. Current alerting rules: page immediately whenever CPU utilization exceeds 80% for any single 1-minute measurement, whenever the error rate exceeds 0.1% for any single 1-minute measurement, and whenever any individual downstream dependency's latency exceeds its p99 baseline by any amount.

  1. Diagnose specifically what is wrong with each of the three existing alert rules, using the principles of good paging policy.
  2. Redesign the alerting: propose what should page immediately, what should file a ticket instead, and what should be dashboard-only with no alert at all. Justify each choice.
  3. The team is worried that loosening thresholds will mean a real incident gets missed. What design element addresses that concern directly, without just making every threshold looser?

Share this question

← Back to Production Observability & Monitoring practice

We use cookies for product analytics to improve OmniAtlas. See our Privacy Policy.