Match a job Paths Subjects Questions Quizzes Pricing
Intermediate Open Pro

Fix a Dashboard That Failed During an Incident

During a recent incident, the on-call engineer took 25 minutes to identify that a deploy at 14:00 had caused a conversion-rate drop, even though the metrics needed were "somewhere in our Grafana instance" the whole time. The postmortem notes: the team has 14 separate dashboards, organized by which engineering sub-team owns each service; deploy events are recorded in a separate deployment tracking tool with no link to any metrics dashboard; and the main "system health" dashboard shows CPU, memory, and disk for every host, sorted alphabetically by hostname, with no business or RED metrics visible at all.

  1. Identify three specific, independent flaws in this dashboard setup that each contributed to the 25-minute delay.
  2. Redesign the top of a single incident-response dashboard: list, in order top to bottom, what panels should appear and why that order matters for someone investigating under pressure.
  3. Explain why "14 dashboards organized by owning sub-team" is a reasonable way to organize dashboards for some purpose, and what that purpose is — i.e., why the fix isn't simply "delete all but one dashboard."

Share this question

← Back to Production Observability & Monitoring practice

We use cookies for product analytics to improve OmniAtlas. See our Privacy Policy.