Intermediate
Open
Pro
Fix a Dashboard That Failed During an Incident
During a recent incident, the on-call engineer took 25 minutes to identify that a deploy at 14:00 had caused a conversion-rate drop, even though the metrics needed were "somewhere in our Grafana instance" the whole time. The postmortem notes: the team has 14 separate dashboards, organized by which engineering sub-team owns each service; deploy events are recorded in a separate deployment tracking tool with no link to any metrics dashboard; and the main "system health" dashboard shows CPU, memory, and disk for every host, sorted alphabetically by hostname, with no business or RED metrics visible at all.
- Identify three specific, independent flaws in this dashboard setup that each contributed to the 25-minute delay.
- Redesign the top of a single incident-response dashboard: list, in order top to bottom, what panels should appear and why that order matters for someone investigating under pressure.
- Explain why "14 dashboards organized by owning sub-team" is a reasonable way to organize dashboards for some purpose, and what that purpose is — i.e., why the fix isn't simply "delete all but one dashboard."
Share this question