Practice — Production Observability & Monitoring (5 questions)
Intermediate
Open
Free
Use Logs, Metrics, and Traces to Diagnose a Latency Spike Permalink →
Your team's internal recommendation-scoring API has a dashboard showing p99 latency jumped from 55ms to 1.1s starting at 09:14, with no change in request rate or error rate. You have metrics, structured logs, and sampled traces available.
- Describe, in order, how you would use each of the three observability pillars to go from "p99 latency jumped" to a specific, actionable root cause. Be concrete about what each pillar would show you and why you look at them in that order.
- Suppose traces show the extra time is spent in a "feature_fetch" span, but the fetch itself isn't failing (no errors, correct data returned). What would you look at next, and what are two plausible root causes consistent with "same data, just slower"?
- A colleague suggests skipping straight to grepping the raw application logs from 09:14 onward instead of following the metric → trace → log order. Explain what makes that approach slower or less reliable here.
Share this question
Intermediate
Open
Pro