Match a job Paths Subjects Questions Quizzes Pricing
Overview Read Practice

Practice — Production Observability & Monitoring (5 questions)

Pro content

Sign up free, then start a 14-day Pro trial — no card needed.

Intermediate Open Free

Use Logs, Metrics, and Traces to Diagnose a Latency Spike Permalink →

Your team's internal recommendation-scoring API has a dashboard showing p99 latency jumped from 55ms to 1.1s starting at 09:14, with no change in request rate or error rate. You have metrics, structured logs, and sampled traces available.

  1. Describe, in order, how you would use each of the three observability pillars to go from "p99 latency jumped" to a specific, actionable root cause. Be concrete about what each pillar would show you and why you look at them in that order.
  2. Suppose traces show the extra time is spent in a "feature_fetch" span, but the fetch itself isn't failing (no errors, correct data returned). What would you look at next, and what are two plausible root causes consistent with "same data, just slower"?
  3. A colleague suggests skipping straight to grepping the raw application logs from 09:14 onward instead of following the metric → trace → log order. Explain what makes that approach slower or less reliable here.

Share this question

Intermediate Open Pro

Compute an Error Budget and Decide a Release Policy

Unlock this question →
Intermediate Open Pro

Redesign a Paging Policy That's Causing Alert Fatigue

Unlock this question →
Intermediate Open Pro

Fix a Dashboard That Failed During an Incident

Unlock this question →
Intermediate Open Pro

Where Does General Observability End and ML Monitoring Begin?

Unlock this question →

We use cookies for product analytics to improve OmniAtlas. See our Privacy Policy.