Match a job Paths Subjects Questions Quizzes Pricing
Intermediate Open Free

Use Logs, Metrics, and Traces to Diagnose a Latency Spike

Your team's internal recommendation-scoring API has a dashboard showing p99 latency jumped from 55ms to 1.1s starting at 09:14, with no change in request rate or error rate. You have metrics, structured logs, and sampled traces available.

  1. Describe, in order, how you would use each of the three observability pillars to go from "p99 latency jumped" to a specific, actionable root cause. Be concrete about what each pillar would show you and why you look at them in that order.
  2. Suppose traces show the extra time is spent in a "feature_fetch" span, but the fetch itself isn't failing (no errors, correct data returned). What would you look at next, and what are two plausible root causes consistent with "same data, just slower"?
  3. A colleague suggests skipping straight to grepping the raw application logs from 09:14 onward instead of following the metric → trace → log order. Explain what makes that approach slower or less reliable here.
Solution

1. Order of investigation

Start from the metric that already told you something is wrong and roughly when (09:14) — that's given. Next, pull traces for a sample of slow requests from around 09:14: the trace's span breakdown shows which specific hop in the request path — feature fetch, model inference, response serialization — is consuming the extra time, narrowing "latency is up" to "latency is up specifically in span X." Only then go to logs, filtered to that specific span and time window, to read the structured fields (e.g., cache hit/miss, specific error codes, retry counts) that explain why that span got slow. The order matters because each step narrows the search space for the next: traces tell you where to look in the logs, and without that step you'd be grepping an undifferentiated flood of logs across every span and every request.

2. Same data, slower — next steps and causes

Look at the feature-fetch dependency's own RED metrics (its rate, error rate, and duration as measured from the caller's side) and, if available, resource-level (USE) metrics for whatever backs that fetch — a cache, a database, a connection pool. Two plausible causes consistent with "same data, just slower": (a) a cache miss/eviction — the feature store's cache was evicted or invalidated (e.g., by a deploy or a TTL expiry), so requests are now hitting a slower underlying store for correct-but-slow data, which would show up as a drop in cache-hit-rate in the logs or a dedicated metric; (b) connection pool saturation — enough concurrent requests are waiting for a free connection to the feature store that queueing time dominates even though each individual query is fast, which the USE method would catch as high saturation despite moderate utilization, something an averaged latency metric on the pool wouldn't reveal on its own.

3. Why skip-to-logs is slower/less reliable

Without the trace to narrow down which span is slow, grepping raw logs from 09:14 onward means wading through every log line from every part of the request path — most of which are perfectly normal, since only one hop is actually slow — looking for a signal you don't yet know the shape of. It also risks missing the answer entirely if the relevant structured field (e.g., cache_hit: false) isn't the first thing a human skimming free-text-adjacent logs happens to notice. The metric-then-trace-then-log order isn't just a convention; it's what lets each step eliminate most of the search space before the next, more detailed step begins — going straight to logs throws away the narrowing traces would have given you for free.

Share this question

← Back to Production Observability & Monitoring practice

We use cookies for product analytics to improve OmniAtlas. See our Privacy Policy.