Intermediate
Open
Pro
Reading Latency: Averages, Percentiles and Fan-out
A dashboard shows the checkout API's average latency as a steady 120 ms, but support tickets say "checkout sometimes hangs for 5+ seconds". A sample of 1,000 requests from one minute shows: 950 requests at ~90 ms, 40 requests at ~400 ms, and 10 requests at ~5,000 ms.
- Compute the mean, p50, p95 and p99 for that sample. What does each tell you and why did the dashboard hide the problem?
- The checkout page renders by calling 20 backend services in parallel and waits for all of them. Each backend has p99 = 300 ms. What fraction of page loads will experience at least one backend at or above its p99? What does that imply for the page's own p99?
- Given the sample above, which observability signals (metrics, logs, traces) would you use in what order to find the cause of the 5 s requests, and what specifically would you look for?
Share this question