System Design
Intermediate
Pro
Reliability, Resilience & Observability Patterns
Design services that stay up when their dependencies don't — and prove it with the right signals
30 min read
12 views
Learn availability math, SLOs and error budgets, timeouts, retries with backoff and jitter, circuit breakers, bulkheads, load shedding, safe deploys, and the three pillars of observability, then apply them to a checkout service with a flaky payment provider.
Practice questions (5)
-
View →
Availability of a Dependency Chain
Intermediate · Free -
View →
Diagnosing a Retry Storm
Intermediate -
View →
Setting an SLO and Alerting on Burn Rate
Intermediate -
View →
The Readiness Check That Took Down Everything
Intermediate -
View →
Reading Latency: Averages, Percentiles and Fan-out
Intermediate