Practice — Rate Limiting (6 questions)
Fixed Window Boundary Burst Permalink →
An internal team implements a "100 requests per minute per API key"
limiter using a fixed window: key = f"{api_key}:{floor(now/60)}",
INCR the key, allow if the result is <= 100.
A client sends 100 requests at 08:59:58–08:59:59 and another 100 at 09:00:01–09:00:02.
- How many requests does the limiter allow in that 4-second span, and why?
- Redesign the check using the sliding window counter algorithm and show the estimate calculation for a request arriving at 09:00:02 (2 seconds into the new window), given the previous window ended with 100 requests and the current window has 5 so far.
- Would a sliding window log have been a better fix here? What would it cost?
Share this question
Configuring a Token Bucket for Burst and Sustained Rate
Unlock this question →Your '1,000 req/min' Limit Is Actually 40,000 Permalink →
An engineer implements "1,000 requests/minute per API key" as an in-process token bucket: a plain dictionary living inside each gateway process, refilled and checked entirely in memory, with no shared datastore. It's deployed across a fleet of 40 identical, stateless gateway instances behind a load balancer that spreads each client's requests round-robin across all 40.
In the worst case — a client whose requests happen to spread evenly across every instance — what is the actual limit enforced on that API key across the whole fleet?
Share this question