Match a job Paths Subjects Questions Quizzes Pricing
Intermediate Open Free

Your '1,000 req/min' Limit Is Actually 40,000

An engineer implements "1,000 requests/minute per API key" as an in-process token bucket: a plain dictionary living inside each gateway process, refilled and checked entirely in memory, with no shared datastore. It's deployed across a fleet of 40 identical, stateless gateway instances behind a load balancer that spreads each client's requests round-robin across all 40.

In the worst case — a client whose requests happen to spread evenly across every instance — what is the actual limit enforced on that API key across the whole fleet?

Solution

~40,000 req/min — the limit multiplies by the instance count.

Each of the 40 gateway processes independently maintains its own token bucket for the key, with no shared state and no communication between instances. Every single one of them will happily allow up to 1,000 requests/minute through itself, because from any one instance's local point of view, it has never seen more than its own share of the traffic. A client whose requests are spread evenly round-robin across all 40 instances can therefore get up to 1{,}000 \times 40 = 40{,}000 requests/minute admitted in total, before any single instance's own local bucket would deny anything — the configured number describes what one process enforces, not what the fleet enforces, and nothing in this design makes those the same thing.

This is a different failure from the classic distributed check-then-act race (two processes racing on a shared Redis counter, briefly overshooting by a small margin before converging): here there's no race at all, no bug in any single instance's logic, and no transient overshoot — it's a systematic, permanent multiplication baked into the architecture, worse the more instances you run, and it gets worse over time as the fleet autoscales up under load, which is exactly when the limit matters most.

The fix is a shared source of truth: a centralized store (Redis, as in the fleet-wide INCR-based limiter) that every instance checks against, so the limit is enforced once, globally, rather than once per instance. If a shared store is genuinely unavailable and a local fallback is required (for example, during an outage), the local rate must be explicitly divided by the fleet size — sustained_rate / instance_count per instance — so the sum across instances approximates the intended global limit, rather than each instance independently re-granting the full configured number.

Share this question

← Back to Rate Limiting practice

We use cookies for product analytics to improve OmniAtlas. See our Privacy Policy.