The Attack Your Cache Hit Ratio Can't See
A public product-lookup API uses cache-aside: on a request for
product:{id}, check the cache; on a miss, query the database, and if
the product exists, write it into the cache. Your dashboard reports a
healthy 99% cache hit ratio and normal database load.
Overnight, a buggy partner integration (or an attacker probing for
valid IDs) starts sending millions of requests for random,
non-existent product IDs — product:9182739123, product:1029384756,
and so on. The database's CPU spikes hard even though the dashboard's
hit ratio barely moves. Why does the cache do nothing to absorb this
traffic, and what's the fix?
Cache negative results ('not found') too, with a short TTL (negative caching).
Cache-aside as described only ever writes an entry when the database lookup succeeds. Every request for a non-existent ID is, by definition, a miss on every single attempt — there is nothing to cache after checking, because the product was never found, so the exact same "miss → query DB → still not found → cache stays empty" cycle repeats identically for every one of the millions of requests. Each one reaches the database directly, at full volume, with no absorption at all.
This is why the dashboard's hit ratio doesn't move: it's typically computed only over legitimate, cacheable traffic (or the random IDs are simply a separate, uncounted flood), so a cache that is technically "working" for real products can coexist with a database being hammered by a completely different traffic pattern the cache was never built to intercept. This failure mode has a name — cache penetration — and it's distinct from a cache stampede (many concurrent requests racing to repopulate one real, expired key) or a hot key (many requests for one real, existing key): here, the requested keys don't exist at all, so nothing ever gets cached under the naive design.
The fix: when a lookup genuinely returns "not found," write a
sentinel value (e.g. NULL or a small marker) into the cache for that
key with a short TTL (seconds to low minutes — long enough to absorb
a repeat flood of the same bogus ID, short enough that a product
created moments later becomes visible quickly). Subsequent requests
for the same non-existent ID hit the cache's negative entry and never
reach the database. A complementary defense for a large, adversarial
ID space is a Bloom filter of valid IDs in front of the cache, so
even the first request for a never-valid ID can be rejected before
touching the database at all.
Share this question