Intermediate
Open
Pro
Sizing Partitions and Draining a Backlog
A clickstream topic receives a steady 30,000 events/s. Each consumer
in the enrichment group processes about 2,500 events/s (after batching).
The topic was created with 8 partitions. On Monday morning a downstream
outage stops the group for 40 minutes; when it recovers, on-call adds
consumers to catch up quickly.
- Before the outage, is the group keeping up? Show the arithmetic.
- How large is the backlog after 40 minutes, in events and in "seconds of lag"?
- On-call scales the group to 20 consumers. What throughput does the group actually achieve, and how long does it take to drain? What would you have to change to drain in under 10 minutes?
- What lag metric and alert would have caught this earlier?
Share this question