Practice — Case Study: Design a Streaming Event Pipeline (5 questions)
Advanced
Open
Free
A Consumer Group Is Falling Behind Under Load Permalink →
Your events.commerce topic has 200 partitions, and the Flink job
computing revenue-per-minute reads from all of them with parallelism
200 (one subtask per partition). During a flash-sale campaign,
throughput spikes to 1.6x the normal daily peak for about two hours.
Consumer lag on this job climbs steadily throughout the spike and does
not recover until roughly 40 minutes after traffic returns to normal —
meaning the revenue dashboard was stale by up to that much during and
after the sale, exactly when the business cared most about it.
- Walk through the likely root causes of the lag growth, and how you'd confirm which one is actually happening rather than guessing.
- Propose a fix, or combination of fixes, and explain the trade-off each one makes.
- The case study sizes 200 partitions off the sustained peak (~70,000 events/sec), not the campaign spike (~120,000 events/sec). Was that the wrong call? Justify your answer either way.
Share this question
Advanced
Open
Pro