Practice — Streaming Fundamentals & Kafka (6 questions)
Debugging Duplicate Events Despite an 'Exactly-Once' Pipeline Permalink →
Your team runs a Kafka-based order pipeline that reads from an
orders topic, computes a per-order total, and writes the result to a
Postgres order_totals table via a plain JDBC INSERT. Kafka
transactions are enabled end to end (idempotent producer,
transactional.id set, consumers read with read_committed), and the
team's documentation confidently describes the system as
"exactly-once." Finance reports that a small number of orders show up
twice in order_totals with identical totals, always following
brief periods of consumer-group rebalancing (deploys, pod restarts).
- Explain precisely why Kafka's transactional guarantees do not prevent this duplication, given what's described above.
- Identify the specific moment in the consume-process-write sequence where a rebalance-triggered crash produces a duplicate row.
- Propose a concrete fix, and explain why it closes the gap that Kafka's own exactly-once mechanisms don't cover.
Share this question
A Windowed Revenue Aggregate Is Undercounting After a Network Blip
Unlock this question →Consumer Lag Keeps Growing Even After Adding More Consumers
Unlock this question →Should This Inventory Sync Actually Be a Streaming Pipeline?
Unlock this question →Sizing Kafka Disk for a Billion Events a Day Permalink →
Your pipeline ingests 1 billion events/day, averaging 1 KB each, with 7-day retention and replication factor 3.
Roughly how much disk does the Kafka cluster need?
Show the arithmetic, and name the lever that cuts the number most in practice.
Share this question