Match a job Paths Subjects Questions Quizzes Pricing
Intermediate Open Free

Sizing Kafka Disk for a Billion Events a Day

Your pipeline ingests 1 billion events/day, averaging 1 KB each, with 7-day retention and replication factor 3.

Roughly how much disk does the Kafka cluster need?

Show the arithmetic, and name the lever that cuts the number most in practice.

Solution

C) ~21 TB — retention and replication, not throughput, are what eat the disk.

Step by step:

Step Math Result
Daily raw volume 10⁹ events × 1 KB ~1 TB/day
× retention 1 TB × 7 days 7 TB
× replication 7 TB × 3 21 TB

Note what didn't matter: 1B events/day is only ~11,600 events/sec — about 12 MB/s of write throughput, trivial for a modest cluster. The disk bill comes entirely from how long you keep the data and how many copies you hold.

The rule of thumb to say out loud: Kafka disk ≈ daily volume × retention days × replication factor — then add ~40% headroom so a traffic spike doesn't fill the volumes.

The biggest lever:

  • Compression — lz4/zstd on typical JSON events compresses 3–5×, taking 21 TB down to ~5–7 TB. Producers compress in batches; Kafka stores and replicates the compressed form.
  • Also worth naming: shorter retention with tiered storage (offload old segments to object storage), and log compaction for keyed topics where only the latest value per key matters.

Share this question

← Back to Streaming Fundamentals & Kafka practice

We use cookies for product analytics to improve OmniAtlas. See our Privacy Policy.