Match a job Paths Subjects Questions Quizzes Pricing
Intermediate Open Free

What Happens to a 1 TB CSV in Parquet

You convert a 1 TB CSV of typical tabular data (IDs, timestamps, categories, amounts) to Parquet with a standard codec like snappy or zstd.

Roughly how big is the Parquet output — and why is even that number not the real win?

Name the encoding tricks doing the work.

Solution

C) ~130 GB — a 5–10× shrink is the norm, and it's still not the point.

Parquet is columnar, so each column's values sit together — and values within one column are highly self-similar, which is exactly what encoders love:

  • Dictionary encoding — repeated strings (countries, statuses, categories) become small integer codes.
  • Run-length and delta encoding — sorted IDs and timestamps store as tiny deltas; repeated values collapse into runs.
  • General compression on top — snappy/zstd compress each already-encoded column chunk far better than they compress row-mixed CSV text.

Typical tabular data lands 5–10× smaller — around 100–200 GB from 1 TB of CSV.

Why storage still isn't the real win: queries read only the columns they touch. A query on 3 of 50 columns scans ~6% of the column chunks, and min/max statistics per row group let it skip most of those. Your 1 TB scan becomes a few GB read — that's the 100× that shows up on the query bill, not the 7× on the storage bill.

The rule of thumb to say out loud: CSV → Parquet ≈ 5–10× smaller at rest, but scan bytes drop by column pruning × row-group skipping — size the win by what queries read, not what disks hold.

Share this question

← Back to File Formats, Partitioning & Storage Layout practice

We use cookies for product analytics to improve OmniAtlas. See our Privacy Policy.