What Happens to a 1 TB CSV in Parquet
You convert a 1 TB CSV of typical tabular data (IDs, timestamps, categories, amounts) to Parquet with a standard codec like snappy or zstd.
Roughly how big is the Parquet output — and why is even that number not the real win?
Name the encoding tricks doing the work.
C) ~130 GB — a 5–10× shrink is the norm, and it's still not the point.
Parquet is columnar, so each column's values sit together — and values within one column are highly self-similar, which is exactly what encoders love:
- Dictionary encoding — repeated strings (countries, statuses, categories) become small integer codes.
- Run-length and delta encoding — sorted IDs and timestamps store as tiny deltas; repeated values collapse into runs.
- General compression on top — snappy/zstd compress each already-encoded column chunk far better than they compress row-mixed CSV text.
Typical tabular data lands 5–10× smaller — around 100–200 GB from 1 TB of CSV.
Why storage still isn't the real win: queries read only the columns they touch. A query on 3 of 50 columns scans ~6% of the column chunks, and min/max statistics per row group let it skip most of those. Your 1 TB scan becomes a few GB read — that's the 100× that shows up on the query bill, not the 7× on the storage bill.
The rule of thumb to say out loud: CSV → Parquet ≈ 5–10× smaller at rest, but scan bytes drop by column pruning × row-group skipping — size the win by what queries read, not what disks hold.
Share this question