Practice — File Formats, Partitioning & Storage Layout (6 questions)
Diagnosing a Query That Got Slower After 'Better' Partitioning Permalink →
Your team's events table stores clickstream data and used to be
partitioned only by event_date (roughly 50 GB of Parquet data per
day, in a few dozen files per day, each 200–500 MB). A well-meaning
engineer, trying to speed up a common query that filters by
event_date and country, repartitions the table by
event_date, country, device_type (device_type has ~40 distinct
values, country has ~195). After the change:
- The table now has roughly 700,000 partitions total across its history.
- Listing files for a single day's query now takes noticeably longer than the actual data scan used to take.
- Most files are now a few hundred kilobytes to a few megabytes.
- The Glue Catalog is showing elevated latency on planning for this table specifically.
- Explain exactly why this change made things worse instead of better, in terms of partition cardinality.
- Propose a better partitioning and storage layout for this table
that still serves the
event_date+countryquery pattern well. - If
device_typegenuinely needs to be a fast filter for a different set of queries, how would you support that without repeating this mistake?
Share this question
Designing a Compaction Strategy for a Streaming Landing Zone
Unlock this question →Picking a Compression Codec for Two Very Different Tables
Unlock this question →What Happens to a 1 TB CSV in Parquet Permalink →
You convert a 1 TB CSV of typical tabular data (IDs, timestamps, categories, amounts) to Parquet with a standard codec like snappy or zstd.
Roughly how big is the Parquet output — and why is even that number not the real win?
Name the encoding tricks doing the work.
Share this question