Match a job Paths Subjects Questions Quizzes Pricing
Overview Read Practice

Practice — File Formats, Partitioning & Storage Layout (6 questions)

Pro content

Sign up free, then start a 14-day Pro trial — no card needed.

Intermediate Open Free

Diagnosing a Query That Got Slower After 'Better' Partitioning Permalink →

Your team's events table stores clickstream data and used to be partitioned only by event_date (roughly 50 GB of Parquet data per day, in a few dozen files per day, each 200–500 MB). A well-meaning engineer, trying to speed up a common query that filters by event_date and country, repartitions the table by event_date, country, device_type (device_type has ~40 distinct values, country has ~195). After the change:

  • The table now has roughly 700,000 partitions total across its history.
  • Listing files for a single day's query now takes noticeably longer than the actual data scan used to take.
  • Most files are now a few hundred kilobytes to a few megabytes.
  • The Glue Catalog is showing elevated latency on planning for this table specifically.
  1. Explain exactly why this change made things worse instead of better, in terms of partition cardinality.
  2. Propose a better partitioning and storage layout for this table that still serves the event_date + country query pattern well.
  3. If device_type genuinely needs to be a fast filter for a different set of queries, how would you support that without repeating this mistake?

Share this question

Intermediate Open Pro

Designing a Compaction Strategy for a Streaming Landing Zone

Unlock this question →
Intermediate Open Pro

Choosing a Format for a New Event Pipeline, End to End

Unlock this question →
Intermediate Open Pro

Predicate Pushdown Isn't Helping — Diagnose the Layout

Unlock this question →
Intermediate Open Pro

Picking a Compression Codec for Two Very Different Tables

Unlock this question →
Intermediate Open Free

What Happens to a 1 TB CSV in Parquet Permalink →

You convert a 1 TB CSV of typical tabular data (IDs, timestamps, categories, amounts) to Parquet with a standard codec like snappy or zstd.

Roughly how big is the Parquet output — and why is even that number not the real win?

Name the encoding tricks doing the work.

Share this question

We use cookies for product analytics to improve OmniAtlas. See our Privacy Policy.