Paths Subjects Questions Quizzes Pricing Search
Data Engineering Intermediate Pro

File Formats, Partitioning & Storage Layout

Row vs columnar storage, Parquet internals and predicate pushdown, Avro schema evolution, compression codec trade-offs, partitioning strategy, clustering, and the small-files problem

20 min read 6 views

A practitioner's tour of how data is physically laid out on disk and why that layout is often the single biggest lever on query cost: row-oriented vs columnar storage and why analytics workloads favor the latter, Parquet's row-group and column-chunk structure and how embedded statistics enable predicate pushdown, Avro's role in streaming and schema evolution, a brief look at ORC, the speed-vs-ratio trade-off across snappy, gzip, and zstd, partitioning strategy and the cardinality pitfalls of over-partitioning, clustering and sort order within files, and the small-files problem with concrete compaction strategies.

Practice questions (5)

  • Diagnosing a Query That Got Slower After 'Better' Partitioning

    Intermediate · Free
    View →
  • Designing a Compaction Strategy for a Streaming Landing Zone

    Intermediate
    View →
  • Choosing a Format for a New Event Pipeline, End to End

    Intermediate
    View →
  • Predicate Pushdown Isn't Helping — Diagnose the Layout

    Intermediate
    View →
  • Picking a Compression Codec for Two Very Different Tables

    Intermediate
    View →

We use cookies for product analytics to improve OmniAtlas. See our Privacy Policy.