Match a job Paths Subjects Questions Quizzes Pricing
Overview Read Practice

Practice — Spark Performance Tuning (5 questions)

Pro content

Sign up free, then start a 14-day Pro trial — no card needed.

Advanced Open Free

A GroupBy Aggregation Hangs at 799/800 Tasks Permalink →

A nightly job aggregates total watch-time per device_id from a 120GB events table:

result = (
    events
    .groupBy("device_id")
    .agg(F.sum("watch_seconds").alias("total_watch_seconds"))
)
result.write.parquet(output_path)

In the Spark UI, the aggregation stage shows 799 of 800 tasks completing in 20-40 seconds each. The 800th task is still running after 90 minutes. Its "Shuffle Read Size / Records" column shows 340,000,000 records, versus a median of roughly 600,000 records for the other 799 tasks. Data-team context: device_id is nullable, and roughly 15% of raw events fail device attribution and are landed with device_id = NULL.

  1. Diagnose the root cause precisely — not just "it's skewed," but why this specific key ended up this large.
  2. Propose a fix, including whether salting is appropriate here and why (or why not), and show the code.
  3. Is there a fix that's simpler than salting for this specific case, given what you know about the NULL key? Would you use it instead, or in addition to a general skew-handling strategy?

Share this question

Advanced Open Pro

A Join Stage Is Spilling Tens of GB to Disk Per Task

Unlock this question →
Advanced Open Pro

Adding .cache() Made the Pipeline Slower, Not Faster

Unlock this question →
Advanced Open Pro

Yesterday's Output Created 40,000 Tiny Files and Today's Read Job Is Slow

Unlock this question →
Advanced Open Pro

Raising the Broadcast Threshold Caused Executor OOMs

Unlock this question →

We use cookies for product analytics to improve OmniAtlas. See our Privacy Policy.