Practice — Spark Performance Tuning (5 questions)
Advanced
Open
Free
A GroupBy Aggregation Hangs at 799/800 Tasks Permalink →
A nightly job aggregates total watch-time per device_id from a
120GB events table:
result = (
events
.groupBy("device_id")
.agg(F.sum("watch_seconds").alias("total_watch_seconds"))
)
result.write.parquet(output_path)
In the Spark UI, the aggregation stage shows 799 of 800 tasks
completing in 20-40 seconds each. The 800th task is still running
after 90 minutes. Its "Shuffle Read Size / Records" column shows
340,000,000 records, versus a median of roughly 600,000 records for
the other 799 tasks. Data-team context: device_id is nullable, and
roughly 15% of raw events fail device attribution and are landed with
device_id = NULL.
- Diagnose the root cause precisely — not just "it's skewed," but why this specific key ended up this large.
- Propose a fix, including whether salting is appropriate here and why (or why not), and show the code.
- Is there a fix that's simpler than salting for this specific case, given what you know about the NULL key? Would you use it instead, or in addition to a general skew-handling strategy?
Share this question
Advanced
Open
Pro