Practice — Spark Architecture & Execution Model (5 questions)
Intermediate
Open
Free
Why Does Nothing Happen Until `.count()` Runs? Permalink →
A new engineer on your team writes the following PySpark script and is confused by what they see in the Spark UI:
df = spark.read.parquet("s3://bucket/clickstream/")
step1 = df.filter(F.col("event") == "click")
step2 = step1.withColumn("session_min", F.col("ts") / 60)
step3 = step2.groupBy("page_id").agg(F.count("*").alias("clicks"))
step4 = step3.orderBy(F.desc("clicks"))
print("built the pipeline") # prints almost instantly
top_pages = step4.limit(10).collect() # this line takes 4 minutes
They ask: "Why does building four chained operations on a huge S3
dataset finish in milliseconds, but the very last line takes minutes?
Doesn't each .filter() and .groupBy() have to touch the data to
produce its result?"
- Explain precisely what happens (and doesn't happen) at each of the
four
step*lines. - Explain what happens once
.collect()is called, including which Spark components get involved that weren't involved before. - If they instead wanted to see intermediate progress after each step for debugging, what would you tell them to do, and what would it cost them?
Share this question
Intermediate
Open
Pro
A Job That Ran Fine Across the Cluster Crashes on the Last Line
Unlock this question →
Intermediate
Open
Pro
Diagnosing a 40-Minute Straggler Task in an Otherwise 2-Minute Job
Unlock this question →
Intermediate
Open
Pro