Match a job Paths Subjects Questions Quizzes Pricing
Overview Read Practice

Practice — Spark Architecture & Execution Model (5 questions)

Pro content

Sign up free, then start a 14-day Pro trial — no card needed.

Intermediate Open Free

Why Does Nothing Happen Until `.count()` Runs? Permalink →

A new engineer on your team writes the following PySpark script and is confused by what they see in the Spark UI:

df = spark.read.parquet("s3://bucket/clickstream/")
step1 = df.filter(F.col("event") == "click")
step2 = step1.withColumn("session_min", F.col("ts") / 60)
step3 = step2.groupBy("page_id").agg(F.count("*").alias("clicks"))
step4 = step3.orderBy(F.desc("clicks"))

print("built the pipeline")   # prints almost instantly

top_pages = step4.limit(10).collect()   # this line takes 4 minutes

They ask: "Why does building four chained operations on a huge S3 dataset finish in milliseconds, but the very last line takes minutes? Doesn't each .filter() and .groupBy() have to touch the data to produce its result?"

  1. Explain precisely what happens (and doesn't happen) at each of the four step* lines.
  2. Explain what happens once .collect() is called, including which Spark components get involved that weren't involved before.
  3. If they instead wanted to see intermediate progress after each step for debugging, what would you tell them to do, and what would it cost them?

Share this question

Intermediate Open Pro

A Job That Ran Fine Across the Cluster Crashes on the Last Line

Unlock this question →
Intermediate Open Pro

Diagnosing a 40-Minute Straggler Task in an Otherwise 2-Minute Job

Unlock this question →
Intermediate Open Pro

Should This Nightly Report Run on Spark or a Single Machine?

Unlock this question →
Intermediate Open Pro

Reading an `explain()` Plan to Explain a Slow Join

Unlock this question →

We use cookies for product analytics to improve OmniAtlas. See our Privacy Policy.