CertKeen

Databricks Data Engineer Professional · Free practice question 9 of 12

Driver memory and toPandas

A notebook at Ripley Outdoor ends with pdf = spark.read.table("sales.silver.orders").toPandas() on a 900 million-row table, and the job fails with a driver out-of-memory error even though the executors have plenty of free memory. What is the best fix?

  1. A.Increase the number of worker nodes
  2. B.Raise spark.sql.shuffle.partitions
  3. C.Keep the processing in Spark, aggregating, filtering or writing results to a table, and bring only small results to the driver
  4. D.Enable Photon on the cluster
Show answer and explanation

Correct answer: C. Keep the processing in Spark, aggregating, filtering or writing results to a table, and bring only small results to the driver

Why: toPandas() and collect() move the entire result to the driver's memory, so a large table exhausts the driver regardless of executor capacity. The work should stay distributed, with only small, aggregated or limited results converted to pandas, or pandas logic moved into pandas UDFs or the pandas API on Spark. More workers, shuffle partitions or Photon do not change how much data lands on the driver.

More free Databricks Data Engineer Professional questions