CertKeen

Databricks Data Engineer Professional · Free practice question 2 of 12

Partitioning guidance for mid-size tables

A new Delta table at Cragside Analytics will hold about 400 GB, and an engineer proposes partitioning it by customer_id, which has around 2 million distinct values. What does Databricks guidance recommend?

  1. A.Partition by customer_id, because more partitions always improve data skipping
  2. B.Partition by customer_id and run OPTIMIZE ZORDER BY on the same column
  3. C.Do not partition a table of this size, especially on a high-cardinality column; use liquid clustering instead
  4. D.Partition by a hash of customer_id into 2 million buckets
Show answer and explanation

Correct answer: C. Do not partition a table of this size, especially on a high-cardinality column; use liquid clustering instead

Why: Databricks recommends against partitioning most tables under about 1 TB, and a partition column should leave each partition with at least about 1 GB of data; a high-cardinality column produces millions of tiny partitions and small files. Liquid clustering gives data skipping on customer_id without those problems. Z-ordering cannot be applied to partition columns, and hash buckets create the same small-file problem.

More free Databricks Data Engineer Professional questions