Databricks Data Engineer Associate · Free practice question 4 of 12
collect_set vs collect_list aggregation
For each customer, an analyst at Wicklow Grocers needs an array of the distinct store IDs where the customer shopped, with no repeated values. Which aggregate function should be used with GROUP BY customer_id?
- A.collect_set(store_id)
- B.collect_list(store_id)
- C.array(store_id)
- D.explode(store_id)
Show answer and explanation
Correct answer: A. collect_set(store_id)
Why: collect_set aggregates values into an array without duplicates. collect_list keeps every value, including repeats. array builds an array from its arguments within a single row, and explode turns array elements into rows instead of aggregating.
More free Databricks Data Engineer Associate questions
- Delta shallow vs deep clone
- Auto Loader schema hints
- Reading notebook task parameters
- Triggered vs continuous pipeline mode
- Job maximum concurrent runs
- Job task timeout setting
- SHOW GRANTS on Unity Catalog objects
- Transferring object ownership
- Compute policies for cluster governance
- ALTER TABLE ADD COLUMNS on Delta
- Spark lazy evaluation transformations vs actions