Databricks Data Engineer Professional Practice Exam
Practice questions for the Databricks Certified Data Engineer Professional certification: Structured Streaming in Python and SQL (triggers including real-time mode, watermarks, stateful operations and transformWithState, checkpoints, foreachBatch, stream-static and stream-stream joins), pandas UDFs and Unity Catalog Python UDFs, parameterised SQL, ingestion with Auto Loader at scale, Kafka and Lakeflow Connect, MERGE patterns, SCD Type 1 and 2, change data feed and AUTO CDC (formerly APPLY CHANGES) in Lakeflow Spark Declarative Pipelines (formerly Delta Live Tables), expectations and quarantine tables, OpenSharing (formerly Delta Sharing), Lakehouse Federation and Clean Rooms, monitoring with system tables, data quality monitoring, the Spark UI and job health rules, cost and performance tuning with liquid clustering, data skipping, deletion vectors, file sizing, skew and spill, and serverless performance modes, security and compliance with column masks, ABAC policies, secrets and GDPR deletion, governance with tags, lineage and managed storage, deployment with Declarative Automation Bundles (formerly Databricks Asset Bundles), job parameters and unit tests, and Delta table design for star schemas. Every question includes a written explanation.
100 questions · 12 free preview
Studying more than one? All Databricks exams for $29 · every exam for $79
Free sample questions
- Sample · question 1 · Databricks SQL alerts with notifications
The operations team at Bramwell Energy wants a Slack message whenever the number of rows with a NULL meter_id loaded into energy.silver.readings in the past hour exceeds 500. The check should run every 15 minutes on a SQL warehouse. Which feature fits best?
- A.A NOT NULL constraint on meter_id
- B.A job duration warning threshold
- C.A file arrival trigger on the bronze landing path
- D.A Databricks SQL alert on a query that counts the NULL rows, with a threshold condition, a 15-minute schedule and a Slack notification destinationcorrect
Why: Databricks SQL alerts run a query on a schedule, evaluate a condition on its result, and notify configured destinations such as email, Slack or webhooks when the condition is met. A NOT NULL constraint would reject the writes rather than count and report them, duration thresholds monitor how long jobs take, and file arrival triggers start jobs.
Open this question on its own page → - Sample · question 2 · Partitioning guidance for mid-size tables
A new Delta table at Cragside Analytics will hold about 400 GB, and an engineer proposes partitioning it by customer_id, which has around 2 million distinct values. What does Databricks guidance recommend?
- A.Partition by customer_id, because more partitions always improve data skipping
- B.Partition by customer_id and run OPTIMIZE ZORDER BY on the same column
- C.Do not partition a table of this size, especially on a high-cardinality column; use liquid clustering insteadcorrect
- D.Partition by a hash of customer_id into 2 million buckets
Why: Databricks recommends against partitioning most tables under about 1 TB, and a partition column should leave each partition with at least about 1 GB of data; a high-cardinality column produces millions of tiny partitions and small files. Liquid clustering gives data skipping on customer_id without those problems. Z-ordering cannot be applied to partition columns, and hash buckets create the same small-file problem.
Open this question on its own page → - Sample · question 3 · VACUUM retention duration safety check
To reclaim storage quickly, an engineer at Thurlby Media runs VACUUM media.silver.views RETAIN 24 HOURS, and the command fails with a safety check error. What does this indicate?
- A.Retention below the 7-day default is blocked by a safety check, because removing recent files can break concurrent readers, writers and streams; it can be overridden only by disabling spark.databricks.delta.retentionDurationCheck.enabledcorrect
- B.VACUUM cannot be run on tables with liquid clustering
- C.The RETAIN clause accepts only days, not hours
- D.VACUUM requires the table to have no active streaming readers
Why: Delta refuses VACUUM retention intervals shorter than the default 168 hours unless the retentionDurationCheck safety setting is turned off, because long-running queries and streams may still need files that are no longer current. Disabling the check should be done only when no operation can depend on older files. RETAIN takes a number of hours, and VACUUM works on clustered tables.
Open this question on its own page → - Sample · question 4 · SQL warehouse sizing versus scaling
On a Databricks SQL warehouse at Elmstead Finance, a single complex month-end query runs slowly while nothing else is queued, and dozens of short dashboard queries on another warehouse wait in a queue at 9 a.m. each day. Which changes address each problem?
- A.Increase the maximum cluster count for the month-end warehouse and the cluster size for the dashboard warehouse
- B.Enable auto stop on both warehouses
- C.Move both workloads to a single all-purpose cluster
- D.Increase the cluster size for the month-end warehouse and the maximum cluster count (scaling) for the dashboard warehousecorrect
Why: A larger warehouse size adds compute to each cluster, which helps an individual complex query, while raising the maximum number of clusters lets the warehouse scale out to serve many concurrent queries and shorten queues. Swapping them gives the dashboards bigger clusters that still queue and gives the month-end query more clusters it cannot use. Auto stop saves cost when idle but does not improve performance.
Open this question on its own page → - Sample · question 5 · BROWSE privilege for data discovery
Business users at Holmfirth Retail should be able to discover the tables in the catalog sales in Catalog Explorer, seeing their names, descriptions and tags, so they can request access, without being able to read any data. Which privilege fits?
- A.SELECT on the catalog sales
- B.BROWSE on the catalog salescorrect
- C.USE CATALOG on the catalog sales
- D.MANAGE on the catalog sales
Why: BROWSE lets principals view metadata of objects in a catalog, for example in Catalog Explorer, search results and lineage, without data access and without needing USE CATALOG or USE SCHEMA. SELECT grants data access, USE CATALOG alone is a prerequisite for using objects rather than a discovery grant, and MANAGE allows managing privileges on the objects.
Open this question on its own page → - Sample · question 6 · UNDROP TABLE for managed tables
A data engineer at Wigton Freight accidentally dropped the Unity Catalog managed table ops.gold.route_costs two days ago. How can it be recovered with its data and history?
- A.RESTORE TABLE ops.gold.route_costs TO VERSION AS OF 0
- B.UNDROP TABLE ops.gold.route_costscorrect
- C.Recreate the table and run a full refresh from bronze
- D.Read the table's files from the cloud storage path and run CONVERT TO DELTA
Why: Unity Catalog keeps dropped managed tables for a recovery period (7 days by default), during which UNDROP TABLE restores them with their data, history and permissions. RESTORE works only on a table that still exists, rebuilding from bronze may not reproduce the same results, and managed table storage paths are not meant to be accessed directly.
Open this question on its own page → - Sample · question 7 · Workspace-catalog binding isolation
Selby Health's prod catalog must be accessible only from the production workspace, even though development workspaces share the same Unity Catalog metastore. How should this be enforced?
- A.Revoke USE CATALOG from all users in development workspaces
- B.Create a separate metastore for the production workspace in the same region
- C.Make the catalog isolated and bind it to the production workspace only (workspace-catalog binding)correct
- D.Use a cluster policy in the development workspaces that blocks the catalog
Why: Workspace-catalog binding lets an admin restrict a catalog to specific workspaces, so it cannot be accessed from other workspaces attached to the same metastore, whatever privileges users hold. Grants apply to principals, not workspaces, and the same users often work in both. A region supports one metastore per account, and cluster policies do not control catalog access.
Open this question on its own page → - Sample · question 8 · Bundle generate and deployment bind
Ludgate Analytics has a production job that was built in the UI, and it now wants to manage the job in a Declarative Automation Bundle (formerly called a Databricks Asset Bundle) without creating a duplicate job. Which approach fits?
- A.Use databricks bundle generate job with the existing job ID to create the configuration, then link the bundle resource to the existing job with databricks bundle deployment bindcorrect
- B.Export the job JSON and paste it into databricks.yml as a new job, then delete the original after the first deployment
- C.Run databricks bundle destroy and recreate the job from the bundle
- D.Clone the job in the UI and point the clone at the bundle folder
Why: bundle generate creates bundle configuration, and downloads referenced notebooks, from an existing job or pipeline, and bundle deployment bind links a bundle resource to that existing workspace object, so the next deployment updates it in place instead of creating a duplicate. Pasting definitions creates a new job with a new ID and run history, and destroying or cloning the job loses its continuity.
Open this question on its own page → - Sample · question 9 · Driver memory and toPandas
A notebook at Ripley Outdoor ends with pdf = spark.read.table("sales.silver.orders").toPandas() on a 900 million-row table, and the job fails with a driver out-of-memory error even though the executors have plenty of free memory. What is the best fix?
- A.Increase the number of worker nodes
- B.Raise spark.sql.shuffle.partitions
- C.Keep the processing in Spark, aggregating, filtering or writing results to a table, and bring only small results to the drivercorrect
- D.Enable Photon on the cluster
Why: toPandas() and collect() move the entire result to the driver's memory, so a large table exhausts the driver regardless of executor capacity. The work should stay distributed, with only small, aggregated or limited results converted to pandas, or pandas logic moved into pandas UDFs or the pandas API on Spark. More workers, shuffle partitions or Photon do not change how much data lands on the driver.
Open this question on its own page → - Sample · question 10 · Continuous job trigger for streaming
Ottery Media has a Lakeflow Job that runs a Structured Streaming query that should always be running. If the run fails, a new run should start automatically, and there should never be more than one active run. Which trigger type fits?
- A.A continuous trigger on the jobcorrect
- B.A cron schedule every minute with maximum concurrent runs set to 1
- C.A file arrival trigger on the source path
- D.A table update trigger on the target table
Why: The continuous trigger keeps exactly one run of the job active at all times: when a run ends or fails, a new one starts, with backoff applied after repeated failures. A one-minute schedule creates pointless scheduling attempts and delays restarts, and file arrival or table update triggers start runs in response to events rather than keeping a stream running.
Open this question on its own page → - Sample · question 11 · Table update triggers for jobs
A downstream Lakeflow Job at Kelmscott Foods should start only after both sales.gold.orders and sales.gold.returns have been updated by other teams' pipelines, whose finish times vary. Which trigger does this without polling code?
- A.A cron schedule set after the latest expected finish time
- B.A table update trigger on both tables, configured to fire when all of the tables are updatedcorrect
- C.A continuous trigger that checks the tables in a loop
- D.A file arrival trigger on the tables' storage location
Why: Table update triggers start a job when monitored Unity Catalog tables are updated; with several tables, the trigger can fire when any table or when all tables have been updated, and options such as the wait after the last change smooth out bursts of commits. A fixed schedule either waits too long or runs too early, a continuous polling loop wastes compute, and file arrival triggers watch storage paths rather than table commits.
Open this question on its own page → - Sample · question 12 · Serverless environment dependencies
A Python task at Burwell Logistics is moving from classic job compute to serverless compute for jobs. On classic compute, it installed a private wheel and two PyPI packages with a cluster init script. How should the dependencies be provided on serverless?
- A.Keep the init script, because serverless runs it at start-up
- B.Install the packages with an apt-get command in the first notebook cell
- C.Ask a workspace admin to install the packages on the serverless control plane
- D.Declare the wheel and packages as dependencies in the job's serverless environmentcorrect
Why: Serverless compute for jobs does not support init scripts; Python dependencies are declared in the task's environment, which specifies an environment version and a dependency list such as PyPI packages and wheel files in volumes or workspace files, and Databricks installs them when the environment is created. System-level installs such as apt-get are not a supported way to provide dependencies on serverless, and admins do not install packages on serverless infrastructure.
Open this question on its own page →
Like the sample?
Other practice exams
- AnthropicClaude Certified Architect — Foundations100 questions · $19
- CompTIACompTIA Security+ (SY0-701)100 questions · $19
- ISC2CISSP100 questions · $19
- DatabricksDatabricks Data Engineer Associate100 questions · $19
- Google CloudGoogle Cloud Associate Cloud Engineer100 questions · $19
- Google CloudGoogle Cloud Professional Cloud Architect100 questions · $19
- Google CloudGoogle Cloud Professional Data Engineer100 questions · $19
- HashiCorpTerraform Associate (004)100 questions · $19
- Microsoft Power BI & FabricPower BI Data Analyst (PL-300)100 questions · $19
- Microsoft Power BI & FabricFabric Analytics Engineer (DP-600)100 questions · $19
- SnowflakeSnowPro Core (COF-C03)250 questions · $19
- SnowflakeSnowPro Advanced: Data Engineer100 questions · $19
- SnowflakeSnowPro Advanced: Architect100 questions · $19
- AWSAWS Cloud Practitioner (CLF-C02)100 questions · $19
- AWSAWS Solutions Architect Associate (SAA-C03)100 questions · $19
- AWSAWS AI Practitioner (AIF-C01)100 questions · $19
- Microsoft AzureAzure Fundamentals (AZ-900)100 questions · $19
- Microsoft AzureAzure Administrator (AZ-104)100 questions · $19
- Microsoft AzureAzure AI Fundamentals (AI-901)100 questions · $19
- Microsoft AzureAzure Solutions Architect Expert (AZ-305)100 questions · $19