Microsoft Implementing Data Engineering Solutions Using Azure Databricks DP-750 Exam Questions

Page: 1 / 14
Total 91 questions
Question 1

You have an Azure Databricks workspace that contains an all-purpose cluster named Cluster! You need to configure Cluster1 to meet the following requirements;

* The cluster must scale up automatically when workloads increase.

* The cluster must scale down automatically when workloads decrease.

The solution must minimize costs.

Which two actions should you perform? Each correct answer presents part of the solution.

NOTE: Each correct selection is worth one point.



Answer : C, D

The correct answers are C and D. Together they deliver cost-efficient autoscaling:

D (Enable autoscaling) allows the cluster to grow when workloads increase and shrink when they ease off. This satisfies both scale-up and scale-down requirements without manual intervention.

C (Auto-termination after 30 minutes of inactivity) ensures the cluster stops entirely when no work is running, eliminating the cost of an idle cluster. This is the cheapest possible state.

Option A (disable Photon) reduces compute acceleration --- that's a performance regression with no meaningful cost benefit for autoscaling. Option B (compute policy that lets users manage settings) adds governance overhead and doesn't address scaling behaviour. Option E (fixed number of workers) is the opposite of autoscaling --- a static worker count that either over-provisions during quiet periods or under-provisions during peaks.


Question 2

You have a Lakeflow Spark Declarative Pipelines {SDP) pipeline in Azure Databricks. The pipeline ingests transaction data into a table named Table1.

You need to ensure that in the event of an invalid record, the pipeline continues to run. The solution must meet the following requirements:

* Invalid records must NOT be written to Table 1.

* Invalid records must be preserved for review.

* Minimize development effort

What should you do?



Answer : B

The correct answer is B --- define a pipeline expectation.

SDP expectations with @dlt.expect_or_drop are built precisely for this scenario: the pipeline keeps running, bad records are excluded from Table1, and those records are automatically captured in the pipeline's event log as expectation violations --- available for review without any extra code.

Option A (custom quarantine logic) would work but requires writing and maintaining additional pipeline tables and routing logic. The whole point of SDP expectations is to handle this pattern declaratively, with far less code.

Option C (WHERE clauses in downstream queries) is a read-time filter, not a write-time guard. Invalid records would still land in Table1 and would simply be hidden from downstream views --- they're not preserved for review in any structured way. Option D (check constraint on Table1) would throw an exception on write and halt the pipeline, violating the 'pipeline continues to run' requirement.


Question 3

You have an Azure Databricks workspace that is enabled for Unity Catalog and contains a managed Delta table named Sales. Sales stores transaction data and contains the following columns:

* transactionjd (string)

* transaction date (date)

* amount (decimal)

You need to implement the following data quality requirements by using table-level data quality enforcement:

* amount must be greater than 0.

* transaction id must never be null.

* Invalid records must be rejected when data is written to the Sales table.

What should you do?



Answer : D

The correct answer is D --- a NOT NULL constraint on transaction_id and a CHECK constraint on amount.

Delta Lake table constraints are enforced at write time by the Delta engine itself. A NOT NULL constraint rejects any INSERT or UPDATE that would place a null in transaction_id. A CHECK constraint with amount > 0 rejects any row where amount is zero or negative. Combined, they implement exactly the stated quality rules: bad rows are rejected when data is written, not filtered away at read time.

Options A and C (SELECT with WHERE / views) are read-time constructs --- they don't prevent invalid data from entering the table. A clever pipeline bypass could write directly to the table and skip the view entirely. Option B (row-level security with WHERE conditions) is an access-control feature for restricting which rows users see, not for enforcing data quality on writes. Table constraints are the only mechanism that genuinely blocks bad data at the storage layer.


Question 4

You have an Azure Databricks workspace that is enabled for Unity Catalog.

You need to implement a daily batch data process that requires complex and highly customized Python transformations. The solution must minimize additional complexity.

What should you include in the solution?



Answer : A

A Databricks notebook provides the flexibility required to implement complex, highly customized Python and PySpark transformations. Scheduling that notebook as a Lakeflow Jobs task supplies native daily orchestration, monitoring, retries, and compute management without introducing another service. Azure Data Factory data flows are oriented toward visually designed transformations and would add external orchestration complexity for logic already implemented most naturally in Python. A continuous job is inappropriate because the workload runs once per day rather than continuously. Spark Declarative Pipelines is effective for declarative batch and streaming ETL, but it is less direct when the core requirement emphasizes highly customized procedural Python transformations. A notebook task therefore provides the necessary programming freedom while keeping scheduling and operation inside Azure Databricks.


Question 5

You have an Azure Databricks workspace named Workspace1 that contains a lakehouse and is enabled for Unity Catalog.

You have a connection to a Microsoft SQL Server database named DB1.

You need to expose the schemas and tables of DB1 to meet the following requirements:

* The schemas and tables can be queried in Databricks.

* The schemas and tables appear alongside other Unity Catalog objects.

* The data is NOT copied into Databricks-managed storage.

Solution: You create a foreign catalog in Catalog Explorer.

Does this meet the goal?



Answer : A

The correct answer is A --- Yes.

A foreign catalog created through Lakehouse Federation in Catalog Explorer is the correct solution for all three requirements. Here's why it works:

The schemas and tables of DB1 can be queried in Databricks --- Lakehouse Federation pushes the query down to the external SQL Server and returns results, so analysts write normal SQL in Databricks.

They appear alongside other Unity Catalog objects --- the foreign catalog sits in the same three-tier hierarchy as native catalogs, schemas, and tables, visible in Catalog Explorer alongside all other Unity Catalog assets.

The data is NOT copied into Databricks-managed storage --- foreign catalogs query data in place at the source; nothing is replicated or materialised in Databricks storage.

This is exactly the scenario Lakehouse Federation was built for.


Question 6

You have an Azure Databricks workspace that is enabled for Unity Catalog

You have an Apache Spark Structured Streaming job that writes data to a Delta table.

After the cluster restarts, the streaming job reprocesses previously ingested data

You need to prevent the streaming job from reprocessing the data after the cluster restarts.

What should you do?



Answer : B

The correct answer is B --- configure a checkpoint location.

A checkpoint is the Structured Streaming mechanism for fault tolerance. Databricks writes the committed offset (i.e., how far through the source stream the job has successfully read and processed) to a durable path in ADLS Gen2 or DBFS after each micro-batch. When the cluster restarts, the engine reads that offset and resumes from the next unprocessed record --- nothing is reprocessed, nothing is skipped.

Option A (increase trigger interval) affects how frequently micro-batches run but does nothing to record progress between runs. Option C (watermark) handles late-arriving events in event-time windows but doesn't control source offset tracking. Option D (enable CDF on the target table) tracks changes made to a Delta table for downstream consumers --- it has no bearing on the streaming job's own fault tolerance or offset management.

Checkpointing is a required configuration for any production streaming job. Without it, every cluster restart triggers a full replay from the source.


Question 7

You have an Azure Databricks workspace named Workspace1 that contains a takehouse and is enabled for Unity Catalog.

You have a connection to a Microsoft SQL Server database named DB1.

You need to expose the schemas and tables of DB1 to meet the following requirements:

* The schemas and tables can be queried in Databricks.

* The schemas and tables appear alongside other Unity Catalog objects.

* The data is NOT copied into Databricks-managed storage.

Solution: You create a new native catalog in Unity Catalog. Does this meet the goal?



Answer : B

The correct answer is B --- No.

A native catalog in Unity Catalog is a standard Databricks-managed catalog. Creating one provisions a metadata namespace inside Unity Catalog, but it has no connection whatsoever to DB1 or any external SQL Server database. There is no mechanism to point a native catalog at an external database --- it simply doesn't serve that purpose.

The requirement is to expose an external SQL Server's schemas and tables inside Unity Catalog without copying the data. That requires a foreign catalog, which is created through Lakehouse Federation using a registered connection to the external database. A foreign catalog acts as a read-only, virtual mirror of the external database --- queries run against the live external data in place.

Creating a native catalog and then trying to manually recreate DB1's structure inside it would involve copying all the data, violating the 'data is NOT copied' requirement.


Page:    1 / 14   
Total 91 questions