Amazon AWS Certified Data Engineer - Associate DEA-C01 Exam Questions

Page: 1 / 14
Total 302 questions
Question 1

A company stores sensitive transaction data in an Amazon S3 bucket. A data engineer must implement controls to prevent accidental deletions.



Answer : A

''To protect data from accidental deletion, enable S3 Versioning and MFA Delete, which requires MFA for object deletions and prevents unintentional loss.''

-- Ace the AWS Certified Data Engineer - Associate Certification - version 2 - apple.pdf

This is the AWS best practice for securing critical S3 data.


Question 2

A company stores CSV files in an Amazon S3 bucket. A data engineer needs to process the data in the CSV files and store the processed data in a new S3 bucket.

The process needs to rename a column, remove specific columns, ignore the second row of each file, create a new column based on the values of the first row of the data, and filter the results by a numeric value of a column.

Which solution will meet these requirements with the LEAST development effort?



Answer : D

The requirement involves transforming CSV files by renaming columns, removing rows, and other operations with minimal development effort. AWS Glue DataBrew is the best solution here because it allows you to visually create transformation recipes without writing extensive code.

Option D: Use AWS Glue DataBrew recipes to read and transform the CSV files.DataBrew provides a visual interface where you can build transformation steps (e.g., renaming columns, filtering rows, creating new columns, etc.) as a 'recipe' that can be applied to datasets, making it easy to handle complex transformations on CSV files with minimal coding.

Other options (A, B, C) involve more manual development and configuration effort (e.g., writing Python jobs or creating custom workflows in Glue) compared to the low-code/no-code approach of DataBrew.


AWS Glue DataBrew Documentation

Question 3

A company stores server logs in an Amazon 53 bucket. The company needs to keep the logs for 1 year. The logs are not required after 1 year.

A data engineer needs a solution to automatically delete logs that are older than 1 year.

Which solution will meet these requirements with the LEAST operational overhead?



Answer : B

Problem Analysis:

The company uses AWS Glue for ETL pipelines and requires automatic data quality checks during pipeline execution.

The solution must integrate with existing AWS Glue pipelines and evaluate data quality rules based on predefined thresholds.

Key Considerations:

Ensure minimal implementation effort by leveraging built-in AWS Glue features.

Use a standardized approach for defining and evaluating data quality rules.

Avoid custom libraries or external frameworks unless absolutely necessary.

Solution Analysis:

Option A: SQL Transform

Adding SQL transforms to define and evaluate data quality rules is possible but requires writing complex queries for each rule.

Increases operational overhead and deviates from Glue's declarative approach.

Option B: Evaluate Data Quality Transform with DQDL

AWS Glue provides a built-in Evaluate Data Quality transform.

Allows defining rules in Data Quality Definition Language (DQDL), a concise and declarative way to define quality checks.

Fully integrated with Glue Studio, making it the least effort solution.

Option C: Custom Transform with PyDeequ

PyDeequ is a powerful library for data quality checks but requires custom code and integration.

Increases implementation effort compared to Glue's native capabilities.

Option D: Custom Transform with Great Expectations

Great Expectations is another powerful library for data quality but adds complexity and external dependencies.

Final Recommendation:

Use Evaluate Data Quality transform in AWS Glue.

Define rules in DQDL for checking thresholds, null values, or other quality criteria.

This approach minimizes development effort and ensures seamless integration with AWS Glue.

AWS Glue Data Quality Overview

DQDL Syntax and Examples

Glue Studio Transformations


Question 4

A company is developing a product recommendation system that uses Amazon OpenSearch Service. The system needs to perform k-nearest neighbors (k-NN) vector searches on 10 million product embeddings with 768-dimensional vectors. The system must maintain high recall accuracy and support incremental updates without reindexing as new products are added each day. The system must also accommodate complex filtering based on product categories and inventory status.

Which vector index type will meet these requirements?



Answer : B

The correct answer is B because the scenario requires scalable approximate vector search, high recall, incremental updates, and complex filtering. Amazon OpenSearch Service supports k-NN vector search for recommendation use cases, and OpenSearch supports k-NN vector fields for high-dimensional vector search. The Lucene HNSW option is strongest here because OpenSearch documentation specifically states that Lucene supports k-NN searches using HNSW graphs and supports Lucene filters for k-NN searches. That directly matches the need for category and inventory filtering. Exact k-NN with Painless script scoring is accurate but too slow for 10 million vectors. IVF can be efficient but is less ideal for frequent incremental updates and complex filtering. Binary quantization reduces memory but sacrifices accuracy.


Question 5

A data engineer develops an AWS Glue Apache Spark ETL job to perform transformations on a dataset. When the data engineer runs the job, the job returns an error that reads, ''No space left on device.''

The data engineer needs to identify the source of the error and provide a solution.

Which combinations of steps will meet this requirement MOST cost-effectively? (Select TWO.)



Answer : B, D

Options B and D are correct. AWS Prescriptive Guidance states that shuffle is a major cause of Spark performance issues and can exhaust local disk space on executors, leading to ''No space left on device'' failures. AWS specifically recommends assessing shuffle performance in CloudWatch metrics and in the Spark UI, and for data skew, it says to examine metrics in the Spark UI, including task duration and spill behavior across executors. That makes B the correct low-cost way to identify the source of the problem.

For remediation, AWS Glue documentation states that Spark can throw ''No space left on device'' when there is insufficient local disk on the executor, and that you can use Amazon S3 to store Spark shuffle data by enabling the AWS Glue Spark shuffle plugin. The S3 shuffle approach is specifically presented as a way to run shuffle-intensive jobs more reliably when they are bound by local disk capacity. In skew scenarios, using a skew-mitigation technique such as salting is also a standard way to distribute hot keys more evenly. That makes D the most cost-effective corrective step.

Options A and C increase cost by adding capacity before diagnosing or directly fixing the shuffle bottleneck. Option E is less precise than the Spark UI and Glue metrics for identifying skew.


Question 6

A company uses AWS Glue ETL pipelines to process data. The company uses Amazon Athena to analyze data in an Amazon S3 bucket.

To better understand shipping timelines, the company decides to collect and store shipping dates and delivery dates in addition to order data. The company adds a data quality check to ensure that the shipping date is later than the order date and that the delivery date is later than the shipping date. Orders that fail the quality check must be stored in a second Amazon S3 bucket.

Which solution will meet these requirements in the MOST cost-effective way?



Answer : D


Question 7

A data engineer needs to debug an AWS Glue job that reads from Amazon S3 and writes to Amazon Redshift. The data engineer enabled the bookmark feature for the AWS Glue job. The data engineer has set the maximum concurrency for the AWS Glue job to 1.

The AWS Glue job is successfully writing the output to Amazon Redshift. However, the Amazon S3 files that were loaded during previous runs of the AWS Glue job are being reprocessed by subsequent runs.

What is the likely reason the AWS Glue job is reprocessing the files?



Answer : A

The issue described is that the AWS Glue job is reprocessing files from previous runs despite the bookmark feature being enabled. Bookmarks in AWS Glue allow jobs to keep track of which files or data have already been processed to avoid reprocessing. The most likely reason for reprocessing the files is missing S3 permissions, specifically s3

s3

is a permission required by AWS Glue when bookmarks are enabled to ensure Glue can retrieve metadata from the files in S3, which is necessary for the bookmark mechanism to function correctly. Without this permission, Glue cannot track which files have been processed, resulting in reprocessing during subsequent runs.

Concurrency settings (Option B) and the version of AWS Glue (Option C) do not affect the bookmark behavior. Similarly, the lack of a commit statement (Option D) is not applicable in this context, as Glue handles commits internally when interacting with Redshift and S3.

Thus, the root cause is likely related to insufficient permissions on the S3 bucket, specifically s3

, which is required for bookmarks to work as expected.


AWS Glue Job Bookmarks Documentation

AWS Glue Permissions for Bookmarks

Page:    1 / 14   
Total 302 questions