Associate Data Practitioner Free Practice Questions — Page 5
Question 35
You are designing a pipeline to process data files that arrive in Cloud Storage by 3:00 am each day. Data processing is performed in stages, where the output of one stage becomes the input of the next. Each stage takes a long time to run. Occasionally a stage fails, and you have to address
the problem. You need to ensure that the final output is generated as quickly as possible. What should you do?
A. Design a Spark program that runs under Dataproc. Code the program to wait for user input when an error is detected. Rerun the last action after correcting any stage output data errors.
B. Design the pipeline as a set of PTransforms in Dataflow. Restart the pipeline after correcting any stage output data errors.
C. Design the workflow as a Cloud Workflow instance. Code the workflow to jump to a given stage based on an input parameter. Rerun the workflow after correcting any stage output data errors.
D. Design the processing as a directed acyclic graph (DAG) in Cloud Composer. Clear the state of the failed task after correcting any stage output data errors.
Show Answer
Correct Answer: D
Explanation: Cloud Composer lets you model the stages as a DAG and rerun failed work without restarting the entire pipeline. After fixing the issue, clear the failed task’s state (and rerun dependent downstream tasks if needed), so completed long-running stages can be reused and the final output is produced sooner.
Question 36
Another team in your organization is requesting access to a BigQuery dataset. You need to share the dataset with the team while minimizing the risk of unauthorized copying of data. You also want to create a reusable framework in case you need to share this data with other teams in the future. What should you do?
A. Create authorized views in the team’s Google Cloud project that is only accessible by the team.
B. Create a private exchange using Analytics Hub with data egress restriction, and grant access to the team members.
C. Enable domain restricted sharing on the project. Grant the team members the BigQuery Data Viewer IAM role on the dataset.
D. Export the dataset to a Cloud Storage bucket in the team’s Google Cloud project that is only accessible by the team.
Show Answer
Correct Answer: B
Explanation: Create a private Analytics Hub exchange with data egress restrictions and grant the team access. Analytics Hub provides a reusable framework for sharing BigQuery data, and egress restrictions help limit subscribers’ ability to copy or export it. Authorized views can control what data is visible, but they do not provide the same data-egress controls.
Question 37
You need to create a new data pipeline. You want a serverless solution that meets the following requirements:
• Data is streamed from Pub/Sub and is processed in real-time.
• Data is transformed before being stored.
• Data is stored in a location that will allow it to be analyzed with SQL using Looker.
Which Google Cloud services should you recommend for the pipeline?
A. 1. Dataproc Serverless 2. Bigtable
B. 1. Cloud Composer 2. Cloud SQL for MySQL
C. 1. BigQuery 2. Analytics Hub
D. 1. Dataflow 2. BigQuery
Show Answer
Correct Answer: D
Explanation: Dataflow can read streaming messages from Pub/Sub and transform them in real time. BigQuery provides SQL analytics and integrates with Looker, making it a suitable destination for the processed data.
Question 38
Your team wants to create a monthly report to analyze inventory data that is updated daily. You need to aggregate the inventory counts by using only the most recent month of data, and save the results to be used in a Looker Studio dashboard. What should you do?
A. Create a materialized view in BigQuery that uses the SUM( ) function and the DATE_SUB( ) function.
B. Create a saved query in the BigQuery console that uses the SUM( ) function and the DATE_SUB( ) function. Re-run the saved query every month, and save the results to a BigQuery table.
C. Create a BigQuery table that uses the SUM( ) function and the _PARTITIONDATE filter.
D. Create a BigQuery table that uses the SUM( ) function and the DATE_DIFF( ) function.
Show Answer
Correct Answer: B
Explanation: Use SUM() to aggregate the counts and DATE_SUB() to limit the query to the most recent month, then re-run it monthly and save the results in a BigQuery table for Looker Studio. A rolling-date filter would require a function such as CURRENT_DATE(), which is not supported in a BigQuery materialized view definition.
Sources:
https://cloud.google.com/blog/products/gcp/how-to-build-a-bi-dashboard-using-google-data-studio-and-bigquery
Question 39
You have a BigQuery dataset containing sales data. This data is actively queried for the first 6 months. After that, the data is not queried but needs to be retained for 3 years for compliance reasons. You need to implement a data management strategy that meets access and compliance requirements, while keeping cost and administrative overhead to a minimum. What should you do?
A. Use BigQuery long-term storage for the entire dataset. Set up a Cloud Run function to delete the data from BigQuery after 3 years.
B. Partition a BigQuery table by month. After 6 months, export the data to Coldline storage. Implement a lifecycle policy to delete the data from Cloud Storage after 3 years.
C. Set up a scheduled query to export the data to Cloud Storage after 6 months. Write a stored procedure to delete the data from BigQuery after 3 years.
D. Store all data in a single BigQuery table without partitioning or lifecycle policies.
Show Answer
Correct Answer: B
Explanation: Partitioning by month lets you manage and export older data in manageable units. After six months, moving it to Coldline reduces storage costs, and a Cloud Storage lifecycle rule can delete it when the compliance retention period ends. This keeps recent data readily queryable in BigQuery while avoiding the cost of retaining inactive data there. Configure the lifecycle age so deletion occurs three years after the data’s original retention start date.
Question 40
You have created a LookML model and dashboard that shows daily sales metrics for five regional managers to use. You want to ensure that the regional managers can only see sales metrics specific to their region. You need an easy-to-implement solution. What should you do?
A. Create a sales_region user attribute, and assign each manager’s region as the value of their user attribute. Add an access_filter Explore filter on the region_name dimension by using the sales_region user attribute.
B. Create five different Explores with the sql_always_filter Explore filter applied on the region_name dimension. Set each region_name value to the corresponding region for each manager.
C. Create separate Looker dashboards for each regional manager. Set the default dashboard filter to the corresponding region for each manager.
D. Create separate Looker instances for each regional manager. Copy the LookML model and dashboard to each instance. Provision viewer access to the corresponding manager.
Show Answer
Correct Answer: A
Explanation: Create a user attribute containing each manager’s assigned region, then apply it to the region dimension with an Explore `access_filter`. This enforces row-level access in queries while keeping one model and dashboard; dashboard filters alone do not securely restrict data.
Question 41
You need to design a data pipeline that ingests data from CSV, Avro, and Parquet files into Cloud Storage. The data includes raw user input. You need to remove all malicious SQL injections before storing the data in BigQuery. Which data manipulation methodology should you choose?
A. EL
B. ELT
C. ETL
D. ETLT
Show Answer
Correct Answer: C
Explanation: Choose ETL: extract the CSV, Avro, and Parquet data, transform it to detect and handle malicious input, then load the cleaned data into BigQuery. Since the transformation must occur before the data is stored in BigQuery, this is ETL rather than ELT.
Question 42
Your retail organization stores sensitive application usage data in Cloud Storage. You need to encrypt the data without the operational overhead of managing encryption keys. What should you do?
A. Use Google-managed encryption keys (GMEK).
B. Use customer-managed encryption keys (CMEK).
C. Use customer-supplied encryption keys (CSEK).
D. Use customer-supplied encryption keys (CSEK) for the sensitive data and customer-managed encryption keys (CMEK) for the less sensitive data.
Show Answer
Correct Answer: A
Explanation: Use Google-managed encryption keys (GMEK). Google handles key management and encryption automatically, avoiding the operational overhead of managing keys yourself.
Question 43
You work for a financial organization that stores transaction data in BigQuery. Your organization has a regulatory requirement to retain data for a minimum of seven years for auditing purposes. You need to ensure that the data is retained for seven years using an efficient and cost-optimized approach. What should you do?
A. Create a partition by transaction date, and set the partition expiration policy to seven years.
B. Set the table-level retention policy in BigQuery to seven years.
C. Set the dataset-level retention policy in BigQuery to seven years.
D. Export the BigQuery tables to Cloud Storage daily, and enforce a lifecycle management policy that has a seven-year retention rule.
Show Answer
Correct Answer: A
Explanation: Partition the table by transaction date and set a seven-year partition expiration. Each partition then expires only after its data has reached seven years, rather than expiring the entire table at once. This is an efficient way to manage BigQuery data over time. If the requirement also means users must be prevented from deleting data early, partition expiration alone does not provide that protection; an enforced retention control would be needed.
Question 44
You need to create a weekly aggregated sales report based on a large volume of data. You want to use Python to design an efficient process for generating this report. What should you do?
A. Create a Cloud Run function that uses NumPy. Use Cloud Scheduler to schedule the function to run once a week.
B. Create a Colab Enterprise notebook and use the bigframes.pandas library. Schedule the notebook to execute once a week.
C. Create a Cloud Data Fusion and Wrangler flow. Schedule the flow to run once a week.
D. Create a Dataflow directed acyclic graph (DAG) coded in Python. Use Cloud Scheduler to schedule the code to run once a week.
Show Answer
Correct Answer: D
Explanation: Dataflow can run a Python batch pipeline that aggregates large volumes of sales data in parallel. Scheduling the pipeline weekly provides a scalable, repeatable reporting process.
Sources:
https://cloud.google.com/blog/products/gcp/google-announces-cloud-dataflow-with-python-support
https://cloud.google.com/dataflow/docs/quickstarts/create-pipeline-python
$19
Get all 90 questions with detailed answers and explanations
Instant download HTML + PDF delivered the moment payment clears.
Secure Stripe checkout we never see or store your card details.
7-day refund if files are defective see our refund policy.