Associate Data Practitioner Free Practice Questions — Page 4
Question 25
Your organization uses scheduled queries to perform transformations on data stored in BigQuery. You discover that one of your scheduled queries has failed. You need to troubleshoot the issue as quickly as possible. What should you do?
A. Navigate to the Logs Explorer page in Cloud Logging. Use filters to find the failed job, and analyze the error details.
B. Set up a log sink using the gcloud CLI to export BigQuery audit logs to BigQuery. Query those logs to identify the error associated with the failed job ID.
C. Request access from your admin to the BigQuery information_schema. Query the jobs view with the failed job ID, and analyze error details.
D. Navigate to the Scheduled queries page in the Google Cloud console. Select the failed job, and analyze the error details.
Show Answer
Correct Answer: D
Explanation: Open the Scheduled queries page in the Google Cloud console and select the failed scheduled query to view its execution details and error message. This is the most direct way to troubleshoot the failure quickly.
Question 26
Your organization has decided to migrate their existing enterprise data warehouse to BigQuery. The existing data pipeline tools already support connectors to BigQuery. You need to identify a data migration approach that optimizes migration speed. What should you do?
A. Create a temporary file system to facilitate data transfer from the existing environment to Cloud Storage. Use Storage Transfer Service to migrate the data into BigQuery.
B. Use the Cloud Data Fusion web interface to build data pipelines. Create a directed acyclic graph (DAG) that facilitates pipeline orchestration.
C. Use the existing data pipeline tool’s BigQuery connector to reconfigure the data mapping.
D. Use the BigQuery Data Transfer Service to recreate the data pipeline and migrate the data into BigQuery.
Show Answer
Correct Answer: C
Explanation: Use the existing pipeline tool’s BigQuery connector and reconfigure the data mapping. This reuses tools and pipelines already in place, avoiding the time and complexity of introducing a new transfer or orchestration service.
Question 27
Your organization has several datasets in BigQuery. The datasets need to be shared with your external partners so that they can run SQL queries without needing to copy the data to their own projects. You have organized each partner’s data in its own BigQuery dataset. Each partner should be able to access only their data. You want to share the data while following Google-recommended practices. What should you do?
A. Use Analytics Hub to create a listing on a private data exchange for each partner dataset. Allow each partner to subscribe to their respective listings.
B. Create a Dataflow job that reads from each BigQuery dataset and pushes the data into a dedicated Pub/Sub topic for each partner. Grant each partner the pubsub. subscriber IAM role.
C. Export the BigQuery data to a Cloud Storage bucket. Grant the partners the storage.objectUser IAM role on the bucket.
D. Grant the partners the bigquery.user IAM role on the BigQuery project.
Show Answer
Correct Answer: A
Explanation: Create a separate Analytics Hub listing for each partner’s dataset on a private data exchange, then let each partner subscribe to only their listing. This supports sharing BigQuery data for in-place querying while limiting each partner to the data intended for them.
Question 28
Your team needs to analyze large datasets stored in BigQuery to identify trends in user behavior. The analysis will involve complex statistical calculations, Python packages, and visualizations. You need to recommend a managed collaborative environment to develop and share the analysis. What should you recommend?
A. Create a Colab Enterprise notebook and connect the notebook to BigQuery. Share the notebook with your team. Analyze the data and generate visualizations in Colab Enterprise.
B. Create a statistical model by using BigQuery ML. Share the query with your team. Analyze the data and generate visualizations in Looker Studio.
C. Create a Looker Studio dashboard and connect the dashboard to BigQuery. Share the dashboard with your team. Analyze the data and generate visualizations in Looker Studio.
D. Connect Google Sheets to BigQuery by using Connected Sheets. Share the Google Sheet with your team. Analyze the data and generate visualizations in Gooqle Sheets.
Show Answer
Correct Answer: A
Explanation: Colab Enterprise provides a managed, collaborative notebook environment with Python support. Connecting it to BigQuery lets the team analyze large datasets using statistical libraries and create visualizations, then share the notebook with collaborators.
Question 29
You used BigQuery ML to build a customer purchase propensity model six months ago. You want to compare the current serving data with the historical serving data to determine whether you need to retrain the model. What should you do?
A. Compare the two different models.
B. Evaluate the data skewness.
C. Evaluate data drift.
D. Compare the confusion matrix.
Show Answer
Correct Answer: C
Explanation: Evaluate data drift by comparing the distribution of current serving data with historical serving data. A significant change can indicate that the model may need retraining. Comparing models or confusion matrices does not directly measure this change in input data.
Question 30
You are a data analyst working with sensitive customer data in BigQuery. You need to ensure that only authorized personnel within your organization can query this data, while following the principle of least privilege. What should you do?
A. Enable access control by using IAM roles.
B. Encrypt the data by using customer-managed encryption keys (CMEK).
C. Update dataset privileges by using the SQL GRANT statement.
D. Export the data to Cloud Storage, and use signed URLs to authorize access.
Show Answer
Correct Answer: A
Explanation: Use IAM to grant only the necessary BigQuery permissions to authorized personnel—for example, dataset-level read access and the required query-job permission. This enforces least privilege. CMEK protects encryption keys, not who can query the data; exporting it or using signed URLs is unnecessary.
Question 31
Your retail company collects customer data from various sources:
Online transactions: Stored in a MySQL database
Customer feedback: Stored as text files on a company server
Social media activity: Streamed in real-time from social media platforms
You are designing a data pipeline to extract this data. Which Google Cloud storage system(s) should you select for further analysis and ML model training?
A. 1. Online transactions: Cloud Storage 2. Customer feedback: Cloud Storage 3. Social media activity: Cloud Storage
B. 1. Online transactions: BigQuery 2. Customer feedback: Cloud Storage 3. Social media activity: BigQuery
C. 1. Online transactions: Bigtable 2. Customer feedback: Cloud Storage 3. Social media activity: CloudSQL for MySQL
D. 1. Online transactions: Cloud SQL for MySQL 2. Customer feedback: BigQuery 3. Social media activity: Cloud Storage
Show Answer
Correct Answer: B
Explanation: BigQuery is suitable for analyzing structured online transaction data and ingesting streamed social media data. Cloud Storage is appropriate for the customer feedback text files and supports storing unstructured data for later analysis and ML workflows.
Question 32
You work for a financial services company that handles highly sensitive data. Due to regulatory requirements, your company is required to have complete and manual control of data encryption. Which type of keys should you recommend to use for data storage?
A. Use customer-supplied encryption keys (CSEK).
B. Use a dedicated third-party key management system (KMS) chosen by the company.
C. Use Google-managed encryption keys (GMEK).
D. Use customer-managed encryption keys (CMEK).
Show Answer
Correct Answer: A
Explanation: Customer-supplied encryption keys (CSEK) provide the most direct manual control: the customer supplies the key for encryption rather than relying on Google-managed keys or keys managed through Cloud KMS. CMEK offers customer control over keys through KMS, but not the same direct, manually supplied key model.
Question 33
You are constructing a data pipeline to process sensitive customer data stored in a Cloud Storage bucket. You need to ensure that this data remains accessible, even in the event of a single-zone outage. What should you do?
A. Set up a Cloud CDN in front of the bucket.
B. Enable Object Versioning on the bucket.
C. Store the data in a multi-region bucket.
D. Store the data in Nearline storage.
Show Answer
Correct Answer: C
Explanation: A multi-region Cloud Storage bucket stores data redundantly across multiple regions, helping keep it accessible if a zone becomes unavailable. CDN, Object Versioning, and Nearline storage do not provide this availability benefit.
Question 34
You work for an ecommerce company that has a BigQuery dataset that contains customer purchase history, demographics, and website interactions. You need to build a machine learning (ML) model to predict which customers are most likely to make a purchase in the next month. You have limited engineering resources and need to minimize the ML expertise required for the solution. What should you do?
A. Use BigQuery ML to create a logistic regression model for purchase prediction.
B. Use Vertex AI Workbench to develop a custom model for purchase prediction.
C. Use Colab Enterprise to develop a custom model for purchase prediction.
D. Export the data to Cloud Storage, and use AutoML Tables to build a classification model for purchase prediction.
Show Answer
Correct Answer: A
Explanation: Use BigQuery ML to train a logistic regression classifier directly on the data in BigQuery. It avoids exporting data or building a custom modeling workflow, minimizing engineering effort and the ML expertise required. Logistic regression can predict whether a customer will make a purchase in the next month as a binary outcome.
$19
Get all 90 questions with detailed answers and explanations
Instant download HTML + PDF delivered the moment payment clears.
Secure Stripe checkout we never see or store your card details.
7-day refund if files are defective see our refund policy.