Comptia

DY0-001 Free Practice Questions

This is the free Comptia DY0-001 practice question bank — 50 of 82 total questions, each with a full explanation, free to read with no signup required. Updated 2026-10-03.

Every answer is verified against official Comptia documentation — see our methodology.

Question 1

Which of the following measures would a data scientist most likely use to calculate the similarity of two text strings?

A. Word cloud
B. Edit distance
C. String indexing
D. k-nearest neighbors
Show Answer
Correct Answer: B
Explanation:
Edit distance measures how many insertions, deletions, or substitutions are needed to transform one text string into another, so it can be used to assess their similarity.

Question 2

A data scientist is analyzing a data set with categorical features and would like to make those features more useful when building a model. Which of the following data transformation techniques should the data scientist use? (Choose two.)

A. Normalization
B. One-hot encoding
C. Linearization
D. Label encoding
E. Scaling
F. Pivoting
Show Answer
Correct Answer: B, D
Explanation:
One-hot encoding represents each category as a separate binary feature, while label encoding assigns a numeric value to each category. Both are common ways to transform categorical features for modeling.

Question 3

Given these business requirements: Needs to most efficiently move 3,000 boxes across a river Has one boat that holds eight boxes, travels at ten nautical miles per hour, and has a fuel economy of six nautical miles per gallon Has another boat that holds two boxes, travels at 50 nautical miles per hour, and has a fuel economy of 18 nautical miles per gallon The river is one nautical mile wide The data scientist only has access to 125 gallons of fuel Which of the following is the most likely optimization technique a data scientist would apply?

A. Constrained
B. Unconstrained
C. Non-iterative
D. Iterative
Show Answer
Correct Answer: A
Explanation:
This is a constrained optimization problem: the data scientist must choose how to use the boats while respecting limits such as the 125-gallon fuel supply, boat capacities, and river distance, and optimize the movement of the boxes.

Question 4

Which of the following is best solved with graph theory?

A. Optical character recognition
B. Traveling salesman
C. Fraud detection
D. One-armed bandit
Show Answer
Correct Answer: B
Explanation:
The traveling salesman problem is classically modeled as a weighted graph: cities are vertices, routes are edges, and the goal is to find a minimum-cost tour visiting each city once.

Question 5

A data scientist would like to model a complex phenomenon using a large data set composed of categorical, discrete, and continuous variables. After completing exploratory data analysis, the data scientist is reasonably certain that no linear relationship exists between the predictors and the target. Although the phenomenon is complex, the data scientist still wants to maintain the highest possible degree of interpretability in the final model. Which of the following algorithms best meets this objective?

A. Artificial neural network
B. Decision tree
C. Multiple linear regression
D. Random forest
Show Answer
Correct Answer: B
Explanation:
A decision tree can model nonlinear relationships and handle categorical, discrete, and continuous predictors, while its sequence of splits is relatively easy to interpret. Neural networks and random forests are less interpretable, and multiple linear regression assumes a linear relationship.

Question 6

A company created a very popular collectible card set. Collectors attempt to collect the entire set, but the availability of each card varies, with because some cards have higher production volumes than others. The set contains a total of 12 cards. The attributes of the cards are below: A data scientist is provided a historical record of cards purchased, which was acquired by a local collectors' association. The data scientist needs to design an initial model iteration to predict whether or not the animal on the card lives in the sea or on land given the provided attributes. Which of the following is the best way to accomplish this task?

A. ARIMA
B. Linear regression
C. Association rules
D. Decision trees
Show Answer
Correct Answer: D
Explanation:
A decision tree is a supervised classification method that can use the card attributes to predict the categorical label—whether the animal lives in the sea or on land. ARIMA is for time-series forecasting, linear regression predicts continuous values, and association rules find co-occurrence patterns.

Question 7

A data scientist receives an update on a business case about a machine that has thousands of error codes. The data scientist creates the following summary statistics profile while reviewing the logs for each machine: Which of the following is the most likely concern with respect to data design for model ingestion?

A. Sparse matrix
B. Granularity misalignment
C. Insufficient features
D. Multivariate outliers
Show Answer
Correct Answer: A
Explanation:
Thousands of possible error codes, with only a few occurring for each machine, would produce a high-dimensional feature matrix containing mostly zeros. This is a sparse matrix concern.

Question 8

Which of the following image data augmentation techniques allows a data scientist to increase the size of a data set?

A. Clipping
B. Cropping
C. Masking
D. Scaling
Show Answer
Correct Answer: B
Explanation:
Cropping can create multiple distinct image samples from different regions of one original image, increasing the dataset’s size.

Question 9

Which of the following types of layers is used to downsample feature detection when using a convolutional neural network?

A. Pooling
B. Input
C. Output
D. Hidden
Show Answer
Correct Answer: A
Explanation:
Pooling layers downsample feature maps by reducing their spatial dimensions, while retaining salient features.

Question 10

Which of the following distributions would be best to use for hypothesis testing on a data set with 20 observations?

A. Power law
B. Normal
C. Uniform
D. Student's t-
Show Answer
Correct Answer: D
Explanation:
For a small sample (20 observations), when the population variance is unknown and estimated from the sample, the Student’s t-distribution is generally appropriate for hypothesis tests about a mean, assuming the data are approximately normal.

$19

Get all 82 questions with detailed answers and explanations

  • Instant download HTML + PDF delivered the moment payment clears.
  • Secure Stripe checkout we never see or store your card details.
  • 7-day refund if files are defective see our refund policy.