Under perfect conditions, E. coli bacteria would cover the entire earth in a matter of days. Which of the following types of models is the best for explaining this type of growth?
A. Linear
B. Logarithmic
C. Polynomial
D. Exponential
Show Answer
Correct Answer: D
Explanation: Under ideal conditions, E. coli reproduces by repeated doubling, so its population grows exponentially rather than at a constant rate.
Question 22
The following graphic shows the results of an unsupervised, machine-learning clustering model:
k is the number of clusters, and n is the processing time required to run the model. Which of the following is the best value of k to optimize both accuracy and processing requirements?
A. 2
B. 10
C. 15
D. 20
Show Answer
Correct Answer: B
Explanation: k = 10 is the best trade-off: beyond about 10 clusters, the graph’s improvement in clustering quality is small relative to the additional processing time, so the curve reaches an elbow near this value.
Question 23
A data scientist is using the following confusion matrix to assess model performance:
The model is predicting whether a delivery truck will be able to make 200 scheduled delivery stops. Every time the model is correct, the company saves an hour in planning and scheduling of maintenance work. Every time the model is wrong, the company loses four hours of delivery time for the truck. Which of the following is the net model impact for the company?
A. 25 hours lost
B. 25 hours saved
C. 165 hours lost
D. 165 hours saved
Show Answer
Correct Answer: B
Explanation: There are 165 correct predictions (80 + 85), saving 165 hours. There are 35 incorrect predictions (20 + 15), costing 140 hours. Net impact: 165 − 140 = 25 hours saved.
Question 24
A team is building a spam detection system. The team wants a probability-based identification method without complex, in-depth training from the historical data set. Which of the following methods would best serve this purpose?
A. Logistic regression
B. Random forest
C. Naive Bayes
D. Linear regression
Show Answer
Correct Answer: C
Explanation: Naive Bayes is a probabilistic classification method that is relatively simple and efficient to train, making it well suited to spam detection when extensive or complex training is not desired.
Question 25
A model's results show increasing explanatory value as additional independent variables are added to the model. Which of the following is the most appropriate statistic?
A. Adjusted R2
B. p value
C. x2
D. R2
Show Answer
Correct Answer: A
Explanation: Adjusted R² is most appropriate when comparing models with different numbers of independent variables. Unlike ordinary R², it adjusts for the number of predictors, so adding variables improves the statistic only when they contribute enough explanatory value.
Question 26
A data scientist is standardizing a large data set that contains website addresses. A specific string inside some of the web addresses needs to be extracted. Which of the following is the best method for extracting the desired string from the text data?
A. Regular expressions
B. Named-entity recognition
C. Large language model
D. Find and replace
Show Answer
Correct Answer: A
Explanation: Regular expressions are well suited to extracting strings that follow a specific pattern in website addresses. Named-entity recognition identifies entities such as people or organizations, while an LLM or find-and-replace is less direct for this structured extraction task.
Question 27
A data scientist is building a proof of concept for a commercialized machine-learning model. Which of the following is the best starting point?
A. Literature review
B. Model performance evaluation
C. Hyperparameter tuning
D. Model selection
Show Answer
Correct Answer: A
Explanation: A literature review is the best starting point: it helps identify relevant prior work, suitable methods, and established practices before selecting or tuning a model or evaluating its performance.
Question 28
Which of the following best describes the minimization of the residual term in a LASSO linear regression?
A. |e|
B. e
C. 0
D. e2
Show Answer
Correct Answer: D
Explanation: LASSO minimizes the sum of squared residuals (e²) together with an L1 penalty on the coefficients. The residual term is therefore represented by e².
Question 29
Which of the following layer sets includes the minimum three layers required to constitute an artificial neural network?
A. An input layer, a pooling layer, and an output layer
B. An input layer, a convolutional layer, and a hidden layer
C. An input layer, a hidden layer, and an output layer
D. An input layer, a dropout layer, and a hidden layer
Show Answer
Correct Answer: C
Explanation: The minimum three-layer neural network structure consists of an input layer, at least one hidden layer for processing, and an output layer for producing results.
Question 30
A data scientist uses a large data set to build multiple linear regression models to predict the likely market value of a real estate property. The selected new model has an RMSE of 995 on the holdout set and an adjusted R2 of .75. The benchmark model has an RMSE of 1,000 on the holdout set. Which of the following is the best business statement regarding the new model?
A. The model should be deployed because it has a lower RMSE.
B. The model's adjusted R2 is exceptionally strong for such a complex relationship.
C. The model fails to improve meaningfully on the benchmark model.
D. The model's adjusted R2 is too low for the real estate industry.
Show Answer
Correct Answer: C
Explanation: The holdout RMSE decreases only from 1,000 to 995, a 0.5% improvement. Without evidence that this small gain is practically meaningful, it is not enough to justify deploying the new model. Adjusted R² alone does not establish business value.
$19
Get all 82 questions with detailed answers and explanations
Instant download HTML + PDF delivered the moment payment clears.
Secure Stripe checkout we never see or store your card details.
7-day refund if files are defective see our refund policy.