A data scientist is working with a data set that has ten predictors and wants to use only the predictors that most influence the results. Which of the following models would be the best for the data scientist to use?
A. OLS
B. Ridge
C. Weighted least squares
D. LASSO
Show Answer
Correct Answer: D
Explanation: LASSO performs regularization and feature selection by shrinking some predictor coefficients exactly to zero, leaving the predictors that contribute most to the model.
Question 32
A data scientist wants to predict a person's travel destination. The options are:
Branson, Missouri, United States
Mount Kilimanjaro, Tanzania
Disneyland Paris, Paris, France
Sydney Opera House, Sydney, Australia
Which of the following models would best fit this use case?
A. Linear discriminant analysis
B. k-means modeling
C. Latent semantic analysis
D. Principal component analysis
Show Answer
Correct Answer: A
Explanation: The destinations are a fixed set of labeled categories, so predicting one of them is a supervised multiclass classification task. Linear discriminant analysis is a classification method; k-means, LSA, and PCA are primarily unsupervised clustering or dimensionality-reduction techniques.
Question 33
A data scientist is creating a responsive model that will update a product's daily pricing based on the previous day's sales volume. Which of the following resource constraints is the data scientist's greatest concern?
A. Deployment time
B. Training time
C. Development time
D. Data collection time
Show Answer
Correct Answer: B
Explanation: Because the model must be updated from new sales data each day, training needs to finish quickly enough for the next day’s pricing decision. Training time is therefore the greatest concern.
Question 34
A data scientist is building a model to predict customer credit scores based on information collected from reporting agencies. The model needs to automatically adjust its parameters to adapt to recent changes in the information collected. Which of the following is the best model to use?
A. Decision tree
B. Random forest
C. Linear discrimination analysis
D. XGBoost
Show Answer
Correct Answer: D
Explanation: XGBoost can be continued or updated with new data, making it the best of these options for adapting as reporting information changes. It is not strictly an online-learning algorithm, but it can be more readily refreshed than the other listed models.
Question 35
A data analyst wants to generate the most data using tables from a database. Which of the following is the best way to accomplish this objective?
A. INNER JOIN
B. LEFT OUTER JOIN
C. RIGHT OUTER JOIN
D. FULL OUTER JOIN
Show Answer
Correct Answer: D
Explanation: A FULL OUTER JOIN includes matching rows and also unmatched rows from both tables, so it generally returns the most inclusive combined dataset.
Question 36
A data scientist has built an image recognition model that distinguishes cars from trucks. The data scientist now wants to measure the rate at which the model correctly identifies a car as a car versus when it misidentifies a truck as a car. Which of the following would best convey this information?
A. Confusion matrix
B. AUC/ROC curve
C. Box plot
D. Correlation plot
Show Answer
Correct Answer: A
Explanation: A confusion matrix shows the counts or proportions of actual and predicted classes. It reveals both the rate of correctly identifying cars as cars and the rate of misclassifying trucks as cars.
Question 37
A data scientist is working with a data set that covers a two-year period for a large number of machines. The data set contains:
Machine system ID numbers
Sensor measurement values
Daily time stamps for each machine
The data scientist needs to plot the total measurements from all the machines over the entire time period. Which of the following is the best way to present this data?
A. Scatter plot
B. Line plot
C. Histogram
D. Box-and-whisker plot
Show Answer
Correct Answer: B
Explanation: A line plot is best for showing how the total measurement changes over time. Sum the measurements across machines for each daily timestamp, then plot those totals in chronological order.
Question 38
A data scientist is developing a model to predict the outcome of a vote for a national mascot. The choice is between tigers and lions. The full data set represents feedback from individuals representing 17 professions and 12 different locations. The following rank aggregation represents 80% of the data set:
Which of the following is the most likely concern about the model's ability to predict the outcome of the vote?
A. Interpolated data
B. Extrapolated data
C. In-sample data
D. Out-of-sample data
Show Answer
Correct Answer: B
Explanation: The 80% subset represents only a small portion of the professions and locations in the full dataset. Predicting vote outcomes for professions or locations not represented in that subset requires extending beyond the observed data, which is extrapolation.
Question 39
A data scientist is presenting the recommendations from a monthslong modeling and experiment process to the company's Chief Executive Officer. Which of the following is the best set of artifacts to include in the presentation?
A. Methods, data overview, results, recommendations, and charts
B. Results, recommendations, justifications, and clear charts
C. Recommendation, charts, justifications, code reviews, and results
D. Methodology, code snippets, findings, data tables, and p values
Show Answer
Correct Answer: B
Explanation: For a CEO, the presentation should emphasize the results, actionable recommendations, the rationale behind them, and clear visualizations. These communicate the conclusions without burdening the audience with technical details such as code or extensive methodology.
Question 40
Which of the following methods should a data scientist use just before switching to a potential replacement model?
A. A/B testing
B. Performance monitoring
C. CI/CD
D. Containerization
Show Answer
Correct Answer: A
Explanation: A/B testing lets the data scientist compare the potential replacement against the current model on real traffic before switching over.
$19
Get all 82 questions with detailed answers and explanations
Instant download HTML + PDF delivered the moment payment clears.
Secure Stripe checkout we never see or store your card details.
7-day refund if files are defective see our refund policy.