A data scientist is attempting to identify sentences that are conceptually similar to each other within a set of text files. Which of the following is the best way to prepare the data set to accomplish this task after data ingestion?
A. Embeddings
B. Extrapolation
C. Sampling
D. One-hot encoding
Show Answer
Correct Answer: A
Explanation: Embeddings represent sentences as numeric vectors that capture semantic meaning, allowing conceptually similar sentences to be compared by vector similarity. One-hot encoding does not capture semantic relationships; extrapolation and sampling are not suitable preparation methods for this task.
Question 12
Which of the following types of machine learning is a GPU most commonly used for?
A. Deep learning/neural networks
B. Clustering
C. Natural language processing
D. Tree-based
Show Answer
Correct Answer: A
Explanation: GPUs are most commonly used for deep learning and neural networks because their parallel processing architecture efficiently handles the large matrix and tensor computations these models require.
Question 13
Which of the following is the naive assumption in Bayes' rule?
A. Normal distribution
B. Independence
C. Uniform distribution
D. Homoskedasticity
Show Answer
Correct Answer: B
Explanation: Naive Bayes assumes that features are conditionally independent of one another given the class label.
Question 14
An analyst is examining data from an array of temperature sensors and sees that one sensor consistently returns values that are much higher than the values from the other sensors. Which of the following terms best describes this type of error?
A. Synthetic
B. Systematic
C. Heteroskedastic
D. Idiosyncratic
Show Answer
Correct Answer: B
Explanation: A sensor that consistently reports higher values has a persistent bias, which is a systematic error.
Question 15
A data scientist has built a model that provides the likelihood of an error occurring in a factory. The historical accuracy of the model is 90%. At a specific factory, the model is reporting a likelihood score of 0.90. Which of the following explains a confidence score of 0.90?
A. Running this model for all known factory issues, it is expected the model will identify 90 out of 100 known factory issues.
B. Running this model on 100 samples of factories, a certain model performance is expected for 90 out of the 100 samples.
C. Running this model 100 times on a factory, it is expected the model will predict 90 out of 100 factory errors.
D. Running this model 100 times within a factory, it is expected the model will predict error 90 out of 100 times the model is ran.
Show Answer
Correct Answer: D
Explanation: A likelihood score of 0.90 for this factory means the model estimates a 90% probability that an error will occur in this case. The model’s historical accuracy of 90% is a separate measure and does not itself explain this score.
Question 16
A data scientist is preparing to brief a non-technical audience that is focused on analysis and results. During the modeling process, the data scientist produced the following artifacts:
Charts and dashboards
Model performance statistics (accuracy, precision, recall, F1 score, etc.)
Mathematical descriptions of clustering algorithms included in the selected model
Model selection, justification, and purpose
Code documentation
Data dictionary
Which of the following artifacts should the data scientist include in the briefing? (Choose two.)
A. Final charts and dashboards
B. Model selection, justification, and purpose
C. Code documentation
D. Mathematical descriptions of clustering algorithms included in the selected model
E. Model performance statistics (accuracy, precision, recall, F1 score, etc.)
F. Data dictionary
Show Answer
Correct Answer: A, E
Explanation: Final charts and dashboards make the findings accessible to a non-technical audience. Model performance statistics show how reliable the results are, provided the key metrics are explained in plain language. Algorithm mathematics, code documentation, and data definitions are better suited to a technical briefing.
Sources:
https://www.micro1.ai/interview-prep/data-scientist-interview-questions
https://vegavid.com/blog/what-is-accuracy-precision-recall-and-f1-score
Question 17
A movie production company would like to find the actors appearing in its top movies using data from the tables below. The resulting data must show all movies in Table 1, enriched with actors listed in Table 2.
Which of the following query operations achieves the desired data set?
A. Perform an INNER JOIN between Table 1 using column Movie, and Table 2 using column Acted_In.
B. Perform a UNION between Table 1 using column Movie, and Table 2 using column Acted_In.
C. Perform an INTERSECT between Table 1 using column Movie, and Table 2 using column Acted_In.
D. Perform a LEFT JOIN on Table 1 using column Movie, with Table 2 using column Acted_In.
Show Answer
Correct Answer: D
Explanation: A LEFT JOIN keeps every movie from Table 1 and adds matching actor data from Table 2, with null actor fields when no match exists.
Question 18
A data scientist is designing a real-time machine-learning model that classifies a user based on initial behavior. The run times of these models are provided in the following table:
Which of the following models should the data scientist recommend for deployment?
A. XGBoost
B. Random forest
C. Decision trees
D. Artificial neural network
Show Answer
Correct Answer: B
Explanation: Random forest has the shortest run time in the referenced comparison, making it the best fit for a real-time deployment when minimizing classification latency is the priority.
Question 19
A data analyst is examining the correlation matrix of a new data set to identify issues that could adversely impact model performance. Which of the following is the analyst most likely checking for?
A. Undersampling
B. Multicollinearity
C. Oversampling
D. Overfitting
Show Answer
Correct Answer: B
Explanation: A correlation matrix can reveal highly correlated predictor variables, indicating multicollinearity. This can make model estimates unstable and harder to interpret.
Question 20
Which of the following problem-solving approaches is a set of guidelines to handle highly variable and not fully apparent situations?
A. Schedule
B. Plan
C. Heuristic
D. Algorithm
Show Answer
Correct Answer: C
Explanation: A heuristic is a flexible guideline or rule of thumb used to solve problems when situations are variable, uncertain, or not fully defined. An algorithm, by contrast, is a fixed step-by-step procedure.
$19
Get all 82 questions with detailed answers and explanations
Instant download HTML + PDF delivered the moment payment clears.
Secure Stripe checkout we never see or store your card details.
7-day refund if files are defective see our refund policy.