🎉 Limited offer: Get 3 months free when you sign up for an annual plan. See pricing →

Interview guide · Technology

Data Scientist interview questions

A data scientist interview should test whether the candidate can frame a business problem as a modeling problem, choose sensible methods, and validate results honestly. Probe for experimental rigor, awareness of bias and data leakage, and the ability to explain models to decision-makers.

What to assess

Statistical modelingMachine learning fundamentalsExperiment designPython or R programmingProblem framingCommunicating uncertainty

Behavioral questions

  1. Tell me about a model you built that made it into production. What impact did it have?

    What it reveals: Shows end-to-end experience beyond notebooks.

    A strong answer: Covers problem framing, data, model choice, deployment, monitoring, and a measurable business result.

  2. Describe a time a model performed well in testing but poorly in the real world.

    What it reveals: Reveals understanding of drift, leakage, and validation gaps.

    A strong answer: Diagnoses the cause specifically and explains what they changed in their validation process.

  3. Give an example of when a simple approach beat a complex one in your work.

    What it reveals: Tests pragmatism over novelty.

    A strong answer: Explains why the simpler model was more robust, interpretable, or cheaper to maintain.

  4. Tell me about a time you had to tell stakeholders their hypothesis was not supported by the data.

    What it reveals: Shows integrity and communication skill.

    A strong answer: Delivers the result clearly with evidence and suggests next steps rather than burying it.

  5. Describe a project where you worked closely with engineers to deploy your work.

    What it reveals: Tests collaboration across disciplines.

    A strong answer: Shows they wrote reusable, tested code and understood production constraints like latency.

Role-specific questions

  1. How would you detect and prevent data leakage when building a churn prediction model?

    What it reveals: Tests a frequent and costly modeling mistake.

    A strong answer: Mentions time-based splits, excluding features known only after the outcome, and careful pipeline design.

  2. How do you choose an evaluation metric for an imbalanced classification problem like fraud detection?

    What it reveals: Tests metric selection tied to business cost.

    A strong answer: Moves beyond accuracy to precision, recall, PR-AUC, and ties thresholds to the cost of errors.

  3. Walk me through designing an A/B test for a new pricing page.

    What it reveals: Tests experiment design fundamentals.

    A strong answer: Defines the metric, sample size and power, randomization unit, duration, and guardrail metrics.

  4. Explain the bias-variance trade-off and how it shows up in practice.

    What it reveals: Checks core ML understanding.

    A strong answer: Explains it clearly with examples like overfitting and names regularization or cross-validation as tools.

  5. How would you check a hiring or lending model for unfair bias across groups?

    What it reveals: Tests awareness of responsible AI risks.

    A strong answer: Mentions comparing error rates and outcomes across groups, proxy variables, and involving legal or compliance.

Situational questions

  1. Leadership wants a demand forecast in two weeks, but historical data has major gaps. What do you do?

    What it reveals: Shows how they handle imperfect data and deadlines.

    A strong answer: Sets expectations, proposes a simple baseline with clear uncertainty, and plans improvements.

  2. An A/B test shows a statistically significant win, but the effect is tiny and the product team is excited. How do you advise them?

    What it reveals: Tests ability to separate statistical and practical significance.

    A strong answer: Explains effect size, confidence intervals, and whether the gain justifies the cost of the change.

  3. Your model's accuracy drops suddenly after a product change. How do you respond?

    What it reveals: Reveals monitoring and troubleshooting habits.

    A strong answer: Checks input data distributions, pipeline changes, and retrains or rolls back with stakeholder communication.

Motivation and fit

  1. What kind of problems would you be most excited to work on here?

    What it reveals: Measures motivation and interest in the business.

    A strong answer: Connects their interests to specific company problems they have researched.

  2. How do you balance research exploration with delivering results on a schedule?

    What it reveals: Shows working style fit for a business setting.

    A strong answer: Describes timeboxing exploration and shipping incremental value.

Red flags

  • Reports only accuracy without considering the business cost of errors
  • Cannot explain how they validated a model
  • Prefers complex methods without justifying them
  • Shows no concern about bias or fairness in sensitive models

Questions not to ask

  • What country are your parents from? — national origin discrimination risk
  • Do you need a visa now or will you in the future? — ask only the two standard work-authorization questions in a consistent way
  • Are you pregnant or planning to be during the project timeline? — pregnancy discrimination risk
  • How many sick days did you take last year? — can reveal disability and violate ADA pre-offer limits

See legal and illegal interview questions.

Interviewing for data scientist roles?

Generate a structured kit with scoring guidance for your exact role in seconds.

Try the free interview kit

Frequently asked questions

How do I evaluate a data scientist's portfolio?

Look for projects that start with a clear question, show careful validation, and end with a conclusion someone could act on. Ask the candidate to walk through one project and explain what they would do differently now.

Should a data scientist job description require a PhD?

Only if the role involves original research. For most business-focused roles, an advanced degree or equivalent applied experience is a fairer and broader requirement.

What is a fair technical assessment for data scientists?

A short case study using a realistic dataset, followed by a discussion of the approach, tends to work better than puzzle questions. Score every candidate against the same rubric covering framing, method, validation, and communication.

More interview guides

Weekly HR Digest

HR insights in your inbox

Join 2,400+ HR professionals. No spam, ever.

Free forever. Privacy policy