
Google Data Scientist candidates report SQL and Python analysis, experiment design, product and ML cases, then behavioral or Googliness conversations.
$177K
Avg. Base Comp
$320K
Avg. Total Comp
3-5 rounds
Typical Rounds
Not reported
Process Length
Google Data Scientist interviews reported here center on analytical judgment as much as implementation. Candidates describe SQL and Python exercises built around product data, including rolling active-user or retention calculations, joins, messy event data, and explaining edge cases as they work. Window functions appear repeatedly, but the product question behind the query also matters.
Experiment design is a recurring focus. Reported prompts ask candidates to define success metrics, select a randomization approach, reason about sample size and balance, and interpret conflicting results such as stronger click-through rate alongside weaker long-term engagement. Expect follow-up questions about novelty effects, seasonality, sample-ratio mismatch, bias, p-values, and causal interpretation rather than a short definition-based statistics quiz.
Product and ML conversations can begin with an ambiguous Google-product scenario. Candidates report discussing feature success for Maps or YouTube, customer or user segmentation, recommendation design, model choice, cold start, evaluation, and model drift. A clear sequence of assumptions, metrics, data needs, and tradeoffs is more useful than naming complex models.
Behavioral or Googliness sessions are also reported, with examples involving leadership, collaboration, stakeholder disagreement, and explaining work in plain language. The exact sequence varies: reports range from three interviews to a five-round loop, and some say onsite sessions revisit earlier technical formats. End-to-end timing was not consistently reported.
Synthesized from 11 candidate reports by our editorial team.
Had an interview recently?
Share your experience. Unlock the full guide.
Real interview reports from people who went through the Google process.
The biggest surprise was that the recruiter framed the first round as a discussion of my prior ML work, but the interviewer went directly into technical questions. I was transitioning from a Business Intelligence Engineer role and had several business-driven ML projects on my resume, so I expected to spend time connecting that experience to the Data Scientist role. Instead, I had only three days' notice and was not prepared for the format.
The first-round interview included an ML domain portion and a behavioral portion. In the ML portion, I was asked to design a spam classifier and then write pseudocode for a bi-gram model. The questions were more ML-design-focused than a walkthrough of my projects, which caught me off guard and caused me to fumble the round. The behavioral portion included both standard behavioral questions and hypothetical scenarios.
I was rejected after the first round. My main takeaway is to clarify the actual interview format with the recruiter rather than relying on a broad description like "prepare to discuss your experience." For this process, I would specifically practice explaining an end-to-end spam-classifier design and writing clear pseudocode for a bi-gram model under interview pressure.
Prep tip from this candidate
Clarify whether the first round will actually cover your projects or will be technical, then practice designing a spam classifier and writing pseudocode for a bi-gram model. The recruiter description may not match the interviewer's format.
Share your own interview experience to unlock all reports, or subscribe for full access.
Sourced from candidate reports and verified by our team.
Topics based on recent interview experiences.
Featured question at Google
Write a query that returns all neighborhoods that have 0 users.
| Question | |
|---|---|
| 2nd Highest Salary | |
| Top Three Salaries | |
| First Touch Attribution | |
| First to Six | |
| Merge Sorted Lists | |
| Experiment Validity | |
| String Shift | |
| 500 Cards | |
| Last Transaction | |
| Button AB Test | |
| Top 3 Users | |
| Raining in Seattle | |
| Third Purchase | |
| Job Recommendation | |
| Minimum Change | |
| Impression Reach | |
| Jars and Coins | |
| Lazy Raters | |
| WAU vs Open Rates | |
| Bucket Test Scores | |
| Find the First Non-Repeating Character in a String | |
| Network Experiment Design | |
| Complete Addresses | |
| Delivery Estimate Model | |
| RMS Error | |
| Find Bigrams | |
| Daily Retention Summary | |
| Reducing Error Margin | |
| Random Bucketing |
Synthesized from candidate reports. Individual experiences may vary.
Candidates report SQL questions involving joins, rolling active users, retention, and window functions, sometimes paired with Python data manipulation or simulation. Some early rounds are entirely scenario-based, so be ready to clarify the business definition and explain edge cases as well as write the solution.
Candidates report designing experiments end to end: metrics and guardrails, randomization, sample size, balance checks, novelty effects, seasonality, and interpretation of conflicting metrics. Statistics discussions may include p-values, probability, Bayesian reasoning, the central limit theorem, and bias.
Reported cases include measuring a Maps feature, diagnosing engagement changes, and framing user segmentation. Candidates should typically state assumptions, identify user or product context, choose success measures, and explain what data would resolve uncertainty.
Candidates report high-level modeling conversations on YouTube engagement, recommendations, and product scenarios. Discussion may cover feature selection, model selection, cold start, offline versus online evaluation, drift, and how a system can fail in practice.
Candidates report behavioral sessions focused on leadership, work style, collaboration, difficult decisions, and handling pushback. Prepare concise examples that show how you explain analysis to skeptical stakeholders and work through disagreement.