
EY Data Engineer candidates report SQL- and PySpark-focused technical interviews, project discussions, production scenarios, and a behavioral or HR conversation. Some candidates also encountered a timed SQL assessment.
$133K
Avg. Base Comp
$157K
Avg. Total Comp
3 rounds
Typical Rounds
1 days
Process Length
EY Data Engineer interviews reported here are practical and strongly tied to the work candidates describe on their resumes. SQL is the clearest recurring technical theme. Candidates reported window functions, joins, CTEs, nested queries, aggregations, duplicate removal, ranking, and consecutive-record problems. One candidate completed a timed online SQL assessment with 20 multiple-choice questions and three minutes per question, while another reported SQL-heavy technical discussions and practical query scenarios.
PySpark and cloud data engineering also recur. Candidates were asked about partitioning, bucketing, transformations, Spark internals, Dataflow pipeline code, GCP services, BigQuery, Dataproc, Airflow, schema evolution, and pipeline control tables. Several reports describe follow-up questions that move from a direct concept to how the candidate would apply it in a production pipeline, such as handling duplicates, troubleshooting an issue, or designing an idempotent process.
Project communication matters alongside technical depth. Candidates reported explaining architecture, business logic, transformations, past challenges, and the reasoning behind design choices. Behavioral or HR conversations covered background, teamwork, pressure, strengths and weaknesses, motivation, and career changes. One report described two technical rounds followed by HR, while another described a screen and an L1 technical round with later steps still unknown; formats may therefore vary. Prepare concise project stories and practice talking through SQL and pipeline decisions as you solve them.
Synthesized from 6 candidate reports by our editorial team.
Had an interview recently?
Share your experience. Unlock the full guide.
Real interview reports from people who went through the Ey process.
The process began with a 15–20 minute technical call about my experience and fundamentals. I then completed an L1 technical interview on Microsoft Teams that lasted about 45 minutes.
The L1 round went deeper into my data engineering experience, especially GCP, BigQuery, PySpark, SQL, Airflow, and scenario-based questions. I was most comfortable discussing projects I had worked on and explaining why I chose particular approaches. The harder part was moving from direct technical questions to real-time production scenarios involving troubleshooting or design.
Questions included explaining a data engineering project architecture; ingesting data into GCP; PySpark transformations with Dataproc; the Catalyst Optimizer and Tungsten Engine; Spark DAG stages; surrogate versus primary keys; duplicate removal in SQL; schema evolution; audit and control tables; passing data between Airflow DAGs; BigQuery slots and authorized views; materialized-view SQL; and project business logic and transformations.
At the time of writing, I had completed the initial screen and L1 technical round and was waiting for the next step.
Prep tip from this candidate
Prepare a concise walkthrough of your data engineering projects and the tradeoffs behind them. Review GCP, BigQuery, PySpark, SQL, Airflow, and production troubleshooting scenarios.
Share your own interview experience to unlock all reports, or subscribe for full access.
Sourced from candidate reports and verified by our team.
Topics based on recent interview experiences.
Featured question at Ey
Select the 2nd highest salary in the engineering department
| Question | |
|---|---|
| Top 3 Users | |
| Longest Streak Users | |
| Find the First Non-Repeating Character in a String | |
| Size of Joins | |
| Skyscanner Partner ETL | |
| Resumable Fact Table Load | |
| Sort Strings | |
| Xgboost vs Random Forest | |
| Hurdles In Data Projects | |
| Real-Time Hashtag Partitioning | |
| Duplicate Rows | |
| Classification and Regression | |
| Mouse Search | |
| Pipeline Transformation Failures | |
| Azure Kubernetes Infrastructure | |
| Simple Explanations | |
| Relational Migration | |
| Why Do You Want to Work With Us | |
| Marketing Workflow Optimization | |
| Processing Large CSV | |
| Your Strengths and Weaknesses | |
| Stakeholder Communication | |
| Data Cleaning Experiences | |
| Linear vs Logistic Regression | |
| Backpropagation Explanation | |
| Analyzing Multiple Data Sources | |
| Rolling Bank Transactions | |
| Closest SAT Scores | |
| Merge Sorted Lists |
Synthesized from candidate reports. Individual experiences may vary.
Candidates report either an HR conversation about role, skills, and expectations or a 15–20 minute technical screen focused on experience and fundamentals. Be ready to summarize your background and the data engineering work you have done.
One candidate reported an online SQL assessment with 20 multiple-choice questions and three minutes per question, centered on multi-table joins, nested queries, and aggregations. This assessment was not reported by every candidate.
Candidates report technical interviews covering SQL, Python, PySpark, GCP, BigQuery, Airflow, and pipeline architecture. Questions may ask you to solve SQL problems, explain Spark behavior, or discuss how you have used the tools in projects.
Candidates report scenario questions about production issues, schema evolution, duplicate handling, data ingestion, and idempotent pipelines. Expect to explain your project architecture, transformations, and the reasoning behind a design or troubleshooting approach.
Some candidates report a final managerial, behavioral, or HR discussion covering teamwork, pressure, strengths and weaknesses, motivation, prior roles, and project challenges. The exact placement and format may vary.