
Deloitte Data Engineer candidates report practical Spark, SQL, Python, cloud, and pipeline-design interviews, with process structures ranging from three to four rounds.
$102K
Avg. Base Comp
$128K
Avg. Total Comp
3-4 rounds
Typical Rounds
4-5 months
Process Length
Deloitte Data Engineer interviews reported here are practical and stack-specific rather than algorithm-heavy. Candidates describe SQL and Python exercises alongside PySpark or Spark discussions: examples include second-highest-salary queries, joins, window functions, MERGE logic, and explaining how Spark jobs behave. Be ready to narrate assumptions and trade-offs, since one candidate had to formulate SQL verbally without a visible schema.
Pipeline and architecture reasoning is equally prominent. Candidates report troubleshooting a slow Spark job, joining large tables, addressing the small-file problem, describing incremental loads and dimensional modeling, and designing a multi-source pipeline into a cloud warehouse. Cloud emphasis varied across accounts—GCP, Azure, AWS, Databricks, Airflow, and Snowflake all appeared—so prepare the platforms and project decisions most relevant to your background instead of assuming one standard stack.
Several accounts also include a managerial, HR, behavioral, or presentation conversation. Expect questions about your projects, prior experience, motivation, teamwork, ambiguous requirements, and how you communicate a recommendation. One candidate reported a presentation exercise based on a sparse brief. The exact sequence varies, but concise project stories that connect design choices to reliability, monitoring, cost, and data quality will help across the reported formats.
Timing evidence is limited. One candidate described a process lasting more than four months, while other accounts did not provide a complete timeline.
Synthesized from 8 candidate reports by our editorial team.
Had an interview recently?
Share your experience. Unlock the full guide.
Real interview reports from people who went through the Deloitte process.
The round that stood out most was a troubleshooting discussion about a Spark job that was taking much longer than expected. I had to explain how I would identify the bottleneck and then optimize the job, so the emphasis was on working through the diagnosis clearly rather than just naming Spark features. That reflected the interview process overall: it was fairly structured, with three rounds consisting of two technical interviews followed by a managerial interview.
The technical conversations covered Apache Spark, SQL, performance optimization, data architecture, and cloud technologies, particularly Azure and AWS. One of the main design questions asked me to build a scalable pipeline that could ingest data from multiple source systems into a cloud data warehouse. The interviewers wanted to hear the reasoning behind the architecture and understand the role of each component. I was also asked about ETL work and technologies I had used previously, with follow-ups tied closely to my projects and experience. Being able to explain a past architecture in detail was useful because the discussion went beyond the high-level diagram and into why individual components were there.
There were also more direct implementation-style topics in the broader technical portion, including SQL logic such as finding the second-highest salary and implementing a slowly changing dimension Type 2. Python and PySpark fundamentals could come up as well, including small exercises involving anagrams or merging lists. These questions were not especially algorithm-heavy; the harder part was switching between practical coding, Spark performance, and end-to-end system design while keeping the explanations precise.
The managerial round was more conversational and focused on my background, previous experience, projects, and general fit. Overall, the process felt clear and manageable, but the technical scope was broad. I did not receive an offer. My main takeaway is to prepare to defend every part of a data architecture you present and to connect cloud design choices with concrete Spark and SQL performance considerations.
Prep tip from this candidate
Practice diagnosing a slow Spark job step by step, and rehearse designing a multi-source pipeline into an Azure or AWS cloud warehouse while explaining every architecture component. Also review second-highest-salary SQL, SCD Type 2 implementation, and small Python exercises involving anagrams or list merging.
Share your own interview experience to unlock all reports, or subscribe for full access.
Sourced from candidate reports and verified by our team.
Topics based on recent interview experiences.
Featured question at Deloitte
Select the 2nd highest salary in the engineering department
| Question | |
|---|---|
| Merge Sorted Lists | |
| Maximum Profit | |
| Size of Joins | |
| Real-Time Transaction Streaming | |
| Normalize Grades | |
| Cyclic Detection | |
| Skyscanner Partner ETL | |
| Missing Housing Data | |
| Find Duplicate Numbers in a List | |
| Hurdles In Data Projects | |
| Portfolio Platform Architecture | |
| Duplicate Rows | |
| Using R Squared | |
| Classification and Regression | |
| Slow SQL Query | |
| Data Pipelines and Aggregation | |
| Swap Variables | |
| User Event Data Pipeline | |
| Bias vs. Variance Tradeoff | |
| Data Preparation for Imbalanced Data | |
| Assumptions of Linear Regression | |
| Seller Type Modeling | |
| Open Source Reporting Pipeline | |
| String Palindromes | |
| Digital Classroom System Design | |
| Impossibly Iterative Fibonacci | |
| Blob Indexing | |
| Yelp-like System | |
| Text Editor With OOP |
Synthesized from candidate reports. Individual experiences may vary.
Candidates report an opening technical discussion covering Spark or PySpark, SQL, Python, data engineering fundamentals, and prior projects. One account began with a proctored test. Questions may include basic queries, Python tasks, ETL scenarios, or explaining JSON handling and pipeline choices.
Candidates report more in-depth technical rounds on Spark optimization, large-table joins, small files, SQL, dimensional modeling, incremental loads, and warehouse concepts. Some were asked to reason aloud through an incomplete problem or review PySpark code, so explain trade-offs rather than only naming a tool.
Candidates report scenarios such as diagnosing a slow Spark job or designing a scalable pipeline from multiple sources into a cloud warehouse. Cloud and platform topics varied by account, including Azure, GCP, AWS, Databricks, Airflow, and storage design; project follow-ups may probe why components were selected.
Candidates report final conversations about their background, projects, teamwork, motivation, career goals, and fit. Some accounts describe a managerial round, HR discussion, or presentation exercise; prepare concise examples that communicate technical decisions clearly under ambiguity.