
Citi Data Engineer candidates report SQL, Python or PySpark, data-engineering fundamentals, and project discussion across assessment, Karat, and interview stages.
$128K
Avg. Base Comp
$204K
Avg. Total Comp
3 rounds
Typical Rounds
Not reported
Process Length
Citi Data Engineer interview reports point to a practical blend of SQL, programming, data-engineering fundamentals, and discussion of your own work. One candidate described a three-stage path beginning with a 90-minute online assessment containing medium SQL, a medium PySpark task, and general data-engineering questions. That candidate then had a one-hour Karat video interview on SQL and data engineering before an onsite split between technical and behavioral discussion.
A separate candidate also described a Karat screening, but as a tightly timed session: general Ab Initio questions, a live SQL exercise, and coding in a programming language of the candidate’s choice. Practice producing SQL in a shared, timed environment, including joins, window functions, CTEs, CASE expressions, grouping, and derived columns. These topics were named across reports, rather than presented as a universal question set.
Technical conversations may also probe Python, PySpark, Spark internals, testing, and system-design judgment. One offer recipient was asked how to process 10 TB of data in Databricks from a resume project; another said the first interview focused heavily on Python. Be ready to explain architecture choices, scaling tradeoffs, and what you personally built instead of reciting tool names. Behavioral discussion included challenging projects, feedback, conflict, and job preferences. Evidence is limited to a small set of candidate reports, so confirm team-specific expectations with the recruiter.
Synthesized from 4 candidate reports by our editorial team.
Had an interview recently?
Share your experience. Unlock the full guide.
Real interview reports from people who went through the Citi process.
The process was fairly technical but conversational rather than a coding-heavy gauntlet. I had three interview rounds, with the first two focused mainly on PySpark, Spark internals, SQL, and some system-design discussion. The questions tested whether I understood the underlying Spark concepts as well as how I had used them, so I made sure to explain my project decisions clearly instead of only naming tools.
The final round leaned more toward my previous projects, testing, and domain knowledge. One question that stood out was how I would process 10 TB of data in Databricks, based on a project listed on my resume. I was expected to talk through a practical approach and connect it to my hands-on experience. The interviewers also asked about the tools I had worked with, and the overall conversation rewarded being articulate when explaining technical ideas. I received an offer. My main advice is to be ready to defend every Spark and Databricks project on your resume, especially how you would scale it, and to refresh core PySpark, Spark internals, and SQL concepts.
Prep tip from this candidate
Prepare a clear, practical explanation for processing large-scale data in Databricks, including a 10 TB scenario, and be ready to discuss the PySpark, Spark-internals, SQL, and testing choices in projects listed on your resume.
Share your own interview experience to unlock all reports, or subscribe for full access.
Sourced from candidate reports and verified by our team.
Topics based on recent interview experiences.
Featured question at Citi
Write a query to show the number of users, transactions, and total order amount per month in 2020
| Question | |
|---|---|
| Attribution Rules | |
| Size of Joins | |
| Real-Time Transaction Streaming | |
| Hurdles In Data Projects | |
| Replace Words with Stems | |
| Dijkstra implementation | |
| Seller Type Modeling | |
| String Palindromes | |
| Index Fund Return | |
| Concurrent LLM Serving | |
| Above Average Product Prices | |
| International e-Commerce Warehouse | |
| Text Editor With OOP | |
| Unified Event Pipeline | |
| Client Solution Pushback | |
| Why Do You Want to Work With Us | |
| Concentric Circles | |
| Your Strengths and Weaknesses | |
| Data Cleaning Experiences | |
| Student Tests | |
| Measuring Text Difficulty | |
| Singly Linked List | |
| 2nd Highest Salary | |
| Empty Neighborhoods | |
| Rolling Bank Transactions | |
| Employee Salaries | |
| Merge Sorted Lists | |
| Subscription Overlap | |
| Comments Histogram |
Synthesized from candidate reports. Individual experiences may vary.
One candidate reported a 90-minute assessment with medium SQL, a medium PySpark exercise, and ten general data-engineering questions. Their examples included CTEs, CASE expressions, grouping, functions, and data transformations.
Two candidates reported Karat-led video screening rather than an initial Citi interviewer. Reported content included SQL and data engineering; one account also described brief Ab Initio discussion and timed live SQL and programming exercises.
Candidates report technical discussion of Python, PySpark, Spark internals, SQL, testing, and project decisions. One report described an onsite split into technical and behavioral sessions, with questions about projects, feedback, conflict, and role preferences.