
Gsk Data Engineer interview typically runs 4 rounds: Hiring Manager, system design, data modeling, and CI/CD. It usually takes about 1-2 weeks and can feel inconsistent across rounds.
$102K
Avg. Base Comp
$186K
Avg. Total Comp
4
Typical Rounds
2-4 weeks
Process Length
Our candidates report that GSK is looking for someone who can hold the whole data stack in their head, not just one slice of it. The strongest signal in the experience we saw was the system design discussion around an end-to-end unstructured data pipeline from FTP sources: that’s the kind of prompt where they seem to care less about buzzwords and more about whether you can reason through ingestion, processing, and operational tradeoffs in a healthcare context. The repeated appearance of image, PDF, and SFTP pipeline questions suggests they want engineers who are comfortable with messy, real-world data rather than tidy textbook examples.
A recurring theme is that the interviewers may change the problem midstream, and candidates who stayed rigid seemed to struggle. In one round, the discussion moved from images to PDFs, then to Hadoop-to-cloud migration, then into fact and dimension tables, and finally into Python debugging. That tells us GSK is probably testing adaptability as much as technical depth. We’ve also seen Spark behavior questions tied to uneven file sizes, which hints that they care about practical performance awareness, not just architecture diagrams.
What makes or breaks candidates here is the ability to connect domains quickly: storage, modeling, migration, and code-level reasoning. The people who do best are the ones who can keep their answer coherent even when the interviewer pivots, and who can explain why a design choice still holds when the data shape changes. In other words, GSK seems to value broad, composable engineering judgment over a narrow specialty.
Synthesized from 1 candidate report by our editorial team.
Had an interview recently?
Share your experience. Unlock the full guide.
Real interview reports from people who went through the Gsk process.
{ "experience": "The process for this senior data engineer role felt all over the place, and the most frustrating part was that the interview kept changing direction mid-conversation. It started with a Hiring Manager round and then a system design discussion, which were the only two rounds that felt reasonably aligned with the role. The system design question was about designing an end-to-end solution for unstructured data processing from FTP servers, and that part was at least relevant and fairly standard for a data engineering interview.
The later rounds were much less consistent. In the data modelling round, I was first asked how I would process 50 GB of images, and when I started explaining that, the interviewer changed it to PDFs. I walked through the processing flow and mentioned sequencing to avoid zoom-related issues, which he eventually accepted, but then he shifted again and asked how I would migrate a Hadoop pipeline to the cloud. After that, he moved into designing fact and dimension tables, but instead of staying with SQL, he pushed me into Python and gave me a debugging-style problem around sliding windows and nested loops, asking how I would improve the code. That round felt disjointed and hard to follow because the topic kept changing before I could finish a complete answer. There was also a CI/CD discussion where I was asked what I knew about it, and a Spark question about how the system behaves when one file is very small and another is very large.
Overall, the process felt poorly structured and the expectations were not clearly aligned across rounds. I did not get an offer, and the feedback at the end was generic. My main takeaway is to be ready for broad data engineering topics, but also expect the interview to jump between data modelling, cloud migration, Spark behavior, and code debugging without much warning.
outcome": "No offer
outcome_color": "red
prep_tip": "Be ready to explain an end-to-end unstructured data pipeline from FTP sources, then pivot quickly into Hadoop-to-cloud migration and fact/dimension modeling. I’d also review Spark file-size behavior and practice explaining how you’d debug and improve a sliding-window/nested-loop Python problem, since that came up unexpectedly." }
Share your own interview experience to unlock all reports, or subscribe for full access.
Sourced from candidate reports and verified by our team.
Topics based on recent interview experiences.
Featured question at Gsk
Design a data warehouse for a new online retailer
| Question | |
|---|---|
| Hurdles In Data Projects | |
| Moving Window | |
| Rider Discount | |
| Image Classification Pipeline | |
| Why Do You Want to Work With Us | |
| Marketing Workflow Optimization | |
| Your Strengths and Weaknesses | |
| SFTP Pipeline | |
| 2nd Highest Salary | |
| Cumulative Distribution | |
| Experiment Validity | |
| Last Transaction | |
| Monthly Customer Report | |
| Total Spent on Products | |
| RMS Error | |
| Detecting ECG Tachycardia Runs | |
| Size of Joins | |
| Random Forest Explanation | |
| Cumulative Reset | |
| Time Difference | |
| Subscription Retention | |
| Brain Cancer Treatment Outcomes | |
| Always Excited Users | |
| Sum to Zero | |
| Missing Housing Data | |
| P-value to a Layman | |
| Flatten JSON | |
| Valid Anagram | |
| Reducing Error Margin |
Synthesized from candidate reports. Individual experiences may vary.
The process started with a hiring manager interview focused on the candidate’s background and fit for the senior data engineer role. This round also led into technical discussion, so it was not purely behavioral.
A system design discussion followed, centered on designing an end-to-end solution for unstructured data processing from FTP servers. This was the most standard and role-aligned round in the process.
This round covered data modeling and broader engineering topics, but the discussion shifted repeatedly. The interviewer moved between processing image and PDF data, migrating a Hadoop pipeline to the cloud, and designing fact and dimension tables.
The later technical round included a Python debugging-style problem involving sliding windows and nested loops, along with questions about CI/CD and Spark behavior when file sizes are highly uneven. The conversation was broad and somewhat disjointed, with topics changing mid-interview.