All Jobs Vacancy

Senior Data Engineer GSK0JP00109754

Posted 3 days ago by Experis

Role purpose / summary

We see a world in which advanced applications of machine learning (ML) and artificial intelligence (AI) will allow us to develop novel therapies for existing diseases and to respond quickly to emerging or changing diseases with personalised drugs, driving better outcomes at reduced cost with fewer side effects.

It is an ambitious vision and delivering it will require products and solutions at the cutting edge of Machine Learning and AI.

We're looking for a senior data engineer (contractor) to help us make this vision a reality.

Key responsibilities

  • Build, operate and then automate the pipelines behind our medical imaging data and combine that imaging data (DICOM data) with other data modalities including tabular data.
  • Design data storage, transfer and transformation utilities that make it fast to move and harmonise multi-terabyte image datasets on Google Cloud Platform (GCP) by developing standardised onboarding processes with imaging sites and vendors. This includes designing secure inbound and outbound exchange with automated data transfer and Quality Control.
  • Handle derived imaging artefacts. Link segmentations and annotations (for example RTSTRUCT, NIfTI) back to their source series and deliver analysis-ready dataset snapshots into the imaging analysis platforms and ML environments for science teams to work with.
  • Automate the path for curating and harmonising imaging datasets. Parse and normalise metadata, validate incoming manifests against the agreed metadata standard, reconcile images against clinical/tabular and other biomarker data, and build quality and de-identification checks to replace manual review.
  • Write production-grade Python: tested, reviewed, instrumented and documented well enough that the team can run it long-term.
  • Work with ML engineers and software engineers to shape datasets around what the models and the science need.
  • Work with imaging leads to apply ML and AI techniques to ingestion, including by doing anomaly detection over metadata and series structure, using automated detection of missing, duplicate or mismatched studies, applying classifiers for burned-in pixel PHI, and building agent-driven triage of failed loads.

Essential qualifications

  • Strong Python for data engineering, including data transformation, job monitoring and schema management.
  • Experience handling large binary/blob datasets in object storage at scale (Google Cloud Storage, Azure Blob Storage/ADLS, Amazon S3 or equivalent) in a cloud environment.
  • Significant SQL experience, including schema design.
  • One of: a PhD in a computational discipline, MSc in computational discipline + 2 years' experience with imaging data (from any domain), or 2 years' hands-on experience with medical imaging data (such as DICOM data).
  • Experience building and running ETL/ELT pipelines with an orchestration framework.
  • Experience with CI/CD, agile software development and DevOps.
  • Ability to work with ambiguous requirements and decide on the approach independently.

Desirable qualifications

  • Hands-on experience with DICOM data; metadata and tags, multi-frame and series structure, de-identification including burned-in pixel PHI, and the standard Python tooling (pydicom, SimpleITK, dcm2niix, DCMTK).
  • Deep, practical BigQuery and/or SQL: schema design, partitioning and clustering, query and cost optimisation on large tables.
  • Agentic engineering; using coding agents as a day-to-day part of how you build and building agent-driven data workflows.
  • Google Cloud; services such as Cloud Run, GKE, Artifact Registry and Cloud SQL.
  • Modern columnar and array formats (Parquet, Arrow, Zarr) and thoughtful storage layout for large datasets.
  • Infrastructure as Code (IaC); ideally Terraform, Docker containers.
  • Experience working with sensitive data under GDPR, HIPAA or clinical trial data governance.
  • Experience building secure, audited data exchange with external parties like imaging vendors, academic collaborators and Contract Research Organisations, including staging containers, transfer controls and provenance tracking.
  • Background in biology, medicine or biomedical data (genomics, transcriptomics, proteomics, EHR, clinical images).

Key Skills: Python, GCP, DICOM, Parquet, Image Data, Cloud, SQL, BigQuery

Rate:
Not specified
Location:
London
IR35 Status:
Not specified
Remote Status:
Hybrid
Industry:
Data & Analytics
Seniority Level:
Senior

Take-Home Pay

Not Available

Visit calculators for additional details

Share job