Key Responsibilities
- Design, build, and maintain cloud infrastructure within Google Cloud Platform (GCP).
- Deploy, manage, and optimise Kubernetes-based environments and containerised applications.
- Develop and support Back End services and operational tooling using Python.
- Create and maintain Infrastructure as Code using Terraform.
- Implement security hardening, monitoring, observability, and reliability best practices across cloud platforms.
- Troubleshoot and resolve complex production issues, ensuring minimal service disruption.
- Develop and improve CI/CD pipelines to enable efficient and reliable software delivery.
- Work closely with development teams to improve platform resilience, scalability, and operational performance.
- Implement monitoring, alerting, and incident response processes to support production systems.
- Drive automation initiatives to reduce operational overhead and improve system reliability.
We're Looking For
- Proven experience as a Site Reliability Engineer, Platform Engineer, Cloud Engineer, or DevOps Engineer.
- Strong hands-on experience with Google Cloud Platform (GCP).
- Experience deploying, operating, and supporting production cloud applications at scale.
- Strong knowledge of Kubernetes and container orchestration technologies.
- Commercial experience developing with Python.
- Experience building and supporting Back End services and APIs, ideally using FastAPI.
- Strong knowledge of Terraform and Infrastructure as Code principles.
- Experience designing highly available, resilient, and secure cloud environments.
- Strong troubleshooting and incident management skills within production environments.
- Experience with CI/CD pipelines, Git, GitHub, and DevOps best practices.
- Strong understanding of cloud networking concepts, with particular emphasis on GCP networking.
- Ability to take ownership of production systems and drive continuous improvement initiatives.
Desirable
- Microsoft Azure experience.
- Experience deploying and supporting AI/ML applications in production environments.
- Exposure to Large Language Models (LLMs), NLP, or agent-based systems.
- Experience with observability and monitoring tools such as Sentry, Prometheus, Grafana, or similar platforms.
- Strong Docker and multi-container application architecture experience.
- Experience working within scientific, pharmaceutical, genomics, or bioinformatics environments.
- Start-up or high-growth technology company experience.
- Open-source software contributions.
- Technical leadership, mentoring, or coaching experience.
Contract Details
- 6-Month Initial Contract
- Inside IR35
- London-based
- Hybrid Working (2 days onsite per week)
- Competitive Day Rate
- Potential extension opportunities
The Opportunity
This role offers the chance to work on modern cloud-native systems supporting innovative AI-enabled technologies.
You'll be part of a collaborative engineering environment where reliability, automation, security, and scalability are key priorities, with the opportunity to make a significant impact on critical production platforms.