About the Role
To strengthen our AI for Science (AI4S) team, we are looking for Software Engineers with a track record in developing production-grade, data-driven software solutions. You will design, build and operate the scalable cloud infrastructure and services - including the serving of our models - that our AI systems and agentic applications run on, and you will be accountable for keeping them reliable in production. This is hands-on software and platform engineering: building robust, well-tested, high-performance systems that scientists across the client nd on every day, on modern cloud technology and the vast biomedical data sources available to us.
In this role you will
- Design, build and operate scalable infrastructure and services that support our AI models and agentic systems across the entire software development life cycle.
- Own the reliability of what you build - set up CI/CD and release processes, automated testing, monitoring and alerting, and lead the response when things break, so the systems scientists rely on stay dependable.
- Build and operate the model-serving infrastructure that exposes our models in production with efficient use of compute.
- Develop and maintain cloud-native architectures that enable reliable deployment and scaling of AI/ML workloads.
- Deliver robust, tested and high-performance code in an agile environment, and work closely with ML engineers and domain experts to make the infrastructure fit for purpose.
Qualifications & Skills
- Demonstrated advanced programming expertise in Python and in developing and delivering robust, scalable software solutions using frameworks like FastAPI.
- Experience with cloud platforms (GCP, Azure) and cloud-native architectures.
- Passion for software design and commitment to the development of reusable, scalable, and testable software components.
- Basic understanding of at least one major deep learning framework (PyTorch, JAX, TensorFlow).
- Hands-on experience with Google Cloud Platform, in particular the services we build on: Cloud Run, Google Kubernetes Engine, Cloud Storage, Artifact Registry, Cloud SQL.
- Fluency in English.
Preferred Qualifications & Skills: If you have the following characteristics, it would be a plus:
- Familiarity with machine learning principles and state-of-the-art modelling approaches.
- Experience in design, development and deployment of commercial cloud-native software and infrastructure.
- Experience building and deploying large-scale AI models and agentic systems in production environments.
- Experience architecting, developing, and deploying distributed training pipelines for large models with PyTorch or TensorFlow.
- Expertise in performance optimization, cost optimization, and efficient compute resource management in cloud environments.
- Experience running production services at scale, including defining and working to service-level objectives (SLOs/SLIs).
- Experience with incident response and post-incident review, and with building the observability that supports it.
- Infrastructure-as-code (eg Terraform) for provisioning and maintaining cloud environments.
- Experience developing and administering workloads on Kubernetes (eg GKE).
- Familiarity with GCP networking and security controls - VPC, VPC Service Controls (VPC-SC), and private connectivity.
- Contributions to relevant open-source projects.
- Knowledge or interest in disease biology, molecular biology and medicine.
- Experience working with biomedical data (eg, genomics, transcriptomics, proteomics, electronic health records, clinical images).