Role Title: Senior Site Reliability Engineer (GCP)
Location: Manchester, UK
Employment Type: Inside IR35 Contract
Essential skills and experience
- Google Cloud Platform
- Strong hands-on GCP engineering and operations experience across production services.
- Relevant GCP certification is preferred.
- Site Reliability Engineering
- Demonstrable SRE experience improving availability, resilience, scalability, supportability and operational performance.
- Dynatrace & observability
- Instrumentation, dashboards, metrics, logs, traces, service health modelling and SLO-based alerting using Dynatrace.
- Terraform
- Strong Infrastructure as Code capability, including modular design, reusable components, state management, code review and maintainability.
- Kubernetes & containers
- Production Kubernetes administration and Docker knowledge, including deployments, upgrades, networking, capacity and troubleshooting.
- CI/CD engineering
- Hands-on Jenkins, Azure DevOps, GitHub Actions or comparable tooling for build, test, infrastructure and deployment automation.
- Scripting & coding
- Practical automation using Python, Groovy, Bash or PowerShell, with sound software engineering and source-control practices.
- Incident management
- Major incident response, structured troubleshooting, root-cause analysis, problem management and post-incident improvement.
- Cloud security
- Working knowledge of IAM, least privilege, secrets, secure configuration, policy controls and operational risk management.
- Networking & APIs
- Strong understanding of Linux/Unix, DNS, TCP/IP, routing, load balancing, firewalls, VPC connectivity, HTTP and API troubleshooting.
- DevOps & toil reduction
- Automation-first mindset and experience replacing repetitive operational activity with reliable engineering solutions.
- Stakeholder collaboration
- Ability to explain service risk and technical options clearly and work effectively across engineering, product, security and operations teams.