Company Introduction
We have an exciting opportunity now available with one of our sector-leading consultancy clients! They are currently looking for a skilled SRE to join their team for a six-month contract.
Job Responsibilities/Objectives
- Implement SLOs, SLIs, error budgets and service health measures for critical services.
- Build and maintain observability patterns covering monitoring, logging, tracing, alert quality and dashboards.
- Support production readiness reviews, resilience reviews, capacity planning and operational acceptance.
- Automate repetitive operational tasks and reduce manual toil across supported services.
- Contribute to incident response, post-incident reviews and reliability improvement actions.
- Partner with platform and software engineers to embed reliability patterns into delivery pipelines.
- Improve alerting quality by reducing noise, tuning thresholds and linking alerts to actionable runbooks.
- Support AIOps and anomaly detection adoption where it improves service restoration and prediction.
Required Skills/Experience
- Experience in operations, software, cloud, platform or reliability engineering.
- Knowledge of monitoring, logging, tracing, incident response, automation and capacity management.
- Ability to work across technical and service management teams.
- Strong analytical approach to incident trends, service performance and root cause prevention.
- Scripting or automation capability and willingness to improve operability through engineering.
- Practical understanding of ITIL4 service operations and SRE practices.