Key Responsibilities
Platform Reliability & Operations
Azure Cloud Engineering
Infrastructure as Code (Terraform)
DevOps, Automation & AI
Resilience, Disaster Recovery & Failover
Security & Governance
Continuous Improvement
Platform Reliability & Operations
- Ensure the availability, performance, scalability, and reliability of Azure-hosted services.
- Define and manage Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets.
- Proactively monitor platform health and performance using observability tooling.
- Perform root cause analysis and implement permanent fixes for recurring incidents.
- Participate in incident management and on-call support rotations where required.
- Lead blameless post-incident reviews, capture lessons learned, and drive corrective actions through to completion.
- Reduce operational toil by identifying repetitive manual tasks and replacing them with automated, reusable engineering solutions.
- Develop reliability dashboards and actionable alerts that focus on customer-impacting symptoms rather than infrastructure noise.
Azure Cloud Engineering
- Design, deploy, and manage Azure infrastructure services including:
- Virtual Networks
- Application Gateways
- API Management
- Azure Kubernetes Service (AKS)
- Azure Firewall
- Azure Storage
- Key Vault
- Azure Monitor
- Azure AI Services
- Implement cloud platform standards and best practices.
- Support multi-region Azure deployments and platform modernisation initiatives.
- Undertake capacity planning and performance engineering to ensure platforms can scale reliably in line with business growth and peak demand.
Infrastructure as Code (Terraform)
- Develop and maintain Terraform modules and reusable infrastructure patterns.
- Implement Infrastructure as Code (IaC) standards and governance controls.
- Ensure infrastructure is version controlled, peer-reviewed, and fully automated.
- Manage Terraform state securely and consistently across environments.
DevOps, Automation & AI
- Build and maintain Azure DevOps CI/CD pipelines.
- Automate infrastructure provisioning and application deployments using pipelines with automated delivery and testing routines.
- Implement testing, security scanning, policy compliance, and release gates.
- Support DevOps and platform engineering practices.
- Create automation for operational runbooks, self-healing processes, deployment validation, and environment consistency checks.
- Manage and Implement AI platforms and tools such as Claude & Open AI, to develop skills and support business adoption of agentic AI capabilities.
- Implement and enable self-service approach to technology services.
Resilience, Disaster Recovery & Failover
- Design and implement highly available Azure architectures.
- Develop and maintain disaster recovery and business continuity capabilities.
- Implement and test:
- Regional failover strategies
- Active/Passive architectures
- Active/Active deployments
- Traffic Manager and Front Door failover patterns
- Database resiliency and replication
- Backup and recovery solutions
- Conduct regular resilience and recovery testing exercises.
- Identify and reduce single points of failure across platforms.
- Define and execute game days, chaos testing, and controlled failure scenarios to validate operational resilience.
Security & Governance
- Ensure platforms are secure-by-design.
- Work closely with Security and Architecture teams to implement:
- Zero Trust principles
- RBAC controls / Managed Identities
- Network segmentation and Zone based architecture
- Secrets management
- Support compliance requirements and operational audits.
- Help coordinate security updates, patches, maintenance routines, and upgrades of the underlying system across partners and vendors
- Embed reliability, security, and compliance controls into build and release pipelines to support production readiness.
Continuous Improvement
- Drive automation and reduction of manual operational tasks.
- Improve deployment reliability and platform observability.
- Contribute to architecture standards, runbooks, and operational documentation.
- Partner with engineering, architecture, security, and service teams to define production readiness standards and reliability acceptance criteria.