All Jobs Vacancy

Site Reliability Engineer/Manager Azure & AI Platforms

Posted 1 week ago by Dminds Solutions Inc.

Key Responsibilities

Platform Reliability & Operations

Azure Cloud Engineering

Infrastructure as Code (Terraform)

DevOps, Automation & AI

Resilience, Disaster Recovery & Failover

Security & Governance

Continuous Improvement

Platform Reliability & Operations

  • Ensure the availability, performance, scalability, and reliability of Azure-hosted services.
  • Define and manage Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets.
  • Proactively monitor platform health and performance using observability tooling.
  • Perform root cause analysis and implement permanent fixes for recurring incidents.
  • Participate in incident management and on-call support rotations where required.
  • Lead blameless post-incident reviews, capture lessons learned, and drive corrective actions through to completion.
  • Reduce operational toil by identifying repetitive manual tasks and replacing them with automated, reusable engineering solutions.
  • Develop reliability dashboards and actionable alerts that focus on customer-impacting symptoms rather than infrastructure noise.

Azure Cloud Engineering

  • Design, deploy, and manage Azure infrastructure services including:
  • Virtual Networks
  • Application Gateways
  • API Management
  • Azure Kubernetes Service (AKS)
  • Azure Firewall
  • Azure Storage
  • Key Vault
  • Azure Monitor
  • Azure AI Services
  • Implement cloud platform standards and best practices.
  • Support multi-region Azure deployments and platform modernisation initiatives.
  • Undertake capacity planning and performance engineering to ensure platforms can scale reliably in line with business growth and peak demand.

Infrastructure as Code (Terraform)

  • Develop and maintain Terraform modules and reusable infrastructure patterns.
  • Implement Infrastructure as Code (IaC) standards and governance controls.
  • Ensure infrastructure is version controlled, peer-reviewed, and fully automated.
  • Manage Terraform state securely and consistently across environments.

DevOps, Automation & AI

  • Build and maintain Azure DevOps CI/CD pipelines.
  • Automate infrastructure provisioning and application deployments using pipelines with automated delivery and testing routines.
  • Implement testing, security scanning, policy compliance, and release gates.
  • Support DevOps and platform engineering practices.
  • Create automation for operational runbooks, self-healing processes, deployment validation, and environment consistency checks.
  • Manage and Implement AI platforms and tools such as Claude & Open AI, to develop skills and support business adoption of agentic AI capabilities.
  • Implement and enable self-service approach to technology services.

Resilience, Disaster Recovery & Failover

  • Design and implement highly available Azure architectures.
  • Develop and maintain disaster recovery and business continuity capabilities.
  • Implement and test:
  • Regional failover strategies
  • Active/Passive architectures
  • Active/Active deployments
  • Traffic Manager and Front Door failover patterns
  • Database resiliency and replication
  • Backup and recovery solutions
  • Conduct regular resilience and recovery testing exercises.
  • Identify and reduce single points of failure across platforms.
  • Define and execute game days, chaos testing, and controlled failure scenarios to validate operational resilience.

Security & Governance

  • Ensure platforms are secure-by-design.
  • Work closely with Security and Architecture teams to implement:
  • Zero Trust principles
  • RBAC controls / Managed Identities
  • Network segmentation and Zone based architecture
  • Secrets management
  • Support compliance requirements and operational audits.
  • Help coordinate security updates, patches, maintenance routines, and upgrades of the underlying system across partners and vendors
  • Embed reliability, security, and compliance controls into build and release pipelines to support production readiness.

Continuous Improvement

  • Drive automation and reduction of manual operational tasks.
  • Improve deployment reliability and platform observability.
  • Contribute to architecture standards, runbooks, and operational documentation.
  • Partner with engineering, architecture, security, and service teams to define production readiness standards and reliability acceptance criteria.
Rate:
Not specified
Location:
Remote
IR35 Status:
Outside
Remote Status:
Remote
Industry:
IT
Seniority Level:
Not Specified

Take-Home Pay

Not Available

Visit calculators for additional details

Share job