ROLE PURPOSE
Build and evolve secure, resilient cloud platforms and delivery capabilities that enable engineering teams to release software faster, safely and consistently.
This is a hands-on role spanning Kubernetes engineering, infrastructure automation, continuous delivery, observability and platform controls.
What you will do
- Design, build and operate enterprise Kubernetes platforms using GKE and OpenShift, including cluster lifecycle, networking, ingress and workload security.
- Create reusable platform services, golden paths and developer self-service capabilities through Backstage and Internal Developer Platform patterns.
- Build secure, repeatable cloud environments and reusable modules using Terraform and Infrastructure as Code principles.
- Design and maintain continuous delivery pipelines using Harness and Jenkins, with source control and collaboration through GitHub / Git.
- Implement end-to-end observability using Dynatrace, covering metrics, logs, traces, telemetry, dashboards, alerting and service health.
- Define SLI/SLO measures and automate operational runbooks, remediation and repeatable support activities to reduce manual effort.
- Engineer service-to-service connectivity and controls using Istio Service Mesh, including traffic management, ingress/egress and policy enforcement.
- Manage secrets, certificates and workload identity securely using HashiCorp Vault and automated certificate lifecycle practices.
- Embed platform controls, least-privilege access, secure configuration, resilience and auditability into engineering patterns.
- Partner with product, engineering, architecture, security and operations teams; coach colleagues and promote platform adoption and engineering standards.
Essential skills and experience
- Container platforms: Strong hands-on Kubernetes engineering; GKE preferred. OpenShift experience is required.
- Cloud & IaC: Google Cloud Platform experience preferred; strong Terraform and reusable infrastructure automation.
- Delivery engineering: Harness CI/CD and Jenkins; pipeline design, automation, release controls and deployment patterns.
- Developer platform: Backstage and Internal Developer Platform concepts, including self-service and golden paths.
- Observability & SRE: Dynatrace; metrics, telemetry, logs, traces, alerting, SLI/SLO practices and runbook automation.
- Security & controls: HashiCorp Vault, certificate management, IAM/least privilege, platform controls and secure configuration.
- Networking: Istio Service Mesh, ingress/egress, service connectivity, traffic management and network security concepts.
- Engineering practices: GitHub / Git, automation-first mindset, scripting, troubleshooting and production support.
Desirable
- Experience operating platforms within a large, regulated enterprise environment.
- Experience with multi-cloud or hybrid-cloud environments.
- Good working knowledge of Bash, Python or PowerShell automation.
- Understanding of reliability engineering, incident management, capacity, resilience and cost telemetry.
What success looks like
- Secure, supportable and repeatable platform services with clear engineering standards and controls.
- Faster, safer delivery through automation, reusable patterns and an improved developer experience.
- Improved service reliability through measurable SLOs, actionable observability and automated operations.