All Jobs Vacancy

Platform Infrastructure SRE (Kubernetes / Cloud / IaC)

Posted 1 day ago by Bayside Solutions

Job Summary

We are looking for a strong Senior Platform Infrastructure SRE / Software Engineer to support the productionization and operation of large-scale Kubernetes-based platform services across multiple cloud environments. This role is best suited for someone with deep Kubernetes and infrastructure experience who can take capabilities developed by a platform engineering team and make them repeatable, scalable, observable, reliable, and production-ready across many environments.

Duties and Responsibilities

  • Productionize Kubernetes-based platform services developed by platform engineering teams.
  • Deploy and operate platform infrastructure across multiple cloud and production environments.
  • Build repeatable environment provisioning and deployment automation.
  • Develop reusable infrastructure templates, blueprints, and deployment patterns.
  • Provision Kubernetes clusters, cloud infrastructure, networking, and service dependencies.
  • Configure environment-specific infrastructure, connectivity, and platform services.
  • Deploy, validate, upgrade, and maintain platform services throughout their lifecycle.
  • Build monitoring, alerting, dashboards, logging, health checks, and operational controls.
  • Establish and validate production-readiness standards for new platform capabilities.
  • Troubleshoot complex failures across applications, Kubernetes, infrastructure, networking, and distributed systems.
  • Work closely with platform developers to understand application behavior and identify operational gaps before production rollout.
  • Implement and maintain Infrastructure as Code using technologies such as Crossplane, Terraform, Pulumi, or CloudFormation.
  • Build and maintain Helm-based Kubernetes packaging and deployment patterns.
  • Design safe rollout, rollback, upgrade, recovery, and lifecycle-management processes.
  • Validate platform capacity, availability, scalability, and reliability.
  • Configure and troubleshoot DNS, load balancing, VPC networking, routing, service connectivity, security policies, and certificates.
  • Improve operational automation and reduce manual environment-specific work.
  • Document architecture, deployment patterns, operational procedures, and troubleshooting guidance.
  • Participate in production support, incident response, and root-cause analysis as appropriate.
  • Independently own technical work and drive complex problems through resolution with limited supervision.

Requirements and Qualifications

  • Deep experience with Crossplane for infrastructure provisioning and platform automation.
  • Experience with Alibaba Cloud and its Kubernetes, networking, and infrastructure services.
  • Experience with AWS EKS and/or Google Cloud Platform.
  • Experience designing or operating multi-cloud platforms.
  • Experience with service mesh technologies and Kubernetes service networking.
  • Experience with distributed data technologies such as Apache Spark, Apache Flink, or Trino.
  • Experience implementing authentication, authorization, cloud security, and governance controls.
  • Experience building and maintaining CI/CD pipelines for Kubernetes-based platforms.
  • Experience designing highly available and resilient platform architectures.
  • Experience automating provisioning and lifecycle management across a large number of environments.
  • Strong production SRE, incident response, reliability engineering, and operational automation experience.

Preferred Qualifications

  • Deep experience with Crossplane for infrastructure provisioning and platform automation.
  • Experience with Alibaba Cloud and its Kubernetes, networking, and infrastructure services.
  • Experience with AWS EKS and/or Google Cloud Platform.
  • Experience designing or operating multi-cloud platforms.
  • Experience with service mesh technologies and Kubernetes service networking.
  • Experience with distributed data technologies such as Apache Spark, Apache Flink, or Trino.
  • Experience implementing authentication, authorization, cloud security, and governance controls.
  • Experience building and maintaining CI/CD pipelines for Kubernetes-based platforms.
  • Experience designing highly available and resilient platform architectures.
  • Experience automating provisioning and lifecycle management across a large number of environments.
  • Strong production SRE, incident response, reliability engineering, and operational automation experience.
Rate:
Not specified
Location:
Remote
IR35 Status:
Outside
Remote Status:
Remote
Industry:
IT
Seniority Level:
Senior

Take-Home Pay

Not Available

Visit calculators for additional details

Create a free account to view the take-home pay for this contract

Share job