Job Description
- Production on-call experience in a real rotation, with incident command and blameless postmortem practice.
- Production Kubernetes and container experience (Docker), with cloud-native infrastructure patterns.
- Hands-on production ownership on at least one major cloud (AWS, Google Cloud Platform, or Azure).
- Terraform or OpenTofu proficiency.
- Observability depth with Prometheus, Grafana, or equivalent for metrics, logging, and alerting, including dashboard and alert design.
- Strong automation skills in Python, Bash, or Go.
- Networking fundamentals: VPCs, load balancers, DNS, firewalls, cross-cloud connectivity.
- CI/CD experience with GitHub Actions, GitLab CI, Jenkins, or ArgoCD.
- Proven ability to troubleshoot complex distributed systems, largely self-directed.
- Preferred Qualifications
- GPU infrastructure and AI/ML workloads: Ray, Kubeflow, ML flow, or similar.
- NVIDIA GPU orchestration: A100/H100 configuration, driver and CUDA runtime management.
- Distributed training networking: RDMA, InfiniBand, EFA, NCCL.
- Distributed tracing and Open Telemetry instrumentation across services.
- Progressive delivery: canary and blue/green rollouts with automated rollback.
- Chaos or fault-injection testing, game days, and disaster-recovery drills.
- Multi-cloud networking, unified storage abstractions, and disaster recovery.
- FinOps and cost optimization: Spot, Reserved Instances, Savings Plans.
- Establishing an SRE function where one did not previously exist.