All Jobs Vacancy

Sr. Staff Site Reliability/SRE _Remote_$70 per hour on W2_10+ yrs exp

Posted 1 week ago by Xoriant Corporation

Job Description

  • Production on-call experience in a real rotation, with incident command and blameless postmortem practice.
  • Production Kubernetes and container experience (Docker), with cloud-native infrastructure patterns.
  • Hands-on production ownership on at least one major cloud (AWS, Google Cloud Platform, or Azure).
  • Terraform or OpenTofu proficiency.
  • Observability depth with Prometheus, Grafana, or equivalent for metrics, logging, and alerting, including dashboard and alert design.
  • Strong automation skills in Python, Bash, or Go.
  • Networking fundamentals: VPCs, load balancers, DNS, firewalls, cross-cloud connectivity.
  • CI/CD experience with GitHub Actions, GitLab CI, Jenkins, or ArgoCD.
  • Proven ability to troubleshoot complex distributed systems, largely self-directed.
  • Preferred Qualifications
  • GPU infrastructure and AI/ML workloads: Ray, Kubeflow, ML flow, or similar.
  • NVIDIA GPU orchestration: A100/H100 configuration, driver and CUDA runtime management.
  • Distributed training networking: RDMA, InfiniBand, EFA, NCCL.
  • Distributed tracing and Open Telemetry instrumentation across services.
  • Progressive delivery: canary and blue/green rollouts with automated rollback.
  • Chaos or fault-injection testing, game days, and disaster-recovery drills.
  • Multi-cloud networking, unified storage abstractions, and disaster recovery.
  • FinOps and cost optimization: Spot, Reserved Instances, Savings Plans.
  • Establishing an SRE function where one did not previously exist.
Rate:
Not specified
Location:
Remote
IR35 Status:
Outside
Remote Status:
Remote
Industry:
IT
Seniority Level:
Senior

Take-Home Pay

Not Available

Visit calculators for additional details

Create a free account to view the take-home pay for this contract

Share job