Senior Site Reliability Engineer (SRE)
Remote from ONLY LATAM - (Colombia, Mexico, Costa Rica, Brazil, Peru)
We are looking for a Senior Site Reliability Engineer to build, operate, and improve highly scalable, resilient, and secure cloud platforms supporting critical enterprise applications. This is a hands-on technical leadership role focused on AWS/Google Cloud Platform, Kubernetes, observability, automation, and reliability engineering.
Key Responsibilities
- Design and implement reliability strategies for distributed systems across AWS and Google Cloud Platform.
- Define and monitor SLIs, SLOs, error budgets, and reliability metrics.
- Build and enhance monitoring, logging, tracing, alerting, and observability solutions.
- Lead incident response, root cause analysis, and postmortems.
- Improve system performance, scalability, resiliency, and operational readiness.
- Automate operational processes and reduce manual toil.
- Guide engineering teams on reliability architecture, capacity planning, and non-functional requirements.
Required Skills
- 7+ years in SRE, Cloud Engineering, DevOps, or Platform Engineering.
- Strong production experience with AWS and/or Google Cloud Platform.
- Hands-on expertise with Kubernetes (EKS/GKE).
- Strong understanding of SRE principles, SLIs, SLOs, error budgets, and incident management.
- Experience with observability tools such as Prometheus, Grafana, CloudWatch, Cloud Monitoring, Datadog, or Splunk.
- Strong Terraform/IaC experience.
- Proficiency in Python, Bash, or similar scripting languages.
- Strong understanding of cloud networking, distributed systems, security, and performance optimization.
Preferred
- Large-scale cloud migration or modernization experience.
- Chaos engineering/resilience testing experience.
- Istio/service mesh knowledge.
- AWS/Google Cloud Platform certifications.
- Experience in Agile, DevOps, or DevSecOps environments.