Job role: Site Reliability Engineer (SRE)
Exp: 10+
description
- Experience in SRE, platform engineering, infrastructure, distributed systems, or cloud operations.
- Demonstrated technical leadership across multi-team infrastructure or platform programs.
- Strong software engineering skills in Python, Go, Ruby, or similar languages, plus experience with infrastructure-as-code and production automation.
- Deep experience operating large-scale cloud or hybrid-cloud systems with strong reliability, security, and observability requirements.
- Strong understanding of CI/CD, deployment orchestration, Kubernetes, AWS, and production incident response.
- Ability to translate compliance and security requirements into practical engineering systems and operating models.
- Excellent written communication, architecture review, mentoring, and cross-functional leadership skills.
- Experience building or operating FedRAMP High, DoD IL5, GovCloud, FIPS, or other regulated environments.
- Experience with platform APIs, progressive delivery, CI components, continuous delivery, artifact promotion, or deployment control planes.
- Experience leading automation of previously manual operational workflows.