Project Description
- We are looking for a senior AI Platform / SRE Lead to drive AI-enabled operational practices across distributed engineering teams.
- The role will focus on reliability, automation, observability, production support, and adoption of AI within DevOps and SRE processes.
- Strong experience in SRE, DevOps, production operations, automation, and Forward Deployed Engineering (FDE), with practical knowledge of applying AI/GenAI to operational and client-specific use cases.
- Expertise in observability, monitoring, incident management, troubleshooting, CI/CD, operational reliability, and deploying solutions directly into complex enterprise environments is required.
- Must be capable of working closely with client and engineering teams to translate business and operational needs into scalable technical solutions, establishing governance and guardrails, and serving as a senior technical escalation and advisory lead across multiple engineering pods.
- Experience implementing AI-assisted incident response, intelligent monitoring, automated troubleshooting, AI-enabled DevOps solutions, and FDE-led solution deployment and integration.
- Experience with cloud platforms, Infrastructure as Code, enterprise AI tools, customer-facing technical delivery, and large-scale healthcare or regulated environments is preferred.
Responsibilities
- Lead AI-enabled operations, reliability, automation, observability, and production support initiatives.
- Identify and guide AI use cases for incident management, monitoring, troubleshooting, and operational automation.
- Define operational standards, governance, and guardrails.
- Drive adoption of AI-enabled DevOps and SRE practices.
- Improve operational efficiency and reliability through automation and AI.
- Serve as a technical advisor and escalation point across multiple delivery pods.
Mandatory Skills
- Site Reliability Engineering (SRE)
- DevOps
- Generative AI / AI Engineering
- Observability & Monitoring
- Incident Management
- Operational Automation
- CI/CD