What You'll Do
- Build and maintain automation that keeps the application healthy in production — release pipelines, build health checks, crash triage, and incident detection
- Maintain and improve an AI-assisted code repair system that autonomously creates and lands fix diffs
- Develop tooling that automatically identifies broken builds, pinpoints the root-cause change, and recommends or executes fixes
- Monitor weekly deployments of the cloud rendering system, ensuring performance and compatibility between VR and mobile/PC users after every release
- Monitor production quality metrics and respond to regressions and outages
- Drive down recurring on-call and operational work — targeting an 80–90% reduction through automation
- Complete infrastructure and dependency migrations to keep CI/CD pipelines functional as upstream systems are retired
What We're Looking For
- 8+ years of professional software engineering experience, or equivalent
- Proven experience building and operating CI/CD, build, release, and cloud deployment pipelines at scale
- Experience operating cloud services and server-side fleets in production, including reliability, capacity, and latency management
- Experience building or operating AI-assisted developer tooling or agents that generate or repair code
- Experience building tooling that detects broken builds and traces failures back to their root-cause change
- Experience with production monitoring, crash triage, and incident response for a large-scale, multi-surface application
- A track record of reducing operational and on-call load through automation
- Experience completing infrastructure or dependency migrations without breaking downstream CI/CD
Top non-negotiables:
- Experience building autonomous, self-healing AI systems that monitor metrics and take corrective action independently to keep performance within required thresholds
- Experience deploying and maintaining backend services — specifically managing and monitoring deployments of a system like cloud/remote rendering, and verifying cross-platform compatibility (e.g., VR and mobile) after release
Nice to Have
- Experience with cloud game or application streaming, or remote rendering
- Experience with asset delivery or CDN pipelines at scale
- Experience with capacity, latency, or session-orchestration monitoring for streamed workloads
- Experience operating live-service or large-scale production applications (Live Ops)
- Familiarity with large monorepo build systems and dependency management
- Experience designing self-healing or auto-remediation systems