Job description
A global technology organisation is seeking an Automation Engineer to build AI-driven automation and reliability tooling across a large, multi-platform product.
The successful candidate will design systems that detect issues early, stabilise pipelines, and improve overall runtime quality.
This role focuses on creating autonomous workflows that reduce manual effort and strengthen release confidence.
What you'd actually be doing:
Build and run the automation that keeps the product alive in production and reduces manual on-call work.
Keep release pipelines, build health, and cloud rendering/deployment pipelines healthy.
Maintain and improve an AI-assisted code repair system that autonomously creates and lands fixes.
Build tooling that auto-detects broken builds, finds the root-cause change, and fixes or recommends a fix.
Monitor production metrics, triage crashes, and respond to regressions/outages.
Goal: cut recurring operational work by 80-90%.
Handle infra migrations so CI/CD doesn't break when upstream dependencies are retired.
Requirements - they want someone who has:
- 8+ years SWE experience
- Built and operated CI/CD, build, release, and cloud deployment pipelines at scale
- Operated cloud services/server fleets in prod (reliability, capacity, latency)
- Built/operated AI dev tooling/agents that generate or repair code
- Built tooling for broken-build detection + root-cause tracing
- Done prod monitoring, crash triage, and incident response for a large, multi-platform app
- Proven track record of automating away on-call toil
- Done infra migrations without breaking downstream CI/CD
Nice to have:
- Cloud game/app streaming
- Remote rendering CDN/asset delivery at scale
- Live Ops
- Monorepo build systems
- Self-healing / auto-remediation systems.