6-month contract £102,000-£114,000 per annum Remote within the UK We are hiring a Software Engineer to join the run time and reliability organisation of a global technology company, working on the systems that keep a large-scale, high-traffic application stable and operational across VR, mobile and PC environments.
This is a highly autonomous engineering role combining AI-driven automation, back-end reliability, cloud infrastructure and production engineering.
You will solve complex operational problems by building automation and autonomous systems that can identify issues, diagnose failures and increasingly remediate them without human intervention.
This is an opportunity to work on a live product operating at significant scale.
The environment requires someone comfortable with ambiguity.
You'll have significant autonomy and will be expected to take ownership of problems, communicate quickly when blockers arise and drive issues through to resolution.
What You'll Do
- You will build and operate automation responsible for keeping a large live application healthy in production.
- Your work will include:
- Building AI-driven and self-healing systems that monitor production metrics and autonomously respond to problems
- Maintaining and improving AI-assisted developer tooling capable of generating and repairing code
- Developing tooling that automatically detects broken builds, identifies root-cause changes and recommends or executes fixes
- Operating and improving back-end services supporting a large-scale production environment
- Monitoring cloud rendering and deployment pipelines and maintaining compatibility across different platforms
- Monitoring production quality, performance, crashes, regressions and outages
- Improving CI/CD, build, release and cloud deployment pipelines
- Automating repetitive operational work and significantly reducing manual on-call workload
- Completing infrastructure and dependency migrations while protecting downstream CI/CD systems
- Diagnosing complex production behaviour under load and taking ownership of issues through to resolution
What We're Looking For
You are a software engineer who is comfortable owning complex production problems with a high degree of autonomy.
You'll ideally bring:
8+ years of professional software engineering experience or equivalent experience
Strong experience building autonomous, AI-assisted or self-healing engineering systems
Experience building or operating AI-assisted developer tooling or agents that generate, modify or repair code
Deep experience operating back-end and cloud services in production
Strong CI/CD, build, release and deployment pipeline experience
Experience with production monitoring, incident response, crash triage and reliability engineering
Experience building automation that reduces operational and on-call workload
Experience diagnosing failed builds and identifying root-cause changes
Experience managing infrastructure or dependency migrations in complex production environments
Strong ownership, problem-solving and communication skills
Experience in any of the following would be highly valuable: Site Reliability Engineering / Production Engineering Cloud application or game streaming Remote rendering Large-scale distributed backend systems Live-service applications CDN or asset-delivery pipelines Capacity and latency engineering Session orchestration Large monorepo build systems Automated remediation and self-healing infrastructure
Please note that this is a contract role for 6 months . Please apply directly to be considered.