AI Engineer (AI agents focused) | Agentic AI, LLMs, Evals, Langchain | Production Code Generation, Model Routing | Gaming Industry, Start-up | Rate £500-550pd | 6 months | Outside | Remote
The Company
We're partnering with a well-funded, early-stage AI start-up tackling a huge global industry with an ambitious new product.
They're building agentic systems that turn ideas into complex, production-ready software – not another LLM wrapper.
The technology is already working.
The challenge now is making it more reliable, intelligent and scalable, solving difficult problems across agents, evals, context, performance and inference economics.
They're looking for a Founding AI Engineer to own it and help get their MVP released, fast.
The Role
AI Engineer (AI agents focused) | Agentic AI, LLMs, Evals, Langchain | Production Code Generation, Model Routing | Gaming Industry, Start-up | Rate £500-550pd | 6 months | Outside | Remote
You'll own the agent system responsible for turning high-level requirements into working, production-quality software.
Working directly with the founders and engineering team, you'll design agent loops, build production evals, manage context, route between models and optimise for quality, latency and cost.
The challenge isn't getting an LLM to generate code.
It's building an autonomous system that generates software that actually works.
What You'll Be Doing
- Own the agent pipeline from specification to production
- Build workflows across planning, code generation, review, testing and release
- Build evals and regression tooling to measure output quality and catch failures
- Develop model routing based on quality, latency and cost
- Build retrieval and context systems
- Improve or replace third-party agent infrastructure where needed
- Experiment with new models, fine-tuning and post-training
- You'll own one of the hardest and most important technical problems in the business.
Tech & AI Stack
- TypeScript
- LLMs
- Tool-Using Agents
- Evals
- Model Routing
- Retrieval & Context
- Fine-Tuning
- Code Generation
- Langchain
- Langgraph
- Deep TypeScript experience isn't essential – strong systems engineering experience in another language is absolutely fine.
What They're Looking For
- Production experience building LLM/agentic systems
- Multi-step pipelines, tool-using agents or autonomous systems
- Agents that produce output that actually has to run and work , not just generate text
- Strong experience with evals, benchmarks and regression testing
- Strong systems engineering fundamentals
- Understanding of the trade-offs between model quality, latency and cost
- Someone who follows the AI field closely and actively experiments with new models and techniques
- Nice to have: model routing, fine-tuning/post-training, retrieval systems, proprietary agent runtimes, code-generation agents or developer tooling.
- This is particularly suited to engineers who have gone beyond chatbots, RAG and simple LLM integrations and built agentic systems that need to perform reliably in production.
Interview Focus
- Expect to go deep on three areas:
- Production Agents – What have you built where the agent's output actually had to run, and how did you know it worked?
- Evals – What's in your eval suite, and what has it stopped from shipping?
- Performance & Cost – What was the slowest or most expensive part of your agent loop, and how did you improve it?
Why Join
- Founding-level ownership over a core AI system
- Solve genuinely difficult problems across agents, evals and model infrastructure
- Build proprietary AI infrastructure rather than another LLM wrapper
- Work directly with the founders and influence technical direction
- If you've built agentic systems where the output actually has to work – and want to tackle problems at the edge of what's currently possible with LLMs – we'd love to hear from you.