Overview
We're looking for an experienced AI Evaluation Engineer to join a specialist Public Sector Technology team.
This hands-on role focuses on building evaluation frameworks, tooling and harnesses for AI systems, particularly LLM and agentic AI solutions.
The role requires ownership, experimentation and rapid delivery.
Key Responsibilities
- Design evaluation frameworks and tooling
- Develop evaluation harnesses
- Evaluate model and agentic AI systems
- Define metrics and methodologies
- Build repeatable evaluation processes
- Investigate AI failures
- Work directly with government teams
- Prototype and test solutions
- Present findings and recommendations
- Contribute to strategic direction.
What We Are Looking For
- Strong software engineering experience
- Experience building and evaluating AI systems
- Understanding of LLMs and agentic AI
- Experience with AI evaluation frameworks
- Ability to code independently
- Strong problem-solving and communication skills
- Comfortable working autonomously.
Technical Experience
- AI/LLM Evaluation
- RAG Evaluation
- Ragas
- Agentic AI Evaluation
- Evaluation Harnesses
- LLM Testing and Benchmarking
- Prompt Engineering
- Python Development
- Model and Agent Integration.
Ideal Background
- AI Evaluation Engineer
- Applied AI Engineer
- AI Engineer
- LLM Engineer
- Harness Engineer
- Prompt Engineer or Software Engineer specialising in AI.
Team
Small Public Sector Technology team of approximately 2 to 6 people within a wider organisation.
What Will Make Someone Stand Out
Experience building evaluation tooling, evaluating LLMs, using Ragas, identifying AI failure modes, combining engineering with strategic thinking, and thriving in fast-moving environments.
Summary
Hands-on AI Evaluation Engineer role supporting government AI initiatives through evaluation, testing, benchmarking and framework development.