Overview
The team is looking for an experienced, hands-on developer who can create their own base code, set it up and work it out.
The ideal candidate:
- proven experience in quality engineering, testing or evaluation of data-driven and AI-powered products.
- strong understanding of machine learning systems, large language models (LLMs) and AI evaluation methodologies for production.
- experience defining performance metrics, KPIs, benchmarks and acceptance criteria for AI products.
- ability to translate business risks and policy requirements into measurable technical evaluation frameworks.
- experience developing automated testing and evaluation pipelines using Python or similar technologies.
- knowledge of AI safety, model robustness, bias, fairness and responsible AI principles.
- strong analytical and problem-solving skills with the ability to identify trends, risks and performance issues.
- experience working with cross-functional teams in agile product delivery environments.
- ability to communicate complex technical findings clearly to both technical and non-technical audiences.
- experience producing technical reports, dashboards and recommendations for senior stakeholders.
- familiarity with cloud-based AI platforms and modern software development practices.
- experience working in government, public sector or regulated environments is desirable.
NB:
The successful candidate may need to undergo SC while in post, so must be eligible/willing to attain. However the successful candidate can start on BPSS.