About the role
Role/Job title :AI Evaluation Engineer
Work Location :London
Type work : Contract
Mode of work : Hybrid
Responsibilities:
* Define and implement end-to-end evaluation strategy for generative AI conversational systems (LLMs, RAG pipelines, agents).
* Establish evaluation metrics
* Test Dataset & Benchmarking
* Benchmark models (GPT variants, open-source LLMs) and prompt strategies
* Implement automated evaluation pipelines
* Conduct human-in-the-loop evaluation for qualitative validation.
* Prompt & Response Quality Optimization
* RAG & Knowledge Grounding Validation
* Safety, Risk & Compliance Testing
Your Profile
* Strong understanding of LLMs, RAG architecture, prompt engineering
* Python – data analysis, evaluation pipelines
* Prompt evaluation tools (PromptTools, DeepEval, etc.)
Experience with:
- Evaluation Framework Design
- Establish Evaluation metrics & NLP quality assessment
- Test Dataset and benchmarking
- LLM output evaluation and scoring
- Safety, risk and compliance testing
- Python, data analysis, experimentation frameworks
- Tooling and automation
Familiarity with:
- Responsible AI principles
- Customer support / retail conversational flows
- Analytical mindset with ability to translate model behaviour into actionable improvements
About this listing
Screened by Joboru
This role passed our automated spam and quality filters and was active in our feed when last checked. Joboru is an aggregator — here is how we screen listings. If anything looks off, tell us.
Similar jobs you may like
Team Leader
1 day agoCard Factory
Sales Assistant
1 day agoCard Factory
Digital Print Operator
1 day agoGet-Recruited (UK) Ltd
Clinic Manager, Luxury Beauty
1 day agoOffice Angels
Fabric Technician
1 day agoIntegral UK Ltd
Garment Technologist
1 day agoPhoenix Recruitment Consultancy Ltd
Engineering Manager - hybrid
1 day agoBlue Light Card
Deputy Store Manager (Hiring Immediately)
1 day agoLidl
Oribe Retail Education Specialist, UK & Ireland
1 day agoKao