AI Evaluation Engineer

ScreenedHybridFull Time
London
£430 - £500/day
Posted 1 week ago
Apply Now

About the role

Role/Job title :AI Evaluation Engineer Work Location :London Type work : Contract Mode of work : Hybrid Responsibilities: * Define and implement end-to-end evaluation strategy for generative AI conversational systems (LLMs, RAG pipelines, agents). * Establish evaluation metrics * Test Dataset & Benchmarking * Benchmark models (GPT variants, open-source LLMs) and prompt strategies * Implement automated evaluation pipelines * Conduct human-in-the-loop evaluation for qualitative validation. * Prompt & Response Quality Optimization * RAG & Knowledge Grounding Validation * Safety, Risk & Compliance Testing Your Profile * Strong understanding of LLMs, RAG architecture, prompt engineering * Python – data analysis, evaluation pipelines * Prompt evaluation tools (PromptTools, DeepEval, etc.) Experience with: - Evaluation Framework Design - Establish Evaluation metrics & NLP quality assessment - Test Dataset and benchmarking - LLM output evaluation and scoring - Safety, risk and compliance testing - Python, data analysis, experimentation frameworks - Tooling and automation Familiarity with: - Responsible AI principles - Customer support / retail conversational flows - Analytical mindset with ability to translate model behaviour into actionable improvements

About this listing

Screened by Joboru

This role passed our automated spam and quality filters and was active in our feed when last checked. Joboru is an aggregator — here is how we screen listings. If anything looks off, tell us.