New opportunity

Agentic AI Evaluation Engineer

EY · Kolkata, India

About this role

EY GDS Assurance Digital seeks a full-time Agentic AI Evaluation Engineer to build and standardize evaluations for GenAI, RAG-based and agentic solutions before deployment. The role supports evidence-based go/no-go decisions on system quality, safety, reliability and robustness. The stated location is Kolkata, with “Anywhere in Country” listed as another location; salary is described as competitive.

Responsibilities

  • Work with product teams, architects, risk and compliance teams, and assurance leadership to turn business use cases into evaluation plans with datasets, metrics, acceptance thresholds and reporting requirements.
  • Define dataset coverage for core journeys, edge cases, adversarial prompts, and bias and fairness scenarios. Create reusable templates, test libraries, scoring rubrics and guidance on sample sufficiency and statistical coverage.
  • Build reproducible evaluation pipelines and automated harnesses, preferably in Python. Assess answer quality, RAG grounding and citations, agent tool use and goal completion, and operational measures including latency, cost, throughput and recovery. Calibrate LLM-as-judge scoring against human evaluation.
  • Conduct red teaming aligned with the OWASP Top 10 for LLM Applications, including prompt injection, tool hijacking, data disclosure, retrieval poisoning and unauthorized actions. Recommend controls such as validation, filtering, least-privilege tool access, sandboxing and monitoring. Integrate regression evaluations into CI and release gates, and produce auditable reports explaining results, limitations, residual risks and go/no-go recommendations.

Requirements

  • A bachelor’s or master’s degree in data science, statistics, engineering, operational research or a related field, with a strong focus on modern data architectures, processes and environments. The posting requests 4–7+ years of relevant experience in AI or evaluation engineering, applied research, security testing or evaluation-harness design.
  • Strong hands-on Python, understanding of GenAI architectures and enterprise LLM risks, metric and experiment design, evaluation tooling, tracing, basic DevOps, written communication, stakeholder influence and demonstrated project management experience.

Preferred qualifications include security and red-teaming expertise; experience in assurance, finance or regulatory environments; familiarity with responsible AI frameworks; multilingual or domain-heavy assistant evaluation; and hands-on experience with Azure OpenAI, AI Search, Function Apps or App Insights.

Skills for this role

PythonGenAI evaluationAgentic AI evaluationRAG evaluationEvaluation harnessesEvaluation metricsDataset designStatistical samplingLLM-as-judgeHuman evaluationRed teamingAdversarial testingPrompt injection testingOWASP Top 10 for LLM ApplicationsResponsible AILLM securityRAGEmbeddingsVector searchTool callingMulti-agent systemsCI/CDGitDockerRAGASDeepEvalLangSmithPhoenixArizeOpenTelemetry

Your skill match

Checking your profile…

YOUR NEXT STEP

Get interview-ready for this role

A focused preparation guide, built around this job’s responsibilities and requirements.

✦ AI-generated guide
Preparation suggestions, not the employer’s actual interview questions. Always check the original posting for current requirements.

Loading this role’s preparation guide…

KEEP EXPLORING

Similar AI jobs

Related skills and specializations in India. Matched to this role, not your profile.

Explore more jobs
Finding similar opportunities…