About this role
EY GDS Assurance Digital seeks a full-time Agentic AI Evaluation Engineer to build and standardize evaluations for GenAI, RAG-based and agentic solutions before deployment. The role supports evidence-based go/no-go decisions on system quality, safety, reliability and robustness. The stated location is Kolkata, with “Anywhere in Country” listed as another location; salary is described as competitive.
Responsibilities
- Work with product teams, architects, risk and compliance teams, and assurance leadership to turn business use cases into evaluation plans with datasets, metrics, acceptance thresholds and reporting requirements.
- Define dataset coverage for core journeys, edge cases, adversarial prompts, and bias and fairness scenarios. Create reusable templates, test libraries, scoring rubrics and guidance on sample sufficiency and statistical coverage.
- Build reproducible evaluation pipelines and automated harnesses, preferably in Python. Assess answer quality, RAG grounding and citations, agent tool use and goal completion, and operational measures including latency, cost, throughput and recovery. Calibrate LLM-as-judge scoring against human evaluation.
- Conduct red teaming aligned with the OWASP Top 10 for LLM Applications, including prompt injection, tool hijacking, data disclosure, retrieval poisoning and unauthorized actions. Recommend controls such as validation, filtering, least-privilege tool access, sandboxing and monitoring. Integrate regression evaluations into CI and release gates, and produce auditable reports explaining results, limitations, residual risks and go/no-go recommendations.
Requirements
- A bachelor’s or master’s degree in data science, statistics, engineering, operational research or a related field, with a strong focus on modern data architectures, processes and environments. The posting requests 4–7+ years of relevant experience in AI or evaluation engineering, applied research, security testing or evaluation-harness design.
- Strong hands-on Python, understanding of GenAI architectures and enterprise LLM risks, metric and experiment design, evaluation tooling, tracing, basic DevOps, written communication, stakeholder influence and demonstrated project management experience.
Preferred qualifications include security and red-teaming expertise; experience in assurance, finance or regulatory environments; familiarity with responsible AI frameworks; multilingual or domain-heavy assistant evaluation; and hands-on experience with Azure OpenAI, AI Search, Function Apps or App Insights.