About this role
The Responsible AI / Eval Specialist designs, implements and operates evaluation frameworks that determine whether AI use cases can pass governance stage-gates and enter production. Based in Gurgaon, Haryana, India, the role supports McCain’s Responsible AI standards across solutions built on Azure, Anthropic foundation models, Distyl technologies, Databricks, SAP, Salesforce and other enterprise platforms.
Responsibilities
- Design and run evaluations of accuracy, reliability, hallucination rates, calibration, fairness, robustness and overall quality. Build and maintain evaluation harnesses, benchmark suites and golden datasets; deliver actionable findings for stage-gate and risk-classification decisions.
- Test bias and fairness across protected attributes and relevant user groups, quantify concerns and recommend remediation. Lead red-team and adversarial exercises covering vulnerabilities, failure modes, prompt injection and jailbreaks, working with Enterprise Security on AI resilience.
- Embed evaluation capabilities in development workflows with AI Platform & Engineering and Knowledge Engineering. Maintain model cards, evaluation reports, decision logs and governance documentation; contribute to Responsible AI standards and playbooks, coach engineers, and share reusable evaluation practices. Collaborate with Forward Deployment Engineering, the Use Case Governance & Operations Lead and business stakeholders.
Requirements
- Bachelor’s degree in Computer Science, Statistics, Engineering or a related field, and 10+ years of experience in AI/ML evaluation, model validation, model risk management or a related discipline.
- Hands-on knowledge of AI/ML evaluation methods and metrics; experience evaluating LLMs and AI agents with HELM, Eleuther LM Evaluation Harness or equivalent frameworks; strong Python skills and familiarity with machine-learning evaluation and testing tools.
- Experience with Databricks or a comparable enterprise data platform; familiarity with Responsible AI and governance frameworks such as NIST AI RMF, the EU AI Act or ISO/IEC 42001. Strong analytical, problem-solving and communication skills are required to explain findings to technical and non-technical stakeholders.
Preferred
An advanced degree and experience conducting bias testing, red-team exercises, adversarial testing and AI risk assessments.
The position is regular full-time, with 40 hours per week stated. The posting says most office-based roles follow a hybrid model with two remote days weekly, but notes that exceptions depend on the role and location; candidates are advised to confirm the arrangement with a recruiter.