New opportunity

Machine Learning Engineer, Model Evaluation

Xenonstack Private Limited · Mohali, India

About this role

XenonStack seeks a Machine Learning Engineer, Model Evaluation to assess the accuracy, safety, robustness and trustworthiness of large language models and agentic AI systems in enterprise workflows. The full-time role is listed in Mohali, India, with a salary of 8-15 LPA.

Responsibilities

  • Design LLM evaluation pipelines and automated benchmarks for enterprise-relevant tasks. Stress-test models through adversarial and edge-case evaluations, including multi-turn interactions; measure latency, consistency and error recovery.
  • Define metrics for factual accuracy, hallucinations, toxicity and compliance alignment. Establish monitoring for drift, anomalies and performance regressions, and maintain a central repository of test cases, benchmarks and results.
  • Work with ML engineers, product managers, domain experts and Responsible AI teams to align testing with business objectives and ethical, explainable, compliant practices. Feed findings into fine-tuning, RLHF/RLAIF and model selection.
  • Research evaluation techniques and explore agentic test harnesses and synthetic data generation.

Required qualifications

The qualifications specify 3–6 years in AI/ML, NLP or applied model evaluation, while the job-information section lists 2–4 years of work experience. Candidates need an understanding of LLM architectures, prompt engineering and failure modes; hands-on experience with evaluation frameworks such as Ragas, OpenAI Evals and DeepEval; and Python proficiency with libraries such as LangChain, LangGraph, LlamaIndex and Hugging Face. Experience with vector databases, RAG pipelines and knowledge graph integration, plus familiarity with bias/fairness testing and Responsible AI frameworks, is required.

Preferred qualifications

RLHF, RLAIF and reward modeling; agentic evaluation, including multi-agent stress testing and synthetic user simulators; safety and compliance requirements for BFSI, GRC or SOC use cases; and contributions to open-source evaluation libraries or research papers. The role involves enterprise-scale evaluation across domains including BFSI, healthcare, telecom and GRC.

Skills for this role

LLM evaluationModel benchmarkingReliability testingAdversarial testingBias and fairness testingPythonPrompt engineeringRagasOpenAI EvalsDeepEvalLangChainLangGraphLlamaIndexHugging FaceVector databasesRAGKnowledge graphsResponsible AIRLHFRLAIFReward modelingSynthetic data generationPerformance monitoring

Your skill match

Checking your profile…

YOUR NEXT STEP

Get interview-ready for this role

A focused preparation guide, built around this job’s responsibilities and requirements.

✦ AI-generated guide
Preparation suggestions, not the employer’s actual interview questions. Always check the original posting for current requirements.

Loading this role’s preparation guide…

KEEP EXPLORING

Similar AI jobs

Related skills and specializations in India. Matched to this role, not your profile.

Explore more jobs
Finding similar opportunities…