About this role
XenonStack seeks a Machine Learning Engineer, Model Evaluation to assess the accuracy, safety, robustness and trustworthiness of large language models and agentic AI systems in enterprise workflows. The full-time role is listed in Mohali, India, with a salary of 8-15 LPA.
Responsibilities
- Design LLM evaluation pipelines and automated benchmarks for enterprise-relevant tasks. Stress-test models through adversarial and edge-case evaluations, including multi-turn interactions; measure latency, consistency and error recovery.
- Define metrics for factual accuracy, hallucinations, toxicity and compliance alignment. Establish monitoring for drift, anomalies and performance regressions, and maintain a central repository of test cases, benchmarks and results.
- Work with ML engineers, product managers, domain experts and Responsible AI teams to align testing with business objectives and ethical, explainable, compliant practices. Feed findings into fine-tuning, RLHF/RLAIF and model selection.
- Research evaluation techniques and explore agentic test harnesses and synthetic data generation.
Required qualifications
The qualifications specify 3–6 years in AI/ML, NLP or applied model evaluation, while the job-information section lists 2–4 years of work experience. Candidates need an understanding of LLM architectures, prompt engineering and failure modes; hands-on experience with evaluation frameworks such as Ragas, OpenAI Evals and DeepEval; and Python proficiency with libraries such as LangChain, LangGraph, LlamaIndex and Hugging Face. Experience with vector databases, RAG pipelines and knowledge graph integration, plus familiarity with bias/fairness testing and Responsible AI frameworks, is required.
Preferred qualifications
RLHF, RLAIF and reward modeling; agentic evaluation, including multi-agent stress testing and synthetic user simulators; safety and compliance requirements for BFSI, GRC or SOC use cases; and contributions to open-source evaluation libraries or research papers. The role involves enterprise-scale evaluation across domains including BFSI, healthcare, telecom and GRC.