New opportunity

AI Test Engineer - VLabs

Vialto Partners · Bengaluru, India

About this role

Vialto Labs seeks an AI Test Engineer to validate the performance, reliability and integrity of production AI solutions used in tax, immigration and internal operations. Working with the Programme Test Manager and engineering, product and delivery teams, the engineer will turn AI testing strategy into reusable evaluation frameworks and delivery-cycle quality checks for LLMs, OCR pipelines, document classification models and agentic workflows.

Responsibilities

  • Design executable scenarios, including adversarial, boundary and edge-case tests, to assess output structure, consistency, accuracy, hallucinations, misclassification and production readiness against defined thresholds.
  • Build Python evaluation frameworks, parameterized scripts, scoring and hallucination-detection mechanisms. Implement AI-as-Judge assessments with calibrated prompts and scoring, and integrate evaluations into CI/CD pipelines.
  • Monitor drift using fixed baseline datasets and scheduled re-evaluation; set thresholds for acceptable variation and identify regressions for release gating. Develop and update ground truth datasets with subject matter experts and define classification and extraction accuracy standards.
  • Test end-to-end agentic workflows for data integrity, error propagation and fallback behavior. Test AI pipeline APIs with Python and Postman/Newman, and verify persistence and integrity with SQL. Scale reusable evaluation patterns, support Responsible AI and privacy requirements, and communicate model-readiness risks and results to stakeholders.

Requirements

  • 7+ years in software testing, including 2–3 years focused on AI/ML-enabled systems in production. Experience designing AI evaluation strategies and pipelines, building ground truth datasets and drift detection systems, and testing multi-step agentic workflows.
  • Advanced Python; experience with LLM evaluation tools such as deepeval, RAGAS or promptfoo; output validation, grounding checks, statistical evaluation, OCR, VLM and document AI testing; API testing with requests or httpx and Postman/Newman; SQL; and familiarity with LangChain, LlamaIndex or similar frameworks. Independent execution, analytical rigor and communication in fast-moving environments are expected. A bachelor’s degree is required.

Preferred

Experience in regulated or compliance-driven environments, Azure AI Foundry or AWS Bedrock, and an advanced degree in Computer Science, Data Science or a related field. The role is based in the Bangalore office, with the possibility of hybrid work.

Skills for this role

PythonSQLdeepevalRAGASpromptfooPostmanNewmanrequestshttpxLangChainLlamaIndexAzure AI FoundryAWS BedrockCI/CDAPI testingLLM evaluationAI-as-JudgePrompt designHallucination detectionGrounding checksDrift detectionStatistical evaluationOCR testingVLM testingDocument classification testingAgentic workflow testingGround truth dataset developmentResponsible AIData privacy

Your skill match

Checking your profile…

YOUR NEXT STEP

Get interview-ready for this role

A focused preparation guide, built around this job’s responsibilities and requirements.

✦ AI-generated guide
Preparation suggestions, not the employer’s actual interview questions. Always check the original posting for current requirements.

Loading this role’s preparation guide…

KEEP EXPLORING

Similar AI jobs

Related skills and specializations in India. Matched to this role, not your profile.

Explore more jobs
Finding similar opportunities…