About this role
Srijan Technologies, a Material company, seeks a Lead Agentic AI Engineer to design and deliver production-grade agentic systems for client engagements. The role combines technical solution design, multi-agent architecture, production AI engineering and team leadership. The position is listed in Gurgaon/Gurugram, India, and is full-time.
Responsibilities
- Architect orchestrator/sub-agent systems, state machines and tool registries using agent frameworks. Build advanced retrieval pipelines incorporating hybrid search, re-ranking, query expansion, multi-hop reasoning and knowledge graphs.
- Adapt foundation models through quantization, PEFT/LoRA fine-tuning and prompt optimization. Implement output validation, factual grounding, prompt-injection defenses, content filtering and hallucination controls.
- Deploy AI services on AWS or Azure using Docker, Kubernetes and CI/CD. Monitor accuracy, hallucinations, latency, cost and drift; establish LLM evaluation suites aligned with client KPIs. Optimize inference cost and throughput through token budgeting, prompt compression, KV-cache management, model routing, streaming, batching and model-serving tools. Maintain automated unit, contract and model-quality tests.
- Lead client discovery and technical proposals, present designs and prototypes, explain trade-offs to technical and business stakeholders, mentor engineers, review architecture decisions and develop reusable AI patterns.
Required qualifications
- 5–10 years in software engineering or data science, including at least 3 years in applied generative AI or LLM engineering in a services or consulting context. Experience building production agents with LangGraph, CrewAI, AutoGen or Semantic Kernel; deep RAG, vector-database and embedding-pipeline expertise; and knowledge of GPT-4o, Claude, Gemini and open-source models.
- Hands-on AWS or Azure AI services, Docker, Kubernetes, CI/CD, Python, FastAPI and SQL. Proven ability to take LLM systems from prototype to production, including deployment, observability, evaluations, guardrails and ongoing model health. Rigorous software design and testing practices. B.Tech, B.E. or M.Tech in Computer Science or a related discipline.
Preferred qualifications
PEFT/LoRA fine-tuning, high-throughput serving with vLLM, DeepSpeed or Triton Inference Server, GraphRAG or ontology-based retrieval, vision-language models or multimodal agents, and AWS Solutions Architect, Azure AI Engineer Associate or equivalent certification.