About this role
Xenonstack Private Limited seeks a Site Reliability Engineer, AI Platform to build end-to-end observability for AI-native and multi-agent systems. The role focuses on making agents, pipelines, and inference infrastructure measurable, reliable, and continuously optimized.
Responsibilities
- Design observability pipelines for metrics, logs, traces, and cost telemetry; build dashboards and real-time alerts for reliability, performance, and drift.
- Monitor LLM usage, context windows, token allocation, and multi-agent interactions. Add monitoring hooks to LangChain, LangGraph, MCP, and RAG pipelines.
- Define and track SLOs, SLIs, and SLAs for agentic workflows and inference infrastructure. Investigate agent failures, latency issues, and cost spikes.
- Integrate observability into CI/CD and AgentOps pipelines; write plugins and scripts for monitoring LLMs, agents, and data pipelines.
- Collaborate with AgentOps, DevOps, and Data Engineering teams, report reliability, efficiency, and adoption metrics to executives, and implement feedback loops to improve performance and reduce downtime.
Required qualifications
3–6 years of experience in SRE, DevOps, or AI platform site reliability engineering; knowledge of Prometheus, Grafana, ELK, OpenTelemetry, and Jaeger; experience with AWS, GCP, or Azure and Kubernetes monitoring; Python, Go, or Bash scripting; understanding of AI/LLM pipelines, RAG, and vector databases; and hands-on experience with CI/CD and monitoring-as-code.
Preferred
Experience with LangSmith, PromptLayer, Arize AI, or Weights & Biases; AI-specific monitoring of token usage, model latency, or hallucinations; knowledge of Responsible AI monitoring; or a background in BFSI, GRC, SOC, or other regulated industries.
This full-time role is based in Mohali, India, with a listed salary of 8-15 LPA.