About this role
The Agentic Infrastructure Observability Engineer will design and implement end-to-end observability frameworks for AI-native and multi-agent systems, focusing on monitoring reliability, performance, and drift. This role involves building dashboards, tracking LLM usage and token allocation, and integrating observability into AgentOps pipelines to ensure transparency and resilience in enterprise AI systems.
Skills for this role
SREDevOpsObservability EngineeringPrometheusGrafanaELKOpenTelemetryJaegerAWSGCPAzureKubernetesPythonGoBashAI/LLM pipelinesRAGVector databasesCI/CDMonitoring-as-code