About this role
Manage the reliability, performance, and operational health of enterprise AI agents in production, including monitoring agent behavior, optimizing costs, managing releases, and ensuring solutions are secure and business-ready. Design and implement observability, incident response, release management, and operational runbooks while supporting AI FinOps and lifecycle controls.
Skills for this role
DevOpsSite Reliability Engineering (SRE)MonitoringObservabilityIncident ManagementRelease ManagementProduction SupportService OperationsReliability EngineeringPerformance OptimizationAutomationOperational ExcellenceStakeholder ManagementAI agentsCopilotsGenerative AIIntelligent Automation PlatformsAWSMicrosoft AzureGoogle Cloud Platform (GCP)CI/CDInfrastructure as CodeTerraformBicepPlatform EngineeringAI ObservabilityAI FinOpsCost OptimizationPerformance MonitoringSAP