About this role
The Applied AI Site Reliability Engineer II will take a hands-on approach to the reliability, performance, and operational integrity of high-visibility AI products and platforms. This role focuses on operating AI and agentic workloads at scale, ensuring production systems are safe, performant, and cost-effective. Responsibilities include defining and owning Service Level Objectives (SLOs) and error budgets, building production observability, and managing the admission of systems into production. You will lead performance and resilience testing, conduct blameless postmortems, and develop automation to optimize reliability. The role requires deep expertise in cloud-native engineering, AI/ML production failure modes (such as drift, train/serve skew, and latency), and AI control plane management. You will collaborate with cross-functional teams including platform engineering, security, and data governance to uphold production standards. Required qualifications include a bachelor's degree in computer science, software engineering, data science, machine learning, or a related field. Candidates must have at least 5 years of software engineering and SRE experience with large-scale, distributed, cloud-native systems, including 3 years specifically in site reliability or production engineering for large-scale systems and 3 years in cloud-native engineering on hyperscalers like Azure. Prior experience operating AI/ML and agentic workloads in production is required, along with proficiency in load testing, chaos engineering, and AI cost engineering. The role involves 10% travel and may offer limited immigration sponsorship.