New opportunity

AI Reliability Operations Engineer - - 80842

Lenovo · Chicago, United States of America

About this role

Support the operational health of Qira's production and non-production systems, including system monitoring, alert response, incident response, and observability across the AI stack. Work includes monitoring model performance and inference pipelines, maintaining runbooks and tooling, and supporting releases and SDLC operations in staging and pre-production.

Skills for this role

Incident ResponseSite Reliability Engineering (SRE)ObservabilityGrafanaDatadogCloud-native dashboardsPagerDutyOpsGenieJiraServiceNowTroubleshootingMicrosoft AzureModel performance monitoringInference latency monitoringData pipeline monitoringAlertingRunbook developmentScriptingSDLC operationsDeployment verificationIncident managementOn-call supportP50/P95/P99 latency metricsMTTA/MTTM/MTTRGolden signals (latencytrafficerrorssaturation)Cross-functional communication