About this role
Support the operational health of Qira's production and non-production systems, including system monitoring, alert response, incident response, and observability across the AI stack. Work includes monitoring model performance and inference pipelines, maintaining runbooks and tooling, and supporting releases and SDLC operations in staging and pre-production.
Skills for this role
Incident ResponseSite Reliability Engineering (SRE)ObservabilityGrafanaDatadogCloud-native dashboardsPagerDutyOpsGenieJiraServiceNowTroubleshootingMicrosoft AzureModel performance monitoringInference latency monitoringData pipeline monitoringAlertingRunbook developmentScriptingSDLC operationsDeployment verificationIncident managementOn-call supportP50/P95/P99 latency metricsMTTA/MTTM/MTTRGolden signals (latencytrafficerrorssaturation)Cross-functional communication