About this role
Owns the inference and orchestration layer that powers AI interactions, building and operating production systems to deliver fast, stable, observable APIs. Focuses on inference pipeline design, orchestration, monitoring, reliability, and optimizing latency, throughput, and cost for AI-powered features.
Skills for this role
PythonNodeJsPyTorchOpenAIAnthropicOpen-source LLMsSQLnoSQLKubernetesDockerInference pipelinesOrchestrationMonitoringLoggingAlertingIncident responseCachingBatchingStreamingLLMsEmbeddingsMultimodalDistributed systemsLow-latency servicesHigh-throughput systemsBackend engineeringAPIsProduction systems