About this role
Design and implement asynchronous multi-agent orchestration, resilient AI inference pipelines, intelligent request routing, WebSocket/streaming infrastructure, and observability for AI system performance. The role requires experience scaling ML/AI inference in production, event-driven architectures, caching, message queues, multiple LLM/AI models, and AI model serving frameworks.
Skills for this role
Artificial IntelligenceLLMAgentic AIMachine LearningWebSocketsConversational AIAsynchronous multi-agent orchestrationAsync/event-driven architecturesML/AI inferenceRedisIn-memory cachingCDNMessage queuesReal-time communication protocolsAI model servingTensorFlow ServingTritonAI inference optimisationBatchingModel quantisationConversation state managementContext handling