About this role
Lead the design and operation of resilient, enterprise ML platform infrastructure supporting machine learning workloads, LLMs, RAG pipelines, and AI Agents across multi-cloud environments, driving reliability, scalability, observability, and operational excellence.
Skills for this role
Site Reliability EngineeringPlatform EngineeringDevOpsAWSGoogle Cloud Platform (GCP)Microsoft AzureDataikuAmazon SageMaker AIDatabricksGoogle Vertex AILLMsRAG pipelinesAI AgentsAnomaly detectionPredictive analyticsTime-series modelingOperational intelligenceInfrastructure as Code (IaC)TerraformPulumiAWS CDKCI/CDPythonGoBashPrometheusGrafanaDatadogOpenTelemetryDistributed tracing