About this role
Lead the design, implementation, and scaling of the company’s ML infrastructure and production AI systems, working across the model lifecycle (training, evaluation, deployment, observability, and optimization). Partner with AI researchers, GPU systems engineers, backend teams, and product stakeholders to ensure robust, efficient, automated, and production-grade large-scale AI systems.
Skills for this role
ML OpsDistributed systemsCloud infrastructureAWSGCPAzurePythonTypeScriptGoPyTorchTransformersvLLMLlama-factoryMegatron-LMCUDAGPU accelerationDockerKubernetesHelmAutoscalingModel deploymentModel versioningReproducibilityOrchestrationObservabilityMonitoringDataset curationData labelingFeature pipelinesCI/CD''Continuous integration and deployment'