About this role
Build and run the training and inference systems that turn Core AI's research into scalable production systems, owning distributed training jobs, GPU scheduling, and model-serving systems. Partner with Engineering on capacity planning, observability, and reliability, and debug and harden the training and inference stack under real load.
Skills for this role
Distributed trainingMulti-GPU trainingMulti-node trainingGPU schedulingModel servingvLLMSGLangTritonPyTorchRayTraining infrastructureInference infrastructureCapacity planningObservabilityReliability engineeringDebugging under loadTraining-as-a-service APIsInference gatewaysJob schedulersGPU infrastructureCUDAPerformance engineering for ML workloadsGCPAWSTerraformOn-call engineeringUnderstanding of training loopsReward signalsInference-time debugging