New opportunity

Member of Technical Staff, ML Engineer

Physical Superintelligence · Boston, United States

About this role

Build and run the training and inference systems that turn Core AI's research into scalable production systems, owning distributed training jobs, GPU scheduling, and model-serving systems. Partner with Engineering on capacity planning, observability, and reliability, and debug and harden the training and inference stack under real load.

Skills for this role

Distributed trainingMulti-GPU trainingMulti-node trainingGPU schedulingModel servingvLLMSGLangTritonPyTorchRayTraining infrastructureInference infrastructureCapacity planningObservabilityReliability engineeringDebugging under loadTraining-as-a-service APIsInference gatewaysJob schedulersGPU infrastructureCUDAPerformance engineering for ML workloadsGCPAWSTerraformOn-call engineeringUnderstanding of training loopsReward signalsInference-time debugging