About this role
Design, implement, and maintain large-scale model training systems and developer tooling to support reinforcement learning and machine learning research; collaborate with researchers to scale distributed training, manage cloud compute resources, and contribute to JAX model and training code.
Skills for this role
Reinforcement LearningMachine LearningMLOpsJAXPyTorchTensorFlowDistributed trainingMulti-host setupsJob orchestrationSchedulingCheckpointingExperiment trackingDeveloper toolingMonitoringDebuggingExperiment reproducibilityResource managementGCPAWSKubernetesSLURMData pipelinesCI/CD for ML workflowsAutomated testing pipelinesLogging/telemetry stacksSoftware engineering fundamentalsInternal developer platformsScalable systemsSystems designRobotics