About this role
Organize and prioritize work across training clusters, data processing pipelines, and data operations teams while allocating GPU resources, managing data procurement, and vendor relationships to support AI model training. Collaborate with AI engineering leadership to define resourcing objectives, identify compute/data risks, and drive process improvements for AI infrastructure and internal AI tools.
Skills for this role
GPU resource allocationTraining cluster managementGPU/compute cluster managementData pipeline managementData operationsData annotationQuality assurance (QA)Project managementProgram managementJIRAGoogle SuiteResource schedulingStakeholder engagementContract negotiationData procurementVendor managementRisk mitigationProcess improvementAI/ML development lifecycleAI/ML model training workflowsCross-functional collaboration