About this role
Design, build, operate and optimise HPC clusters and GPU orchestration layers for AI training, inference and scientific workloads; maintain hardware, distributed storage, high-speed networking and supporting IT infrastructure, and partner with engineers to automate operations and translate research code into performant distributed workloads.
Skills for this role
LinuxPythonBashComputer architectureKubernetesAnsibleTerraformNVIDIA GPUsGitOpsInfra CI/CDNetworking protocolsDistributed storageHigh-speed networkingParallel computingDistributed systemsQuantum computingGPU orchestrationHPC cluster operationsWorkload characterisation and optimisationAutomationPerformance analysisTroubleshootingContainer orchestrationInfrastructure provisioningFleet automation