About this role
Responsible for hosting, deploying, and operating open-weight LLMs within a sovereign cloud environment, focusing on GPU infrastructure, model serving, performance optimization, and reliable model lifecycle management.
Skills for this role
MLOpsLLM infrastructuremodel servingGPU-based inferenceKubernetesvLLMNVIDIA TritonTensorRT-LLMLLaMAGemmaMistralGPTDockerCI/CDmodel registriesobservabilityLLM quantizationbatchingcachingGPU memory managementinference optimizationprivate cloudon-premisessovereign cloudNVIDIA GPUsCUDAHelmPrometheusGrafanaMLflow