About this role
Trusted technical advisor role focused on inference for Generative AI and Large Language Models (LLMs), working with customers and internal teams on performance analysis, modeling, and MLOps solutions. Requires experience analyzing AI inference workloads on Kubernetes and strong background with PyTorch, TensorFlow, Python, GPU/MIG management, containerization, and NVIDIA software.
Skills for this role
Generative AILarge Language Models (LLMs)Deep LearningAccelerated ComputingPyTorchTensorFlowPythonGPU orchestrationMulti-Instance GPU (MIG)KubernetesContainerizationOrchestrationMonitoringObservabilityLLM inferenceDL inferenceMLOpsNVIDIA GPUsNVIDIA NIMDynamoTensorRTTensorRT-LLMC++CDebuggingProfilingCode optimizationPerformance analysisTest designParallel programming''Distributed computing''Presentation skills''Communicati