New opportunity

Senior Solutions Architect, Generative AI Deployment and AIOps

NVIDIA · Santa Clara, United States

About this role

Serve as a trusted technical advisor to customers, focusing on inference for Generative AI and LLMs and collaborating with internal teams on performance analysis and modeling of inference software. Work with customers to adopt NVIDIA technology and MLOps solutions and analyze performance and power efficiency of AI inference workloads on Kubernetes.

Skills for this role

Generative AILarge Language Models (LLMs)Deep LearningPyTorchTensorFlowPythonC++GPU orchestrationMulti-Instance GPU (MIG)KubernetesContainerizationOrchestrationMonitoringObservabilityMLOpsInference optimizationPerformance analysisPower efficiency analysisNVIDIA GPUsTensorRTTensorRT-LLMNVIDIA NIMDynamoDebuggingProfilingCode optimizationTest designParallel programmingDistributed computingSoftware design''Communication''Presentation