New opportunity

Senior Software Engineer - GPU Local AI Platforms

NVIDIA · Santa Clara, United States of America

About this role

Work on NVIDIA's Local AI team to build the software stack that makes LLMs and generative AI run efficiently on NVIDIA edge AI hardware, including evaluating open-source LLM inference frameworks, analyzing model architectures and inference algorithms for GPU optimization, characterizing multi-node inference behavior, producing performance analysis reports, owning model validation workflows, and developing developer-facing inference recipes.

Skills for this role

Large Language Models (LLMs)LLM inferenceGPU computingML systemsHigh-performance inferencePythonC++GPU kernel developmentCUDATritonAttention mechanismsKV-cache managementContinuous batchingQuantization formatsTensor parallelismNCCLRCCLCollective communicationTopology-aware all-reducePerformance analysisCI/CD pipelinesDockerOCINVIDIA Container ToolkitContainer engineeringModel validationInference recipe developmentMulti-node inferenceParallelism efficiencySoftware design''Software engineering