About this role
Work on NVIDIA's Local AI team to build the software stack that makes LLMs and generative AI run efficiently on NVIDIA edge AI hardware, including evaluating open-source LLM inference frameworks, analyzing model architectures and inference algorithms for GPU optimization, characterizing multi-node inference behavior, producing performance analysis reports, owning model validation workflows, and developing developer-facing inference recipes.
Skills for this role
Large Language Models (LLMs)LLM inferenceGPU computingML systemsHigh-performance inferencePythonC++GPU kernel developmentCUDATritonAttention mechanismsKV-cache managementContinuous batchingQuantization formatsTensor parallelismNCCLRCCLCollective communicationTopology-aware all-reducePerformance analysisCI/CD pipelinesDockerOCINVIDIA Container ToolkitContainer engineeringModel validationInference recipe developmentMulti-node inferenceParallelism efficiencySoftware design''Software engineering