New opportunity

Engineering Manager, Deep Learning Inference

NVIDIA · Santa Clara, United States

About this role

Lead and scale an engineering team building and optimizing open-source GPU-accelerated deep learning inference frameworks (e.g., SGLang, vLLM, FlashInfer) for LLM, multimodal, and generative AI. Oversee performance tuning, profiling, and multi-GPU optimization while partnering with compiler, libraries, and research teams to deliver end-to-end inference pipelines.

Skills for this role

Deep learning inferenceGPU-accelerated softwareSGLangvLLMFlashInferCUDATritonCUTLASSNIXLNCCLNVSHMEMC/C++PythonGPU programmingPerformance optimizationModel deploymentAgile methodologiesPyTorchTensorRT-LLMDistributed inference architecturesPerformance profilingSystem-level optimizationOpen-source developmentCompiler integrationMentoring