About this role
NVIDIA’s LocalAI team seeks a Senior System Software Engineer to architect efficient on-device AI software for RTX and DGX-class systems in India, Pune. The role focuses on low-latency local inference, memory efficiency, reliable infrastructure, and deployment on resource-constrained platforms.
Responsibilities
- Build and optimize the local AI inference stack for RTX, RTX Pro, and DGX GPUs, improving performance, stability, and scalability across hardware architectures.
- Develop inference runtimes and execution stacks using tools such as llama.cpp, vLLM, PyTorch, Windows ML, DXCGC, and TensorRT-RTX for LLM, vision-language, text-to-speech, speech recognition, and diffusion workloads.
- Optimize models, data pipelines, and runtimes end to end. Apply quantization, pruning, sparsity, and distillation to deploy large models locally and on edge devices.
- Lead system-level debugging and performance tuning; assess performance-accuracy trade-offs, build performance and accuracy sweep infrastructure, identify gaps, and establish guidelines for production readiness of new models and inference backends.
- Coordinate technical priorities with software, research, architecture, and product teams and external partners, including Microsoft. Mentor engineers and review architecture proposals across groups.
Required qualifications
- 5+ years of experience with a bachelor’s, master’s, or PhD in Computer Science, Software Engineering, Mathematics, or a related field, or equivalent experience.
- Excellent C++ programming and debugging skills; strong grounding in data structures, algorithms, and machine learning. Experience architecting and optimizing AI inference pipelines and applications with ML/DL frameworks.
- Deep knowledge of inference backend and runtime internals, including scheduling, memory management, KV-cache behavior, graph execution, quantization, and hardware-aware optimization. Ability to set technical direction across teams, solve problems, manage competing priorities, and communicate effectively.
Preferred qualifications
- Deeper expertise in neural networks and generative AI, including exposure to AOT graph compilation such as TensorRT builder or MLIR-based compilers; expert-level GPU kernel programming, CUDA, and high-performance systems development.
- Experience delivering products across geographically distributed teams, significant open-source inference or tooling contributions, and application development with llama.cpp, PyTorch, TensorRT, Vulkan, DirectX, or vLLM.
The posting identifies the position as full-time and mentions competitive salaries and a benefits package without specifying amounts.
Skills for this role
C++Data structuresAlgorithmsMachine learningDeep learningAI inferenceInference runtimesPerformance tuningSystem-level debuggingGPU programmingCUDAllama.cppvLLMPyTorchWindows MLDXCGCTensorRTTensorRT-RTXQuantizationPruningSparsityKnowledge distillationKV-cacheMemory managementGraph executionAOT graph compilationMLIRVulkanDirectXTechnical leadership