New opportunity

Principal Developer, AI Networking

NVIDIA · Santa Clara, United States of America

About this role

Profile, analyze, and optimize AI workloads on large-scale GPU and CPU clusters for distributed deep learning LLM training and inference with a primary focus on collective communication and high-performance networking. Develop PyTorch trace-based profiling, analysis, and replay toolsets, collaborate across hardware and software teams, and define performance test plans to identify bottlenecks and achieve performance targets.

Skills for this role

High-performance networkingRDMAMPINCCLSHARPRoCENVIDIA GPUsCUDATensorFlowPyTorchPyTorch profilingTrace-based profilingBenchmarkingProfilingPerformance analysisPerformance evaluationPerformance test planningPythonBashC++Container-based developmentLLM trainingLLM inferenceCollective communicationHCAsSwitchesCPUsGPUsMemoryPCI