New opportunity

Principal Developer, AI Networking

NVIDIA · Santa Clara, United States of America

About this role

NVIDIA’s AI Networking Codesign and Benchmarking R&D group seeks a Principal Developer to analyze and optimize AI workloads on large-scale GPU and CPU clusters used for distributed deep learning and LLM training and inference. The role focuses on collective communication and high-performance networking across hardware systems, machine learning frameworks, and communication and computing libraries.

Responsibilities

  • Characterize AI workloads and deep learning models on NVIDIA supercomputers, with particular attention to distributed systems and networking.
  • Benchmark and profile performance, identify bottlenecks, and recommend optimizations. Develop a PyTorch trace-based toolset for profiling, analysis, replay, benchmarking, debugging, and network-system co-design for LLM workloads.
  • Collaborate with hardware and software teams to share performance insights. Define performance test plans and expectations for new technologies and work toward performance targets.

Required qualifications

  • B.Sc. in Computer Science or Software Engineering, or equivalent experience, and 15+ years of experience with high-performance networking, including RDMA, MPI, NCCL, or SHARP.
  • Demonstrated performance-evaluation ability; experience with NVIDIA GPUs and CUDA; knowledge of TensorFlow or PyTorch; and expertise in collective communication libraries such as NCCL and protocols such as RoCE and RDMA.
  • Proficiency in Python, Bash, and C++; experience in container-based development; strong analytical, problem-solving, learning, communication, and teamwork skills.

Preferred qualifications

  • Hands-on experience benchmarking AI workloads for distributed LLM training; knowledge of PyTorch, CUDA, and NCCL; broad systems knowledge covering Intel, AMD, or ARM CPUs, NVIDIA GPUs, HCAs, memory, and PCI; and strong performance-evaluation skills using contemporary tools.

Locations listed are Santa Clara, California, and remote roles in Texas, Colorado, and Washington. The listed base salary ranges are 272,000–431,250 USD for Level 6 and 320,000–488,750 USD for Level 7, with equity and benefits eligibility. Applications will be accepted at least until June 16, 2026.

Skills for this role

High-performance networkingDistributed systemsLLM trainingLLM inferenceDeep learningPerformance benchmarkingPerformance profilingCollective communicationRDMAMPINCCLSHARPRoCECUDAPyTorchTensorFlowPythonBashC++Container-based developmentNVIDIA GPUsPerformance analysisAnalytical problem-solvingCommunicationCollaboration

Your skill match

Checking your profile…

YOUR NEXT STEP

Get interview-ready for this role

A focused preparation guide, built around this job’s responsibilities and requirements.

✦ AI-generated guide
Preparation suggestions, not the employer’s actual interview questions. Always check the original posting for current requirements.

Loading this role’s preparation guide…

KEEP EXPLORING

Similar AI jobs

Related skills and specializations in United States of America. Matched to this role, not your profile.

Explore more jobs
Finding similar opportunities…