About this role
Integrate advanced communication libraries (e.g., NCCL, NVSHMEM, GPUDirect) into AI frameworks like PyTorch and JAX, perform deep performance analysis and benchmarking on multi-GPU clusters, and design scalable, fault-tolerant solutions for large-scale AI workloads.
Skills for this role
Deep Learning FrameworksPyTorchJAXTRT-LLMvLLMSGLangNCCLNVSHMEMGPUDirectMulti-GPU communicationHigh Performance Computing (HPC)PythonC++CUDATritoncuTeCompiler technologiestorch.compilePerformance benchmarkingPyTorch profilerNVIDIA Nsight SystemsParallel programmingMPIDistributed inferenceMixture of Experts (MoE)Reinforcement LearningKernel authoringAI compiler pattern matchingMemory hierarchyTensor layout""Fault-tolerant system design