About this role
Lead and scale an engineering team building and optimizing open-source GPU-accelerated deep learning inference frameworks (e.g., SGLang, vLLM, FlashInfer) for LLM, multimodal, and generative AI. Oversee performance tuning, profiling, and multi-GPU optimization while partnering with compiler, libraries, and research teams to deliver end-to-end inference pipelines.
Skills for this role
Deep learning inferenceGPU-accelerated softwareSGLangvLLMFlashInferCUDATritonCUTLASSNIXLNCCLNVSHMEMC/C++PythonGPU programmingPerformance optimizationModel deploymentAgile methodologiesPyTorchTensorRT-LLMDistributed inference architecturesPerformance profilingSystem-level optimizationOpen-source developmentCompiler integrationMentoring