About this role
Own inference research to lower cost per token and latency on customer workloads, focusing on low-level inference optimization (speculative decoding, quantization, KV-cache, memory management), working directly with customers to optimize production workloads, and translating learnings into platform improvements.
Skills for this role
LLM inferenceLLM servingLow-level inference optimizationSpeculative decodingQuantizationKV-cacheMemory managementPyTorchTorchReinforcement LearningGPU programmingKernel-level optimizationScheduler designCloud native technologiesDeveloper toolsOpen source softwareCustomer collaboration