About this role
Perplexity is hiring an AI Inference Engineer to build and run their inference engine, deploy transformer-based, text-generation, and multimodal models, develop a Rust-based serving runtime, port CUDA kernels to CuTe DSL, and optimize performance, reliability, and observability in production.
Skills for this role
RustPythonCUDACuTe DSLGPU programmingTritonCUTLASSGPU kernel developmentPyTorchJAXTensorFlowPyTorch internalstorch.compileCustom operatorsNCCLNVLinkInfiniBandRDMAModel parallelismTensor parallelismINT8 quantizationFP8 quantizationFP4 quantizationMixed-precision servingNsight ComputeNsight SystemsCUDA-GDBPTX/SASS analysisKubernetesGPU scheduling''Autoscaling inference workloads'