New opportunity

Member of Technical Staff (AI Inference Engineer)

Perplexity · San Francisco, United States

About this role

Perplexity is hiring an AI Inference Engineer to build and run their inference engine, deploy transformer-based, text-generation, and multimodal models, develop a Rust-based serving runtime, port CUDA kernels to CuTe DSL, and optimize performance, reliability, and observability in production.

Skills for this role

RustPythonCUDACuTe DSLGPU programmingTritonCUTLASSGPU kernel developmentPyTorchJAXTensorFlowPyTorch internalstorch.compileCustom operatorsNCCLNVLinkInfiniBandRDMAModel parallelismTensor parallelismINT8 quantizationFP8 quantizationFP4 quantizationMixed-precision servingNsight ComputeNsight SystemsCUDA-GDBPTX/SASS analysisKubernetesGPU scheduling''Autoscaling inference workloads'