About this role
Early-career engineering role to optimize end-to-end inference latency for a full-duplex multimodal model stack, including LLMs, audio models, and diffusion components. Responsibilities include quantization, KV cache optimization, kernel-level acceleration, extending inference serving frameworks, profiling/benchmarking, and building tooling to meet strict real-time latency SLAs.
Skills for this role
KV cachingKV cache eviction policiesMemory-efficient attentionAttention kernelsMemory layoutQuantization (INT8)Quantization (INT4)GPTQAWQvLLMSGLangTensorRT-LLMPythonPyTorchCUDATritonKernel optimizationCustom kernel optimizationsBatching strategiesProfilingBenchmarkingLatency optimizationThroughput optimizationDiffusion model inferenceConsistency modelsStep distillationSpeculative decodingModel compressionMultimodal inferenceStreaming inference architectures''Serving infrastructure