New opportunity

Member of Technical Staff — Model Optimization and Inference (New Grad)

Nuance Labs · Seattle, United States

About this role

Early-career engineering role to optimize end-to-end inference latency for a full-duplex multimodal model stack, including LLMs, audio models, and diffusion components. Responsibilities include quantization, KV cache optimization, kernel-level acceleration, extending inference serving frameworks, profiling/benchmarking, and building tooling to meet strict real-time latency SLAs.

Skills for this role

KV cachingKV cache eviction policiesMemory-efficient attentionAttention kernelsMemory layoutQuantization (INT8)Quantization (INT4)GPTQAWQvLLMSGLangTensorRT-LLMPythonPyTorchCUDATritonKernel optimizationCustom kernel optimizationsBatching strategiesProfilingBenchmarkingLatency optimizationThroughput optimizationDiffusion model inferenceConsistency modelsStep distillationSpeculative decodingModel compressionMultimodal inferenceStreaming inference architectures''Serving infrastructure