New opportunity

Member of Technical Staff, Inference

Inferact · Singapore

About this role

Work on the core of vLLM to optimize how LLMs and diffusion models execute across diverse hardware and architectures, improving inference performance, memory, and reliability. Responsibilities include implementing and debugging inference techniques, contributing performant code to the inference engine, and supporting advanced model-serving paradigms.

Skills for this role

Transformer architecturesPythonPyTorch internalsvLLMTensorRT-LLMSGLangTGIInference runtime engineeringModel servingKV-cache memory managementPrefix cachingHybrid model servingReinforcement learning frameworksMultimodal inferencePerformance optimizationDebugging ML codebasesOpen-source contributions