About this role
Systems-minded AI Software Engineer to design and extend a low-level AI inference serving stack, customize open-source serving frameworks, and integrate with a proprietary AI accelerator. Role focuses on optimizing throughput, latency, and scalability across software and accelerator hardware, including model partitioning, scheduling, and runtime performance tuning.
Skills for this role
Systems programmingML infrastructureDistributed inferenceC++PythonDebuggingPerformance analysisvLLMSGLangTensorRT-LLMDeepSpeed-InferenceModel deployment internalsToken schedulingKV cachingBatchingPipelined inferenceCUDAPCIeMemory managementRuntime schedulingModel exportRuntime optimizationQuantizationGraph transformsPyTorchAPI developmentC++ kernelsPCIe driversHardware-aware ML optimizationCompiler/runtime integration