About this role
Lead development and commercialization of the Qualcomm AI Runtime (QAIRT) SDK on Qualcomm SoCs to enable on-device Generative AI inference of LLMs and LVMs, optimizing performance for heterogeneous AI hardware. The role focuses on deploying large C/C++ software stacks, inference optimization for CPU/GPU/NPU, and edge-based GenAI deployment while collaborating across global teams.
Skills for this role
Generative AILarge Language Models (LLMs)Large Vision Models (LVM)Large Multimodal Models (LMMs)TransformersSelf-attentionCross-attentionKV cachingFloating-point representationsFixed-point representationsQuantizationAI hardware acceleratorsCPUGPUNPUCC++Design PatternsOS conceptsPythonSIMD processor architectureKernel developmentllama.cppMLXMLCPyTorchTFLiteONNX RuntimeOpenCLCUDA''Object-oriented programming''Linux''Windows''Debugging''Qualcomm AI (