About this role
Embed with enterprise customers to build POCs, MVPs, and production integrations for GenAI systems, shipping code, running benchmarks, debugging production issues, and architecting scalable inference deployments. Guide model selection and fine-tuning (SFT/DPO/RFT), deploy and validate models on inference frameworks (vLLM, SGLang, TensorRT-LLM), tune for latency/throughput/cost, and own technical customer relationships on-site.
Skills for this role
PythonKubernetesInfrastructure engineeringLLM stackModel servingFine-tuning (SFT)DPORFTAWSAzureGCPGPU infrastructurevLLMSGLangTensorRT-LLMLoad testingLatency optimizationThroughput optimizationQuantizationEvaluation frameworksBenchmarkingPOC developmentMVP developmentProduction integrationsModel selectionFine-tuning pipelinesCommunicationStakeholder managementDiscovery callsExecutive presentation skills