About this role
Engineer scalable AI inference solutions integrated within scientific workflows, optimizing execution on ALCF’s HPC systems and AI accelerators. Work involves developing programmatic interfaces and REST APIs, deploying and adapting AI models (including LLMs), and collaborating with science application teams and partners.
Skills for this role
vLLMSGLangOpenAI APIREST API developmentFastAPIPythonCC++gitDistributed inferenceRequest routingAsynchronous executionQueueingFault tolerancePerformance monitoringSlurmPBSHPC systemsGPUAI acceleratorsBatchingMemory managementModel parallelismQuantizationAccelerator utilizationAuthentication/authorizationAPI securityPrometheusGrafanaLog aggregation (log aggregators)OLAP or data warehousing techniques