New opportunity

AI Engineer, Inference

Firmus Technologies · Sydney, Australia

About this role

Build and improve self-hosted AI inference services and model-serving platforms, including model onboarding, deployment, performance optimization, benchmarking, observability, and secure, scalable endpoint provisioning for internal and customer-facing applications. Contribute inference recipes, workload profiles, and operational controls to support Model-to-Grid capabilities and agentic applications.

Skills for this role

Model servingInference servicesModel onboardingEndpoint provisioningRuntime selectionPerformance benchmarkingObservabilityCapacity managementSecurity for inference servicesOperational lifecycle managementTensorRT-LLMTensorRTSGLangvLLMTriton Inference ServerNVIDIA DynamoNVIDIA NIMCUDAcuDNNNCCLHugging Face Text Generation InferenceQuantizationFP8INT8NVFP4TensorRT compilationKernel optimizationBatchingKV-cache managementSpeculative decoding''Memory optimization''Tensor parallelism''Pipeline paral