About this role
AI Singapore (AISG) seeks an AI Engineer for its AI Products Platform team, hosted at Nanyang Technological University’s main campus in Singapore. The role supports production LLM deployment and inference, customized LLM-based solutions, retrieval-augmented generation services, AI agent orchestration platforms and GPU-enabled infrastructure.
Responsibilities
- Operate SEA-LION API Farm, a multi-cloud LLM inference platform. Monitor system health, GPU capacity, performance, cost and security; troubleshoot reliability issues and optimize inference across modalities.
- Build API services, including batch and Model Context Protocol (MCP) services. Manage high-performance AI clusters and storage across cloud providers using infrastructure as code.
- Develop CI/CD pipelines, container build and registry workflows, and deployment automation. Improve logs, metrics, traces and dashboards, and automate repetitive operational work.
- Collaborate with internal teams on optimized inference workflows and use AI tools appropriately to support engineering and build internal operational tools.
Requirements
- A degree in Computer Science, Information Technology or equivalent, and 1–3 years of DevOps, site reliability or platform engineering experience operating production systems at scale.
- Hands-on experience with multi-cloud workloads, infrastructure as code, containers, orchestration and managed compute, storage and networking services. Knowledge of inference frameworks such as vLLM, SGLang, TensorRT-LLM and Transformers.
- Experience serving LLMs in production, including GPU scheduling, autoscaling, latency and throughput optimization, and inference cost management. Knowledge of REST API design, MCP, authentication patterns such as OAuth, CI/CD, observability and incident response.
- Python or Bash scripting skills, experience using AI tools for engineering tasks, ability to read code across the stack and strong technical communication skills.
Preferred
Experience with multimodal models, C, C++, Rust or Go, or contributions to open-source AI/ML projects.
Skills for this role
LLM inferenceModel servingGPU schedulingAutoscalingInference optimizationInference cost managementMulti-cloud infrastructureInfrastructure as codeTerraformDockerKubernetesCI/CDObservabilityIncident responseREST API designModel Context Protocol (MCP)OAuthvLLMSGLangTensorRT-LLMTransformersPythonBashClaudeCopilotCursorRetrieval-augmented generation (RAG)AI agent orchestrationCC++ evelopment and deployment of reliableefficient LLM services. The engineer