New opportunity

AI Engineer (Evaluation)

Nanyang Technological University · Singapore

About this role

AI Singapore, hosted at Nanyang Technological University, is hiring an AI Engineer (Evaluation) at NTU Main Campus, Singapore. The engineer will work with AI scientists, apprentices, and data and software engineers to build evaluations that test the limits of AI models, particularly their multilingual, multicultural and multimodal capabilities.

Responsibilities

  • Develop and maintain frameworks and pipelines that measure the capabilities of large language models (LLMs). Follow and experiment with research on multilingual, multicultural and multimodal LLM evaluation, including LLM-as-a-Judge methods.
  • Work with partners to collect, translate and verify evaluation datasets. Prepare and analyse data, model AI solutions, and carry out coding, testing, validation and deployment to support reliable, scalable solutions.
  • Collaborate with cross-functional AI Products teams on design and issue resolution. Maintain code repositories and documentation standards, and contribute to technical meet-ups, articles and discussion forums.

Requirements

  • A degree in computer science, AI, data science or a related field, or equivalent practical experience.
  • Deep understanding of LLM evaluation experimental design and the advantages and limitations of different evaluation methods. Experience with LLM inference frameworks such as vLLM and AI or deep-learning frameworks such as PyTorch.
  • Ability to write production-level Python code and use version control systems such as Git. Strong written and verbal communication, independent learning, and the ability to read and understand research papers.
  • Fluency in English and one other Southeast Asian language to support high-quality multilingual and multicultural evaluations.

This is a full-time position. No salary or minimum number of years of experience is specified.

Skills for this role

LLM evaluationEvaluation framework developmentEvaluation pipeline developmentExperimental designMultilingual evaluationMulticultural evaluationMultimodal evaluationLLM-as-a-JudgeEvaluation dataset preparationData analysisAI modellingLLM inferencevLLMPyTorchPythonGitTestingValidationDeploymentTechnical documentationResearch paper analysisWritten communicationVerbal communicationEnglishSoutheast Asian language

Your skill match

Checking your profile…

YOUR NEXT STEP

Get interview-ready for this role

A focused preparation guide, built around this job’s responsibilities and requirements.

✦ AI-generated guide
Preparation suggestions, not the employer’s actual interview questions. Always check the original posting for current requirements.

Loading this role’s preparation guide…

KEEP EXPLORING

Similar AI jobs

Related skills and specializations in Singapore. Matched to this role, not your profile.

Explore more jobs
Finding similar opportunities…