About this role
Join AI Singapore at NTU to develop and maintain evaluation frameworks and pipelines that measure the capabilities of Large Language Models, focusing on multilingual, multicultural and multimodal assessments. The role involves experimentation with recent LLM evaluation research, dataset collection and verification, data preparation, AI modelling, coding, testing, validation and deployment, and collaboration with cross-functional teams.
Skills for this role
LLM evaluationLLM-as-a-JudgevLLMPyTorchPythonGitData preparationData analysisAI modellingModel testingModel validationModel deploymentMultilingual evaluationMultimodal evaluationDataset collectionTranslation verificationCode repository maintenanceTechnical documentationReading and understanding research papersWritten communicationVerbal communicationCollaboration