About this role
This Manager-level role within the AI Verify Foundation focuses on improving open-source AI evaluation products for developers. The foundation develops Project Moonshot for evaluating LLM applications and the AI Verify Toolkit for predictive machine learning models. The Product Analyst will connect AI evaluation research with users’ needs and lead quality assurance for product datasets and evaluators.
Responsibilities
- Design and run offline A/B experiments to improve Project Moonshot’s benchmark evaluation pipelines, including LLM-as-jury approaches.
- Develop and maintain benchmark datasets, including realistic test cases relevant to Singapore.
- Research emerging generative AI evaluation practices and open-source tools; use findings to inform the product roadmap and create actionable prototypes for engineering teams.
- Engage and educate the developer community about benchmark-testing challenges and the solutions offered by the foundation’s library.
Requirements
- A background in Statistics, Data Science or a related technical field, and 1–3 years of experience as a data scientist.
- Strong experimental design and analysis skills, with proficiency in Python or R.
- A practical, user-centred approach and strong communication skills, including storytelling with data and explaining complex concepts to varied audiences.
Preferred
Experience in AI testing is a plus. The position level will be commensurate with the candidate’s qualifications and experience.