Public Service Division

AI Product Analyst (Developer Community), AI Verify Foundation

Public Service Division Singapore

IT Services and IT Consulting · 1,001-5,000 employees

13 h ago
product-analyst Mid (2-5 yrs) Full-time Singapore
Log in to apply, save this posting, or score it against your profile with AI.

About the role

You will design and run offline A/B experiments to optimize evaluation pipelines for LLM applications and benchmark tests. Additionally, you will contribute to the development of benchmark datasets and engage the developer community to promote responsible AI practices.

What they look for

Statistics Data Science Python R Experimental Design Data Analysis AI Testing LLM Evaluation Machine Learning Benchmark Testing Product Analysis Data Storytelling Communication Research Prototyping

Requirements

Candidates must have a background in Statistics or Data Science with 1-3 years of experience in data science and experimental design. Proficiency in Python or R is required, along with strong communication skills and a user-centric mindset.

Full description

[What the role is]

This role is pegged at the Manager level, within the AI Verify Foundation.

About Us AI Verify Foundation is an open-source initiative and a wholly owned subsidiary under the Infocomm Media Development Authority (IMDA), dedicated to advancing trustworthy AI. We develop and maintain open-source tools for evaluating LLM applications (Project Moonshot) and predictive ML models (AI Verify Toolkit). Our mission is to empower developers, data scientists, and organisations to build responsible, transparent, and reliable AI systems.

The Role We’re looking for a Product Analyst to elevate the quality and value of our product offerings to our users. You’ll bridge cutting-edge AI research with real-world evaluation needs, shaping our tools and frameworks for the success of our developer community. You will lead the quality assurance of the datasets and evaluators in our products.

[What you will be working on]

  • Design and run offline A/B experiments to design optimal evaluation pipelines for benchmark tests offered in Project Moonshot, including LLM-as jury.
  • Contribute to the development and maintenance of benchmark datasets, particularly generating realistic test cases in the context of Singapore.
  • Research emerging GenAI evaluation best practices and open-source tools — translate key insights to inform product roadmap, and turn research findings into actionable prototypes for engineering teams.
  • Educate and engage the developer community to build awareness of the challenges in benchmark testing, and champion the solutions provided in our library.

[What we are looking for]

  • Background in Statistics or Data Science, or a related technical field.
  • 1-3 years’ experiences as a data scientist, with strong grounding in experimental design and analysis.
  • Proficiency in Python/ R is required. Experience in AI testing is a plus.
  • Deep intellectual capacity combined with a practical, user-centric mindset.
  • Excellent communication skills— excel in storytelling with data, and able to simplify complex concepts for varied audiences.

Position will be commensurate with the candidate’s qualifications and experience.

Please note that only shortlisted candidates will be notified.

#LI-NA1

Similar roles