Research Engineer, Synthetic Data
Clera Singapore, Singapore · $100K–$170K/yr
Technology, Information and Internet · 11-50 employees
Applying here? Try the free cover letter tool — paste this posting and your résumé, no account needed.
About the role
You will build and maintain pipelines to generate, validate, and improve synthetic training tasks for AI agents. Additionally, you will collaborate with subject-matter experts and analyze agent performance to refine model capabilities.
What they look for
Requirements
Candidates should have two to four years of experience in software or machine learning engineering with hands-on expertise in synthetic data generation. Proficiency in Python, Linux, and containerization tools like Docker is required.
Benefits
Full description
About the Role
Join an engineering team at an early-stage AI company building infrastructure for training and evaluating AI agents. You will develop synthetic data pipelines and methods that turn real-world professional workflows into useful training tasks, helping improve AI capabilities across technical and professional domains.
What You'll Do
- Build pipelines that generate realistic, structured, and challenging synthetic training tasks from domain-specific workflows.
- Collaborate with subject-matter experts to create tasks across professional and technical domains.
- Design methods and tools to generate, mutate, validate, and improve synthetic tasks.
- Analyze agent performance to understand what tasks teach and where models fail.
- Develop metrics for task diversity, realism, learnability, and overall quality.
What We're Looking For
- Two to four years of relevant experience in software engineering, machine learning engineering, or AI research.
- Hands-on experience applying synthetic data methods to build end-to-end data generation pipelines for AI or machine learning applications.
- Proficiency in Python, Linux, and containerization tools such as Docker.
- Experience with synthetic data quality criteria, evaluation metrics, and the limitations of synthetic data.
- Experience designing or maintaining evaluation frameworks, benchmarks, or testing environments for AI agents or large language models.
- A track record of independently delivering technical projects and building automated systems to generate, validate, or process structured data at scale.
- Strong attention to detail, first-principles reasoning, and clear communication for collaboration across time zones.
- Familiarity with reinforcement learning, agentic AI workflows, or LLM post-training is useful.
Compensation & Benefits
Compensation is USD 100,000 to USD 170,000 annually. Visa sponsorship is available.
Location
On-site in Singapore, Singapore.