Research Engineer, Synthetic Data
Clera Singapore, Singapore · $150K–$250K/yr
Technology, Information and Internet · 2-10 employees
About the role
Design and build end-to-end synthetic data pipelines to convert domain-specific workflows into structured training tasks for AI agents. Collaborate with subject-matter experts to develop, validate, and improve synthetic tasks while defining metrics to quantify their quality.
What they look for
Requirements
Requires 2 to 4 years of experience in software engineering, machine learning, or AI research with a focus on data pipelines or synthetic data systems. Proficiency in Python, Docker, and Linux environments is essential, along with a track record of independently delivering technical projects.
Benefits
Full description
About the Role
This is a hands-on research engineering role focused on building synthetic data pipelines that turn real-world, domain-specific workflows into structured training tasks for AI agents. You will join a small, high-caliber engineering team of Olympiad medalists and published researchers, working at the core of a platform that powers reinforcement learning environments and post-training data for AI labs.
What You'll Do
- Design and build end-to-end synthetic data pipelines that convert domain-specific workflows into realistic, challenging training tasks.
- Collaborate with subject-matter experts to produce synthetic tasks for AI agents across professional and technical domains.
- Develop task generation methods that maximize diversity, realism, and learnability.
- Build tooling to mutate, validate, and iteratively improve synthetic tasks at scale.
- Analyze model and agent performance on synthetic tasks to identify what they teach and where they break down.
- Define and implement metrics to quantify synthetic task quality across diversity, realism, and learnability dimensions.
What We're Looking For
- 2 to 4 years of experience in software engineering, machine learning engineering, or AI research, with a focus on data pipelines, ML infrastructure, or synthetic data systems.
- Proficiency in Python and hands-on experience with Docker and Linux environments.
- Demonstrated experience applying synthetic data research methods to build generation pipelines end-to-end.
- Strong understanding of synthetic data quality criteria, including diversity, realism, and learnability, as well as their inherent limitations.
- Experience designing, implementing, or maintaining evaluation frameworks, benchmarks, or testing environments for AI agents or large language models.
- Track record of independently owning and delivering technical projects with minimal predefined requirements or roadmap.
- Sharp eye for edge cases, subtle inconsistencies, and quality issues in synthetic or algorithmically generated datasets.
- Familiarity with reinforcement learning paradigms, agentic AI workflows, or LLM post-training pipelines is a plus.
- Strong communication skills for effective collaboration across time zones.
Compensation & Benefits
Salary range: $150,000 to $250,000 USD annually. Visa sponsorship is available.
Location
On-site in Singapore.