Clera

Research Engineer, Synthetic Data

Clera Singapore, Singapore · $150K–$250K/yr

Technology, Information and Internet · 2-10 employees

Yesterday
Mid (2-5 yrs) Full-time Visa sponsorship Singapore
Create a free account to apply — email only, no card. You can also save this posting or score it against your profile with AI.

About the role

You will design and build end-to-end synthetic data pipelines to create high-quality training tasks for AI agents. Additionally, you will collaborate with subject-matter experts to develop metrics and evaluate model performance across various technical domains.

What they look for

Python Linux Docker Synthetic data Machine learning Data pipelines ML infrastructure AI agents Reinforcement learning LLM post-training Evaluation frameworks Benchmarking

Requirements

The role requires 2 to 4 years of experience in software engineering, machine learning, or AI research with a focus on data pipelines. Candidates must be proficient in Python, comfortable in Linux environments, and experienced with containerization tools like Docker.

Benefits

Visa sponsorship

Full description

About the Role

This is a Research Engineer role focused on synthetic data, sitting within a roughly 15-person engineering team of Olympiad medalists and published researchers. You will design and build the pipelines that turn domain-specific workflows into scalable, high-quality training tasks for AI agents, directly expanding what the models can do.

What You'll Do

  • Build end-to-end synthetic data pipelines that transform domain-specific workflows into realistic, structured, and challenging training tasks.
  • Collaborate with subject-matter experts to create synthetic tasks for AI agents across professional and technical domains.
  • Design task generation methods that produce diverse, realistic, and learnable outputs at scale.
  • Build tooling to mutate, validate, and continuously improve synthetic tasks.
  • Analyze model and agent performance on synthetic tasks to understand what they teach and where they fail.
  • Develop metrics to quantify task diversity, realism, learnability, and overall quality.

What We're Looking For

  • 2 to 4 years of experience in software engineering, machine learning engineering, or AI research, with a focus on data pipelines, ML infrastructure, or synthetic data systems.
  • Hands-on experience applying synthetic data research methods to build end-to-end data generation pipelines for AI/ML applications.
  • Proficiency in Python and comfortable working in Linux environments with containerization tools such as Docker.
  • Strong understanding of synthetic data quality criteria, including diversity, realism, and learnability, and awareness of its inherent limitations.
  • Experience designing, implementing, or maintaining evaluation frameworks, benchmarks, or testing environments for AI agents or large language models.
  • Proven ability to independently own and deliver technical projects end-to-end with minimal predefined requirements.
  • Detail-oriented approach to spotting edge cases and subtle inconsistencies in algorithmically generated datasets.
  • Familiarity with reinforcement learning training paradigms, agentic AI workflows, or LLM post-training pipelines is a plus.
  • Experience creating synthetic tasks or evaluations across multiple distinct professional or technical domains is a plus.
  • Comfortable thriving in unstructured, early-stage startup environments and collaborating across time zones.

Compensation & Benefits

Salary range: $150,000 to $250,000 USD annually. Visa sponsorship is available.

Location

On-site in Singapore.