Clera

Research Engineer, Synthetic Data

Clera San Francisco, California, United States · $150K–$250K/yr

Technology, Information and Internet · 2-10 employees

Yesterday
Mid (2-5 yrs) Full-time Visa sponsorship United States
Log in to apply, save this posting, or score it against your profile with AI.

About the role

You will build and maintain end-to-end synthetic data pipelines to convert domain-specific workflows into high-quality training tasks for AI agents. Additionally, you will collaborate with subject-matter experts to design evaluation metrics and improve synthetic task generation at scale.

What they look for

Python Synthetic data Machine learning Data pipelines Linux Docker Reinforcement learning LLM AI agents Evaluation frameworks Data generation ML infrastructure Benchmarking Algorithm design

Requirements

The role requires 2–4 years of experience in software engineering, ML engineering, or AI research with a focus on data pipelines or synthetic systems. Proficiency in Python and experience with Linux and containerization tools like Docker are essential.

Benefits

Visa sponsorship Equity participation

Full description

About the Role

We're a ~15-person engineering team — made up of Olympiad medalists and published researchers — building infrastructure that aligns AI to real-world workflows through reinforcement learning environments and post-training data. We're hiring Research Engineers to own the synthetic data pipeline: transforming domain-specific workflows into scalable, high-quality training tasks for AI agents.

This is a high-ownership, low-bureaucracy role. You'll be working in genuinely unstructured problem spaces where the roadmap is yours to define. Visa sponsorship is available.

What You'll Do

  • Build and maintain the end-to-end synthetic data pipeline, converting domain-specific workflows into realistic, structured, and challenging training tasks for AI agents.
  • Collaborate with subject-matter experts to generate synthetic tasks across professional and technical domains.
  • Design synthetic task generation methods that produce diverse, realistic, and learnable outputs.
  • Build tooling to mutate, validate, and iteratively improve synthetic tasks at scale.
  • Analyze model and agent performance on synthetic tasks to understand what they teach and where they break down.
  • Develop metrics to quantify synthetic task diversity, realism, learnability, and overall quality.

What We're Looking For

Required:

  • 2–4 years of experience in software engineering, ML engineering, or AI research — with a track record of shipping data pipelines, ML infrastructure, or synthetic data systems.
  • Hands-on experience applying synthetic data research methods to build end-to-end data generation pipelines for AI/ML applications.
  • Proficiency in Python; comfortable working in Linux environments with containerization tools such as Docker.
  • Demonstrated understanding of synthetic data quality criteria and evaluation metrics (diversity, realism, learnability) and their limitations — from production or research work.
  • Experience designing, implementing, or maintaining evaluation frameworks, benchmarks, or testing environments for AI agents or large language models.
  • Experience building automated systems to generate, validate, mutate, or process structured datasets at scale.
  • Proven ability to independently own and deliver technical projects end-to-end with minimal predefined requirements.

Nice to Have:

  • Experience detecting edge cases, inconsistencies, or quality issues in synthetic or algorithmically generated datasets.
  • Experience creating synthetic tasks, data, or evaluations across multiple distinct professional or technical domains.
  • Familiarity with reinforcement learning training paradigms, agentic AI workflows, or LLM post-training pipelines.

You'll thrive here if you:

  • Reason from first principles about task design, scoring, and failure modes.
  • Are detail-oriented and naturally spot subtle inconsistencies in data and systems.
  • Are energised by early-stage, ambiguous environments rather than frustrated by them.
  • Communicate clearly and collaborate effectively across time zones.

Compensation & Benefits

  • Salary: $150,000 – $250,000 USD annually
  • Visa sponsorship available
  • Equity participation (early-stage startup)

Location

This role is on-site in San Francisco, CA. Candidates based in or willing to relocate to San Francisco are strongly preferred. The team also has a presence in Singapore.