Clera

Research Engineer, Synthetic Data

Clera · San Francisco, California, United States · $150K–$250K/yr

Technology, Information and Internet · 2-10 employees

7 h ago
Mid (2-5 yrs) Full-time United States
Log in to apply, save this posting, or score it against your profile with AI.

About the role

Design and build scalable pipelines for generating synthetic post-training datasets and reward signals. Collaborate with researchers to develop evaluations and automated QA systems to ensure model alignment and data quality.

What they look for

Python Machine learning infrastructure Reinforcement learning RLHF Synthetic data generation LLM evaluation Reward modeling Distributed systems Task orchestration Data pipelines Agent frameworks Software engineering

Requirements

Requires strong software engineering fundamentals in Python and hands-on experience with reinforcement learning or post-training data pipelines. Candidates should be comfortable reading ML research papers and translating them into functional systems.

Benefits

Equity participation Cutting-edge AI alignment work Small team environment

Full description

About the Role

We're a fast-growing AI infrastructure company building the foundational platform for reinforcement learning (RL) environments — enabling AI labs and enterprises to train, evaluate, and align AI models to real-world workflows at scale. Backed by early-stage venture funding, our small but highly technical team is tackling some of the hardest problems in post-training data generation and AI alignment.

We're looking for a Research Engineer, Synthetic Data to join our team in San Francisco. In this role, you'll sit at the intersection of research and engineering, designing and building systems that produce high-quality synthetic datasets used to train and align state-of-the-art AI models.

What You'll Do

  • Design, build, and iterate on pipelines for generating synthetic post-training datasets at scale.
  • Develop and refine reward signals, verifiers, and graders to ensure data quality and model alignment.
  • Build tooling to detect reward hacking, false positives/negatives, and grader-prompt misalignment in agent traces.
  • Run large-scale agent tasks across concurrent RL environments and analyze results to improve data quality.
  • Collaborate closely with ML researchers to design evaluations and environment definitions that encode domain expertise.
  • Contribute to automated QA systems that audit agent behavior and surface root causes of failures.
  • Experiment with novel approaches to scalable environment creation and task execution for model training.

What We're Looking For

Required

  • Strong software engineering fundamentals with experience in Python and ML infrastructure.
  • Hands-on experience with reinforcement learning, RLHF, or post-training data pipelines.
  • Familiarity with LLM evaluation frameworks, reward modeling, or synthetic data generation techniques.
  • Ability to move fast in a small-team environment — design, implement, and ship end-to-end.
  • Strong research instincts: comfort reading ML papers and translating ideas into working systems.

Nice to Have

  • Experience building or working with AI agent frameworks and agentic task environments.
  • Background in distributed systems or large-scale task orchestration.
  • Prior work at an AI lab, research group, or AI/ML infrastructure company.
  • Publications or open-source contributions in RL, alignment, or synthetic data.

Compensation & Benefits

  • Salary: $150,000 – $250,000 USD annually, depending on experience.
  • Equity participation in an early-stage, high-growth company.
  • Opportunity to work on cutting-edge AI alignment and RL infrastructure problems.
  • Small team with significant individual impact and ownership.

Please note: visa sponsorship is not available for this role.

Location

This is an on-site role based in San Francisco, CA. Candidates should be able to work from the San Francisco office.