O

Founding Machine Learning - World Models

One Robot · San Francisco, California, United States · $150K–$275K/yr

Research Services · 2-10 employees

20 h ago
Senior (5-10 yrs) Full-time United States
Log in to apply, save this posting, or score it against your profile with AI.

About the role

Develop and train generative world models with action conditioning to simulate manipulation scenes and validate policies. Manage training infrastructure, including multi-GPU clusters and custom CUDA development, while building a data engine to scale learning across tasks.

What they look for

Machine Learning Python PyTorch Generative Models Video Generation Dynamics Models CUDA Multi-GPU Clusters 3D Vision Multi-view Geometry Scene Reconstruction Physical Priors Robotics Data Engineering

Requirements

Requires strong proficiency in Python and PyTorch, along with deep experience in end-to-end training of image or video generation models. Candidates must have a proven track record of operating large-scale training clusters and possess knowledge of 3D vision and physical priors.

Full description

We build world models that simulate manipulation scenes faithfully enough to validate, and one day, train policies without touching a robot. You'll develop generative models that make this work, with the controllability and physical fidelity to match real-robot behavior.

What you'll do:

  • Train video and dynamics models: Develop world models with action conditioning for manipulation policies.
  • Push long-horizon coherence: Develop architectures and training methods that extend rollout quality on hard physical tasks.
  • Own training infrastructure: Run multi-GPU clusters, write custom CUDA, debug at scale.
  • Build the world-model data engine: Design, implement, and improve a data engine that allows the world model to compound learning across customers and manipulation tasks.

Requirements:

  • Very strong coding in Python and PyTorch (or similar).
  • Video generation experience: Deep experience training image or video generation models end-to-end.
  • Large-scale training: Track record operating training runs at cluster scale.
  • 3D vision: Working knowledge of multi-view geometry, scene reconstruction, and physical priors.