Welocalize

Senior Software Engineer - Robot Data Collection

Welocalize Santa Clara, California, United States · $218K/yr

Technology, Information and Internet · 1,001-5,000 employees

11 h ago
Mid (2-5 yrs) Full-time Contractor United States
Create a free account to apply — email only, no card. You can also save this posting or score it against your profile with AI.

About the role

The Senior Software Engineer will design, build, and maintain the software infrastructure for large-scale robot data collection, including teleoperation stacks and data pipelines. They will also direct a team of teleoperators, manage fleet reliability, and ensure high-quality data throughput for research teams.

What they look for

Python Linux Docker Git CI SSH ROS 2 Robot middleware Teleoperation Data pipeline Fleet management Sensor integration Calibration Firmware Robotics QA tooling

Requirements

Candidates must have an MS or BS in a relevant engineering field with at least 3 years of experience building software for real-world robot systems. Proficiency in Python, Linux, and robot middleware like ROS 2 is required, along with experience in fleet administration and hardware integration.

Full description

Role Overview

We are looking for a Senior Software Engineer to own the systems that power a large-scale data collection operation, including the collection platform, operator tooling, data pipeline, and fleet infrastructure.

This is an engineering role. Day-to-day teleoperation is performed by a dedicated operator team that you direct — you build what they run on, and you own what comes out of it.

Project Details

  • Location: 100% Onsite (Santa Clara, CA, US)
  • Employment: Full-time W-2, Monday–Friday, 9:00 AM–5:00 PM
  • Start Date: ASAP
  • Contract Duration: 12-month contract with possibility of extension
  • Rate: $105/hour

What You'll Do

  • Design, build, and maintain the software behind our data collection stations — teleoperation stack, multi-camera recording, and session tooling.
  • Deploy and upgrade that stack across the fleet without stopping collection; keep workstations, containers, and networking healthy.
  • Own the data pipeline end to end: automated validation, conversion into our dataset format, upload, metadata, and versioning.
  • Build the QA tooling that catches dropped frames, stream desync, bad calibration, and mislabeled episodes before data reaches training.
  • Direct a team of teleoperators — define task protocols and SOPs, onboard them on new tasks and embodiments, review session quality, and unblock them daily.
  • Own fleet reliability: calibration and firmware pipelines, sensor and end-effector integration, root-cause analysis of recurring faults, and vendor escalation.
  • Bring up new stations and new embodiments; report throughput, yield, and data health to the research team.

What we need to see

  • MS in CS, Robotics, EE/ME, or a related field (or BS with equivalent hands-on experience).
  • 3+ years building software for real robot systems — not simulation only.
  • A system you designed and shipped that other people depended on daily, and that you kept running.
  • Strong Python and Linux; Docker, git, CI, and SSH-based fleet administration.
  • Robot middleware and real-time multi-sensor recording — ROS 2 or equivalent vendor SDKs.
  • Experience supporting a data collection or teleoperation operation: you have built the tools operators depend on and owned throughput and quality numbers.
  • Comfortable around robot hardware — arms or humanoids, dexterous hands, RGB-D cameras, and the calibration that keeps them honest.
  • Able to lead operators without formal authority: you write clear SOPs, train people, and communicate status without being asked.

Ways to stand out from the crowd

  • Humanoid or bimanual platforms (Unitree, Fourier, ALOHA/ARX-class rigs) and dexterous hands.
  • XR teleoperation: Quest / PICO / Vision Pro, motion-capture gloves, whole-body retargeting.
  • Imitation-learning data tooling — LeRobot, HDF5/Zarr episode stores, dataset versioning, S3 at TB–PB scale.
  • Built internal dashboards or tooling for a production line or high-throughput lab.
  • Open-source contributions to robotics data or teleoperation tooling.