AI Robot Association

R&D-038 Data Engineer

AI Robot Association Ota, Japan

Technology, Information and Internet · 2-10 employees

6 h ago
data-engineer Senior (5-10 yrs) Full-time Japan
Log in to apply, save this posting, or score it against your profile with AI.

About the role

You will design and implement large-scale data pipelines to manage the full lifecycle of high-quality datasets for robotics foundation models. Additionally, you will build and maintain storage solutions and query interfaces to enable researchers to efficiently access and utilize multimodal data.

What they look for

Data Engineering Distributed Systems ETL Pipelines Spark Flink Ray Kubernetes Multimodal Data Python Data Pipelines Cloud Infrastructure Data Curation Robotics Machine Learning Data Schema Design Orchestration Tools

Requirements

Candidates must have a bachelor's degree in a relevant field and at least 5 years of professional experience in data engineering or platform development. Proficiency in designing distributed data systems and handling large-scale unstructured data is essential for this role.

Full description

(日本語が下部に続きます)

About AIRoAThe AI Robot Association (AIRoA) is launching a groundbreaking initiative: collecting one million hours of humanoid robot operation data with hundreds of robots, and leveraging it to train the world’s most powerful Vision-Language-Action (VLA) models.

What makes AIRoA unique is not only the unprecedented scale of real-world data and humanoid platforms, but also our commitment to making everything open and accessible. We are building a shared “robot data ecosystem” where datasets, trained models, and benchmarks are available to everyone. Researchers around the world will be able to evaluate their models on standardized humanoid robots through our open evaluation platform.

For researchers, this means an opportunity to:

  • Work on fundamental challenges in robotics and AI: multimodal learning, tactile-rich manipulation, sim-to-real transfer, and large-scale benchmarking.
  • Access state-of-the-art infrastructure: hundreds of humanoid robots, GPU clusters, high-fidelity simulators, and a global-scale evaluation pipeline.
  • Collaborate with leading experts across academia and industry, and publish results that will shape the next decade of robotics.
  • Contribute to an initiative that will redefine the future of embodied AI—with all results made open to the world.

Key Responsibilities

You will play a critical role in building the data backbone powering next-generation robotics foundation models:

  • Design and implement large-scale data pipelines that cover the full lifecycle of high-quality datasets for robotics foundation models—collection, processing, curation, and publishing.
  • Design, build, and maintain data schemas, storage solutions, and query interfaces to enable VLA researchers to efficiently discover, query, and consume curated datasets.
  • Collaborate closely with VLA researchers to capture evolving data requirements and continuously improve data pipelines through analysis and experimentation.
  • Design and scale distributed data-processing pipelines capable of handling petabyte-scale multimodal datasets (e.g., RGB/Depth, point clouds) with full lineage and reproducibility.
  • Define data-quality metrics and build feedback loops to continuously monitor and improve data quality.

AIRoAについてAI Robot Association(AIRoA)は、数百台のヒューマノイドロボットを用いて100万時間分のロボット操作データを収集し、それを活用して世界最高水準のVision-Language-Action(VLA)モデルを学習させるという、画期的な取り組みを進めています。

AIRoAの特徴は、実世界データとヒューマノイドプラットフォームの規模がこれまでにないものであるだけでなく、すべてをオープンかつ誰もが利用できる形にすることを目指している点にあります。私たちは、データセット、学習済みモデル、ベンチマークを誰もが利用できる共通の「ロボットデータ・エコシステム」を構築しています。また、世界中の研究者が、オープンな評価プラットフォームを通じて、標準化されたヒューマノイドロボット上で自身のモデルを評価できる環境を提供します。

研究者にとって、これは以下のような機会を意味します。

  • マルチモーダル学習、触覚情報を活用したマニピュレーション、Sim-to-Real転移、大規模ベンチマーキングなど、ロボティクスおよびAIにおける本質的な課題に取り組む。
  • 数百台のヒューマノイドロボット、GPUクラスター、高精度シミュレーター、グローバル規模の評価パイプラインといった最先端のインフラを利用する。
  • 学術界・産業界を代表する専門家と協働し、今後10年のロボティクスを形作る研究成果を発表する。
  • Embodied AIの未来を再定義する取り組みに貢献し、その成果をすべて世界にオープンにする。

主な業務内容

次世代のロボティクス基盤モデルを支えるデータ基盤の構築において、重要な役割を担っていただきます。

  • ロボティクス基盤モデル向けの高品質なデータセットについて、収集、処理、キュレーション、公開までのライフサイクル全体をカバーする大規模データパイプラインを設計・実装する。
  • VLA研究者がキュレーション済みデータセットを効率的に発見・検索・利用できるよう、データスキーマ、ストレージソリューション、クエリインターフェースを設計・構築・運用する。
  • VLA研究者と密接に連携し、変化するデータ要件を把握するとともに、分析や実験を通じてデータパイプラインを継続的に改善する。
  • RGB/Depth、ポイントクラウドなどのペタバイト規模のマルチモーダルデータセットを扱い、完全なデータリネージと再現性を確保できる分散データ処理パイプラインを設計し、スケールさせる。
  • データ品質指標を定義し、データ品質を継続的に監視・改善するためのフィードバックループを構築する。

(日本語が下部に続きます)

Required Qualifications【1. Academic & Professional】

  • Bachelor's degree in Computer Science, Engineering, or related field (or equivalent practical experience).
  • 5+ years professional experience in data engineering / data platform development.
  • Proven record of delivering production-grade, distributed data systems.

【2. Large-Scale Multimodal / Unstructured Data Processing】

  • Experience personally designing and operating pipelines that process unstructured data such as video, images, point clouds, or sensor time-series at scale (100TB+ in total, or TB+/day throughput).
  • Working understanding of storage formats (Parquet, WebDataset, MCAP/rosbag, etc.) or codecs (H.264, H.265, AV1, etc.) for such data.
  • Relevant domains include autonomous driving, robotics, video/streaming, industrial IoT, and video analytics. Equivalent experience from other domains is also welcome.

【3. ETL / Distributed Data Processing】

  • 3+ years designing and operating large-scale ETL / ELT pipelines using a distributed engine such as Spark, Flink, or Ray.
  • Experience personally building and operating pipelines with orchestration tools such as Airflow or Dagster.

【4. Collaboration & Data Democratization】

  • Worked with research, product, or analytics teams and designed pipelines and platforms around their goals.
  • Built tools, interfaces, and documentation that let users who are not data-engineering specialists (such as researchers and analysts) find and use data on their own.

Preferred Qualifications• Experience analyzing and evaluating robot data: statistics and visualization of trajectories and manipulation logs, task success/failure assessment, and metrics for dataset diversity and quality.

  • Understanding of sensor characteristics (camera, depth, IMU, force/tactile), time synchronization, and calibration; knowledge of robot learning dataset formats such as LeRobot and Open X-Embodiment (RLDS).
  • Proven optimization of workloads at 10TB+/day or petabyte scale.
  • Experience operating data-processing workloads on Kubernetes (e.g., EKS, GKE).
  • Experience building pipelines that feed ML training jobs, including dataloader/sharding optimization (WebDataset, Mosaic StreamingDataset, Lance, etc.) and dataset versioning for reproducibility.
  • Experience building dataset search and discovery, using full-text search and vector search (e.g., OpenSearch, FAISS, pgvector).
  • Experience building and operating annotation platforms (e.g., CVAT, Label Studio, Labelbox), designing human-in-the-loop workflows, and managing annotation quality.
  • Experience with large-scale telemetry or streaming ingestion from devices and fleets (e.g., Kafka, Kinesis, MQTT).
  • Experience building lakehouses (e.g., Apache Iceberg, Delta Lake, Hudi) with query engines such as Trino or Athena.

Others (linguistic qualification, etc.)【Highly appreciated】 English proficiency at business level; Japanese proficiency a plus.

必須要件【1. 学歴・職務経験】

  • コンピュータサイエンス、工学、または関連分野の学士号を有すること(または同等の実務経験)。
  • データエンジニアリング/データプラットフォーム開発における5年以上の実務経験。
  • 本番環境で利用される分散データシステムを構築・提供した実績。

【2. 大規模マルチモーダル/非構造データ処理】

  • 動画、画像、ポイントクラウド、センサー時系列などの非構造データを大規模に処理するパイプラインを、自ら設計・運用した経験(累計100TB以上、または1日あたりTB規模以上のスループット)。
  • こうしたデータで使用されるストレージフォーマット(Parquet、WebDataset、MCAP/rosbagなど)、またはCodec(H.264、H.265、AV1など)に関する実務上の理解。
  • 関連する分野として、自動運転、ロボティクス、動画/ストリーミング、Industrial IoT、動画解析などを想定しています。他分野における同等の経験も歓迎します。

【3. ETL/分散データ処理】

  • Spark、Flink、Rayなどの分散処理エンジンを用いた大規模ETL/ELTパイプラインの設計・運用経験が3年以上あること。
  • AirflowやDagsterなどのオーケストレーションツールを用いて、自らパイプラインを構築・運用した経験。

【4. コラボレーション&データの民主化】

  • 研究、プロダクト、またはアナリティクスチームと協働し、それぞれの目的に沿ってパイプラインやプラットフォームを設計した経験。
  • 研究者やアナリストなど、データエンジニアリングの専門家ではないユーザーでも、自らデータを発見・利用できるようにするためのツール、インターフェース、ドキュメントを構築した経験。

歓迎要件• ロボットデータの分析・評価経験:軌道やマニピュレーションログの統計分析・可視化、タスクの成功/失敗評価、データセットの多様性や品質を評価する指標の設計・利用。

  • センサー特性(カメラ、Depth、IMU、Force/Tactile)、時刻同期、キャリブレーションへの理解、およびLeRobotやOpen X-Embodiment(RLDS)などのロボット学習用データセットフォーマットに関する知識。
  • 1日あたり10TB以上、またはペタバイト規模のワークロードを最適化した実績。
  • Kubernetes(EKS、GKEなど)上でデータ処理ワークロードを運用した経験。
  • ML学習ジョブにデータを供給するパイプラインの構築経験。WebDataset、Mosaic StreamingDataset、Lanceなどを用いたDataLoader/Shardingの最適化や、再現性を担保するためのデータセットバージョニングを含む。
  • 全文検索やベクトル検索(OpenSearch、FAISS、pgvectorなど)を用いたデータセット検索・発見機能の構築経験。
  • Annotation Platform(CVAT、Label Studio、Labelboxなど)の構築・運用、Human-in-the-Loopワークフローの設計、Annotation Qualityの管理経験。
  • デバイスやFleetからの大規模Telemetry/Streaming Ingestionの経験(Kafka、Kinesis、MQTTなど)。
  • Apache Iceberg、Delta Lake、Hudiなどを用いたLakehouseと、TrinoやAthenaなどのQuery Engineを組み合わせたシステムの構築経験。

その他(語学要件など)【特に歓迎】ビジネスレベルの英語力。日本語力があれば尚可。

There are currently no comparable projects in the world that collect data and develop foundation models on such a large scale. As mentioned above, this is one of Japan’s leading national projects, supported by a substantial investment of 20.5 billion yen from NEDO.

This position will play a crucial role in determining the success of the project. You will have broad discretion and responsibility, and we are confident that, if successful, you will gain both a great sense of achievement and the opportunity to make a meaningful contribution to society.

Furthermore, we strongly encourage engineers to actively build their careers through this project—for example, by publishing research papers and engaging in academic activities.

●Work locationTokyo Ryutsu Center A Bldg. AW4-5/4-6, 6-1-1 Heiwajima, Ota-ku, Tokyo 143-0006, Japan

Similar roles