LionsBot International Pte Ltd

Lead Data Engineer

LionsBot International Pte Ltd · Singapore, Singapore

Automation Machinery Manufacturing · 51-200 employees

4 h ago
Mid (2-5 yrs) Full-time Singapore
Log in to apply, save this posting, or score it against your profile with AI.

About the role

You will own the end-to-end data platform, including ingestion, storage, modeling, and serving for high-volume IoT telemetry. You will also design a medallion architecture and build a semantic layer to ensure a single source of truth for fleet KPIs.

What they look for

Data Engineering SQL PostgreSQL Python Data Modeling Time-series Databases Streaming Systems Kafka Data Quality Observability Medallion Architecture IoT Cloud Infrastructure Analytics Engineering Database Design

Requirements

The role requires 3+ years of experience in data engineering or backend development with heavy data exposure. Proficiency in SQL, Python, and database design is essential, along with experience in event-driven or streaming data systems.

Full description

Your data sources have wheels. LionsBot designs and builds autonomous cleaning robots that work in the real world: malls, airports, offices and industrial sites across 30+ countries. More than 5,000 of them stream telemetry to us in real time: missions, maps, locations, incidents, battery health. We move fast: small team, quick decisions, zero bureaucracy, hardware you can kick.

You'll be our first dedicated data hire, a true 0→1, greenfield ownership role. The foundations are in place: real-time telemetry streams from the fleet into a time-series store, with dashboards on top. Our fleet has now grown to the point where data deserves a full-time owner, so we're making it a first-class function. End-to-end, it's yours.

The mission: take us from "a pipeline that works" to a streaming-first data platform with a proper medallion architecture: bronze raw telemetry, silver cleaned and conformed, gold business-ready marts. On top of it all, a semantic layer where every metric has exactly one definition and everyone trusts the number.

The fun problems, all real

  • Robots report cumulative lifetime odometers on every mission row. Sum the wrong column and your fleet total inflates 1,000×. Design the models that make that mistake impossible.
  • A sensor glitch claims one robot cleaned 2.5 million m² in twenty minutes. Build the data quality and anomaly detection that catches it before a human ever sees it.
  • Robots in basements with bad Wi-Fi send late-arriving, out-of-order data. Make the pipelines idempotent anyway.
  • Real-time fleet health: which robots are sick right now, across 30+ countries and time zones?

What you will do

  • Own the data platform end-to-end: ingestion, storage, modeling, serving, dashboards. Real-time event streams from the fleet land in a time-series database today. Where it goes next is your call.
  • Design the medallion architecture: bronze, silver and gold layers over high-volume IoT telemetry, with clear data contracts agreed with the backend teams so quality is designed in at the source.
  • Build the metrics/semantic layer: canonical, documented, version-controlled definitions for fleet KPIs: cleaning hours, area, mission success, incident rates, robot health. A genuine single source of truth.
  • Run data quality & observability like production software: freshness SLAs, validation, dedup, outlier handling, anomaly alerts. Flag the weird number before leadership does.
  • Design, tune and re-architect databases at scale: schemas, indexes, continuous aggregates, compression, downsampling, retention and partitioning, treating them like the production systems they are.
  • Make analytics self-serve: dashboards and models for ops, product, leadership and OEM partners, plus fast, rigorous answers to the high-stakes ad-hoc questions.
  • Shape the roadmap: we run lean today, so what comes next is genuinely open: OLAP, orchestration, transformation tooling, lakehouse patterns. You evaluate, make the case, and we adopt what earns its keep.
  • Work AI-native: we pair humans with LLM-powered analytics agents daily. You'll design the platform so both humans and AI agents can query it safely and correctly.

What we are looking for

  • 3+ years working with data in production: data engineering, analytics engineering, or backend with heavy data exposure. We hire for trajectory, not year count.
  • Strong SQL, solid PostgreSQL and confident database design: schemas, indexes and data models that hold up as data grows. Time-series databases like TimescaleDB or InfluxDB are a big plus, but you'll learn them fast here.
  • Comfortable with event-driven data: you've worked with streaming or message-queue systems like Kafka, or you're a data-minded backend engineer keen to go deeper on real-time.
  • Solid Python for pipelines and tooling, and comfortable reading Go or Java services.
  • You've shipped dashboards and metrics people actually used, whatever the BI tool.
  • Fast and autonomous, like our robots: high ownership, pragmatic trade-offs, comfortable with ambiguity, ships iteratively.
  • Clear communication: you translate data into decisions, not just charts.

Nice to have:

  • Production streaming chops: you know your at-least-once from your exactly-once.
  • IoT, robotics, or high-volume device telemetry experience.
  • AWS, especially EKS, RDS and S3, with exposure to Azure or GCP.
  • OLAP engines, orchestration or transformation tooling, CDC pipelines.
  • Search engines like Quickwit or Elasticsearch, or graph databases like Neo4j.
  • Geospatial data: our robots navigate real floors, so maps and location streams are first-class citizens.
  • Experience making data platforms LLM/agent-friendly: semantic layers, governed self-serve.
  • Familiarity with OpenRMF, ROS or robotics-related communication stacks

Why join:

  • 0→1 ownership. The architecture, the standards, and the tooling choices are yours to shape. Eventually, so is the team.
  • Grow with the function. We're hiring for trajectory: as data grows from one person into a team, you're first in line to lead it.
  • Physical-world data at real scale. Thousands of robots, 30+ countries, real-time streams. Not clickstream. Not ad attribution. Robots.
  • Visible impact. Small team, direct line to leadership. What you ship this week is used in decisions next week.
  • AI-forward team. We already run AI-assisted analytics in production workflows. You'll multiply it, not fight it.