Lead Data Engineer
LionsBot International Pte Ltd · Singapore, Singapore
Automation Machinery Manufacturing · 51-200 employees
About the role
You will own the end-to-end data platform, including ingestion, storage, modeling, and serving for high-volume IoT telemetry. You will also design a medallion architecture and build a semantic layer to ensure a single source of truth for fleet KPIs.
What they look for
Requirements
The role requires 3+ years of experience in data engineering or backend development with heavy data exposure. Proficiency in SQL, Python, and database design is essential, along with experience in event-driven or streaming data systems.
Full description
Your data sources have wheels. LionsBot designs and builds autonomous cleaning robots that work in the real world: malls, airports, offices and industrial sites across 30+ countries. More than 5,000 of them stream telemetry to us in real time: missions, maps, locations, incidents, battery health. We move fast: small team, quick decisions, zero bureaucracy, hardware you can kick.
You'll be our first dedicated data hire, a true 0→1, greenfield ownership role. The foundations are in place: real-time telemetry streams from the fleet into a time-series store, with dashboards on top. Our fleet has now grown to the point where data deserves a full-time owner, so we're making it a first-class function. End-to-end, it's yours.
The mission: take us from "a pipeline that works" to a streaming-first data platform with a proper medallion architecture: bronze raw telemetry, silver cleaned and conformed, gold business-ready marts. On top of it all, a semantic layer where every metric has exactly one definition and everyone trusts the number.
The fun problems, all real
- Robots report cumulative lifetime odometers on every mission row. Sum the wrong column and your fleet total inflates 1,000×. Design the models that make that mistake impossible.
- A sensor glitch claims one robot cleaned 2.5 million m² in twenty minutes. Build the data quality and anomaly detection that catches it before a human ever sees it.
- Robots in basements with bad Wi-Fi send late-arriving, out-of-order data. Make the pipelines idempotent anyway.
- Real-time fleet health: which robots are sick right now, across 30+ countries and time zones?
What you will do
- Own the data platform end-to-end: ingestion, storage, modeling, serving, dashboards. Real-time event streams from the fleet land in a time-series database today. Where it goes next is your call.
- Design the medallion architecture: bronze, silver and gold layers over high-volume IoT telemetry, with clear data contracts agreed with the backend teams so quality is designed in at the source.
- Build the metrics/semantic layer: canonical, documented, version-controlled definitions for fleet KPIs: cleaning hours, area, mission success, incident rates, robot health. A genuine single source of truth.
- Run data quality & observability like production software: freshness SLAs, validation, dedup, outlier handling, anomaly alerts. Flag the weird number before leadership does.
- Design, tune and re-architect databases at scale: schemas, indexes, continuous aggregates, compression, downsampling, retention and partitioning, treating them like the production systems they are.
- Make analytics self-serve: dashboards and models for ops, product, leadership and OEM partners, plus fast, rigorous answers to the high-stakes ad-hoc questions.
- Shape the roadmap: we run lean today, so what comes next is genuinely open: OLAP, orchestration, transformation tooling, lakehouse patterns. You evaluate, make the case, and we adopt what earns its keep.
- Work AI-native: we pair humans with LLM-powered analytics agents daily. You'll design the platform so both humans and AI agents can query it safely and correctly.
What we are looking for
- 3+ years working with data in production: data engineering, analytics engineering, or backend with heavy data exposure. We hire for trajectory, not year count.
- Strong SQL, solid PostgreSQL and confident database design: schemas, indexes and data models that hold up as data grows. Time-series databases like TimescaleDB or InfluxDB are a big plus, but you'll learn them fast here.
- Comfortable with event-driven data: you've worked with streaming or message-queue systems like Kafka, or you're a data-minded backend engineer keen to go deeper on real-time.
- Solid Python for pipelines and tooling, and comfortable reading Go or Java services.
- You've shipped dashboards and metrics people actually used, whatever the BI tool.
- Fast and autonomous, like our robots: high ownership, pragmatic trade-offs, comfortable with ambiguity, ships iteratively.
- Clear communication: you translate data into decisions, not just charts.
Nice to have:
- Production streaming chops: you know your at-least-once from your exactly-once.
- IoT, robotics, or high-volume device telemetry experience.
- AWS, especially EKS, RDS and S3, with exposure to Azure or GCP.
- OLAP engines, orchestration or transformation tooling, CDC pipelines.
- Search engines like Quickwit or Elasticsearch, or graph databases like Neo4j.
- Geospatial data: our robots navigate real floors, so maps and location streams are first-class citizens.
- Experience making data platforms LLM/agent-friendly: semantic layers, governed self-serve.
- Familiarity with OpenRMF, ROS or robotics-related communication stacks
Why join:
- 0→1 ownership. The architecture, the standards, and the tooling choices are yours to shape. Eventually, so is the team.
- Grow with the function. We're hiring for trajectory: as data grows from one person into a team, you're first in line to lead it.
- Physical-world data at real scale. Thousands of robots, 30+ countries, real-time streams. Not clickstream. Not ad attribution. Robots.
- Visible impact. Small team, direct line to leadership. What you ship this week is used in decisions next week.
- AI-forward team. We already run AI-assisted analytics in production workflows. You'll multiply it, not fight it.