Principal Data Engineer
Bukuwarung Hyderabad, Telangana, India
Technology, Information and Internet · 201-500 employees
About the role
You will architect and build a unified, scalable data platform to serve as the single source of truth for payments, credit, and fraud operations. This role involves leading technical standards, mentoring the team, and delivering real-time streaming pipelines to support company-wide decision-making.
What they look for
Requirements
Candidates must have 8+ years of experience in data engineering with a strong background in building scalable data platforms, ideally within the fintech sector. Proficiency in Python, SQL, streaming architectures, and cloud data infrastructure is essential for this hands-on technical leadership role.
Full description
About BukuWarung
BukuWarung is building the digital and financial infrastructure for micro and small businesses across Southeast Asia. We serve millions of MSMEs through payments (BukuPay), credit (BukuModal), and financial tools — helping underbanked entrepreneurs grow faster and more securely.
The next phase of BukuWarung's growth is data-native: real-time decisioning, smarter underwriting, fraud prevention, and deeply personalized products — built for the 100M+ MSMEs across Southeast Asia who remain underserved by traditional financial institutions. None of that ships without a unified data platform underneath it.
Why This Role Matters
Today, BukuWarung's most valuable signals — payments transactions, device activation and usage, merchant behavior, and a growing lending book — live in fragmented systems across distribution channels, hardware operations, product, and finance. Fraud, credit, and growth teams each stitch together their own view of the truth. That slows decisions and caps how good our models can be.
As Principal Data Platform Engineer, you are the hands-on technical owner of the unified data platform that fixes this. You will architect and build the streaming and batch backbone that consolidates every channel and operation into a single, governed, trustworthy source of truth — the foundation that credit underwriting, real-time fraud detection, and company-wide analytics all depend on. You will:
- Architect and build a unified data platform that consolidates all distribution channels and operations into a single internal system
- Stand up streaming + batch pipelines that deliver low-latency data for fraud, credit, and real-time merchant decisioning
- Make clean, governed, self-serve data a company-wide resource that every business unit can trust
- Set the technical standard for data engineering as the team scales — through architecture, patterns, and mentorship rather than headcount alone
Key Responsibilities
Unified Data Platform & Architecture
- Own the end-to-end architecture of BukuWarung's unified data platform — ingestion, storage, transformation, serving, and access — as a single source of truth across payments, credit, fraud, hardware ops, and GTM
- Consolidate fragmented, channel-specific data sources (EDC/POS, QRIS soundboxes, digital payment flows, lending, field sales, logistics) into a coherent, well-modeled warehouse and lakehouse
- Design the semantic and data-modeling layer so metrics are defined once and consumed consistently across every business unit
- Make deliberate build-vs-buy decisions across the modern data stack (warehouse/lakehouse, orchestration, transformation, catalog, BI) balancing cost, control, and speed
Streaming & Real-Time Data Pipelines
- Build near-real-time streaming pipelines (Kafka, Flink, or Spark Streaming) that power low-latency fraud detection and credit decisioning
- Design and operate reliable batch pipelines and orchestration (e.g. Airflow/Dagster) for reporting, model training, and reconciliation workloads
- Deliver a feature store and serving layer so ML teams can move from a model in a notebook to a production decision with consistent online/offline features
- Engineer for scale and cost efficiency as transaction volume grows, with clear SLAs on latency, freshness, and throughput
Data Quality, Governance & Reliability
- Build data quality, testing, and observability into every pipeline — particularly for credit and fraud data, where errors carry direct financial consequences
- Establish lineage, cataloging, and documentation so data is discoverable, trustworthy, and auditable end to end
- Implement governance, access controls, and PII handling suited to Bank Indonesia regulations, OJK compliance requirements, and cross-border fintech data standards
- Define and enforce SLAs for data freshness, completeness, and reliability, with alerting and clear ownership when things break
Enablement & Self-Serve
- Build reusable frameworks, tooling, and paved paths that let Ops, Finance, and GTM self-serve on routine data questions without engineering bottlenecks
- Deliver a governed BI and metrics layer that drives action — dashboards teams actually decide from, not reports that go unread
- Partner with analysts, ML engineers, and risk scientists to make the platform genuinely productive for the people who build on it
- Evaluate and integrate no-code / low-code tooling where it accelerates business teams safely
Technical Leadership
- Set data engineering standards — architecture reviews, coding patterns, CI/CD for data, and infrastructure-as-code — that the team can scale on
- Mentor engineers and analysts, raising the technical bar as the data organization grows from its current core team
- Be a credible technical partner to product, engineering, and business leaders, translating platform decisions into their impact on capital allocation, risk, and speed
- Contribute hands-on to the hardest parts of the system — this is a builder role, not a purely supervisory one
Requirements
Must-Have
- 8+ years in data engineering, with a track record of architecting and building data platforms at scale — ideally including fintech (payments, lending, or fraud/risk)
- Deep, hands-on expertise designing unified data platforms: warehouse/lakehouse, ingestion, transformation (e.g. dbt/Spark), and orchestration (e.g. Airflow/Dagster)
- Strong experience with streaming architectures (Kafka, Flink, or Spark Streaming) for real-time, low-latency decisioning
- Expert in SQL and Python; comfortable with distributed data processing and cloud data infrastructure (AWS/GCP)
- Proven ownership of data quality, governance, lineage, and reliability in a production, high-stakes environment
- Experience enabling ML and analytics teams — feature stores, serving layers, and self-serve tooling
- Ability to lead through architecture and mentorship, and to communicate technical trade-offs clearly to non-technical stakeholders
Nice-to-Have
- Experience in Indonesia or Southeast Asia fintech, with familiarity with Bank Indonesia and OJK regulatory frameworks
- Exposure to credit underwriting or real-time fraud data systems as a platform consumer
- Prior work with IoT / device telemetry data (relevant to EDC/POS and QRIS soundbox fleets)
- Familiarity with field sales / agent-distribution data and logistics (3PL tracking, kit dispatch, serial mapping)
- Experience with data catalog, observability (e.g. Monte Carlo, Great Expectations), and infrastructure-as-code (Terraform)
Key Impact Areas — First 18 Months
- Unified Data Platform — Ship the foundation of a single source of truth that consolidates all channels and operations, replacing fragmented, per-team data views
- Real-Time Backbone — Deliver production streaming pipelines that make low-latency data available for fraud and credit decisioning
- Trustworthy Data — Establish quality, lineage, and governance so credit, fraud, and finance teams can rely on the platform without manual reconciliation
- Self-Serve Enablement — Give Ops, Finance, and GTM governed, reusable tools to answer routine questions independently
- Scalable Foundation — Set the architecture, standards, and patterns that let the data organization grow without accumulating technical debt
Similar roles
-
Data Engineer
Pixelogic Media Partners, LLC Cape Town, Western Cape, South Africa
-
Data Engineer
Serko Ltd Bengaluru, Karnataka, India
-
Lead Data Engineer
Eltropy Inc. India
-
Senior Data Engineer- Managed services
Telefonica Tech pune, Maharashtra, India
-
Senior Data Engineer
WalletConnect Berlin, Germany
-
Senior Data Engineer
Station United Kingdom · £65K–£70K/yr