Caliente Interactive: Data Engineer — Data Lakehouse Senior
CommIT Warsaw, Masovian Voivodeship, Poland
Software Development · 501-1,000 employees
About the role
You will own the end-to-end data lakehouse architecture, ensuring accurate and reconciled data for analytics, finance, and regulatory reporting. Responsibilities include managing CDC streaming ingestion, optimizing data layout for performance, and maintaining strict governance and data quality standards.
What they look for
Requirements
The role requires at least 5 years of experience in data engineering with a focus on large-scale lakehouse architectures and cloud-based data warehousing. Candidates must possess strong skills in SQL, Python, streaming technologies like Kafka, and data governance practices.
Full description
We are looking for Data Engineer in Kraków, Poland who will own the company data lake — the system of record for millions of financial events a day (bets, wallet movements, live odds) across 12M+ active users. You decide how that data lands, is stored, retained, and governed on S3 + Snowflake/Databricks, so analytics, finance, and regulators all see accurate, reconciled data with zero drift from source.
Domain: Regulated iGaming / wallet & ledger data. Audit-heavy: regulators, finance and analytics all consume the same tables. Millions of financial events per day, terabyte-plus scale.
What you'll be doing:
- Own the lakehouse architecture: bronze/silver/gold layers, Iceberg/Delta tables, schema evolution.
- Land operational data via CDC streaming (Kafka, Debezium), handling late and duplicate events.
- Design data layout for speed and cost: partitioning, compaction, file sizing, query performance on Trino/Athena/Snowflake.
- Own retention and archival: storage tiering, regulatory retention, immutability, GDPR deletion.
- Guarantee correctness: freshness SLAs, drift detection, reconciliation against the source wallet and ledger systems.
- Own governance: catalog and lineage, row/column access control, PII masking, encryption, audit trails.
- Monitor ingestion health, data anomalies, and cloud storage/compute spend.
Requirements
Must-have:
- 5+ years in data engineering, with real ownership of a large-scale data lake or lakehouse.
- Lakehouse architecture — bronze/silver/gold layering, an open table format (Iceberg, Delta, or Hudi), schema evolution.
- Data layout & query optimization at TB+ scale — partitioning, compaction, file sizing, query performance on Trino/Athena/Snowflake.
- Cloud lakehouse/DWH in production — Snowflake, Databricks, or BigQuery.
- CDC & streaming ingestion — Kafka + Debezium or equivalent; late, duplicate and out-of-order events.
- Strong SQL and data modeling — enough relational grounding to reason about the OLTP systems you capture from. Critical for financial ledgers.
- Correctness — freshness SLAs, drift detection, reconciliation against source wallet/ledger systems.
- Governance — catalogs, lineage, row/column access control, PII masking, retention, GDPR deletion.
- Cloud object storage — S3 or GCS, plus storage tiering and archival.
- Python and an orchestrator — Airflow or Dagster, as tools.
Location & work model:
Kraków, Poland. Hybrid — 2 days per week from the office.
Similar roles
-
Data Engineer - fully remote within Europe (m/f/d)
JobLeads Careers Hamburg, Germany
-
Senior Data Engineer
Publicis Groupe Holdings B.V Bengaluru, Karnataka, India
-
Lead Data Engineer
Publicis Groupe Holdings B.V Bengaluru, Karnataka, India
-
Senior Data Engineer
Talan Málaga, Andalusia, Spain
-
Data Engineer
MerQube Inc Bengaluru, Karnataka, India
-
GCP Data Engineer
EXL Gurugram, Haryana, India