T

Sr. Data Engineer - Flink

Techsa Cairo, Cairo, Egypt

Software Development · 11-50 employees

Aug 10
Remote data-engineer Senior (5-10 yrs) Full-time
Create a free account to apply — email only, no card. You can also save this posting or score it against your profile with AI.

About the role

Design and operate high-volume, low-latency real-time data systems using Apache Flink as the core engine. Build and maintain high-scale streaming data pipelines on self-managed on-premise infrastructure.

What they look for

Apache Flink Java Scala Python Data Engineering Kubernetes Kafka Streams SQL Distributed Systems ETL Data Modeling Data Warehousing Apache Iceberg Docker Airflow Stream Processing

Requirements

Requires 7+ years of experience in data engineering with deep hands-on production expertise in Apache Flink and its state management. Candidates must be proficient in Java, Scala, or Python and have experience operating data workloads on self-managed infrastructure like Kubernetes or YARN.

Full description

Building high-scale, low-latency streaming data pipelines deployed on infrastructure we run ourselves (on-prem), not managed cloud services. You will design and operate high-volume real-time data systems end to end, with Apache Flink as the core stream-processing engine.

Requirements

  • 7+ years

of experience in data engineering and software development

  • Ability to

write high-quality code in Java/Scala, Python, or equivalent languages

  • Deep,

hands-on production experience with Apache Flink — DataStream API and Table API / Flink SQL (core requirement)

  • Demonstrated

experience with Flink state management: keyed state, state backends (e.g., RocksDB), large state sizes, and state TTL

  • Hands-on

experience with checkpointing, savepoints, and fault tolerance — exactly-once vs. at-least-once semantics, recovery, and savepoint-based job upgrades

  • Strong

grasp of event-time processing: watermarking, windowing strategies, allowed lateness, and late-data handling

  • Experience

diagnosing and resolving backpressure — parallelism, operator chaining, and network buffer tuning

  • Experience

operating Flink on self-managed infrastructure (Kubernetes or YARN) — application vs. session mode, high availability, and rolling upgrades

  • Practical

experience with stream processing (Kafka Streams or equivalent) and messaging systems for high-volume workloads, including exactly-once sinks and schema registry usage

  • Practical

experience with distributed query engines (e.g., Trino/Presto or similar)

  • Practical

experience with ETL / data integration tools, commercial or open-source (e.g., Datastage, Informatica, Apache NiFi, or similar)

  • Practical

experience with SQL-based transformation frameworks (e.g., dbt or others)

  • Strong SQL

skills and understanding of data modeling and data warehousing for analytical workloads

  • Hands-on

experience with real-time / low-latency analytical stores (columnar or OLAP engines, e.g., Apache Pinot/ClickHouse or similar)

  • Practical

experience with big-data platforms and distributions (e.g., Cloudera, Hadoop ecosystem, or similar)

  • Practical

experience containerizing and operating data workloads (Docker; Kubernetes a plus)

  • Experience

with workflow orchestration tools (e.g., Airflow or similar)

  • Familiarity

with data lake table formats (e.g., Apache Iceberg or similar), including streaming ingestion, compaction, and small-file management

  • Familiarity

with data governance / cataloging tools (e.g., DataHub or similar)

  • Familiarity

with lakehouse management systems (e.g., Apache Amoro or similar)

  • Familiarity

using AI tools for development and debugging (Claude, Cursor, Codex)

Similar roles