About the role
Design and operate high-volume analytical data systems end to end using Apache Spark. Build and maintain high-scale batch and near-real-time data pipelines on self-managed infrastructure.
What they look for
Requirements
Requires 7+ years of experience in data engineering and software development with deep production expertise in Apache Spark. Candidates must be proficient in Java, Scala, or Python and have experience operating Spark on self-managed clusters.
Full description
Building high-scale batch and near-real-time data pipelines deployed on infrastructure we run ourselves (on-prem), not managed cloud services. You will design and operate high-volume analytical data systems end to end, with Apache Spark as the core processing engine for both batch and streaming workloads.
Requirements
- 7+ years
of experience in data engineering and software development
- Ability to
write high-quality code in Java/Scala, Python, or equivalent languages
- Deep,
hands-on production experience with Apache Spark — batch and Spark Structured Streaming (core requirement)
- Demonstrated
Spark performance tuning: partitioning, caching and persistence, broadcast joins, shuffle reduction, data-skew handling, and Adaptive Query Execution
- Experience
operating Spark on self-managed clusters (YARN, Kubernetes, or standalone) — executor sizing, resource allocation, and multi-tenant workloads
- Practical
experience with Kafka (or equivalent messaging systems) as a Spark source and sink for high-volume workloads, including offset and checkpoint management
- Practical
experience with distributed query engines (e.g., Trino/Presto or similar)
- Practical
experience with ETL / data integration tools, commercial or open-source (e.g., Datastage, Informatica, Apache NiFi, or similar)
- Practical
experience with SQL-based transformation frameworks (e.g., dbt or others)
- Strong SQL
skills and understanding of data modeling and data warehousing for analytical workloads
- Hands-on
experience with real-time / low-latency analytical stores (columnar or OLAP engines, e.g., Apache Pinot/ClickHouse or similar)
- Practical
experience with big-data platforms and distributions (e.g., Cloudera, Hadoop ecosystem, Databricks, or similar)
- Practical
experience containerizing and operating data workloads (Docker; Kubernetes a plus)
- Experience
with workflow orchestration tools (e.g., Airflow or similar)
- Familiarity
with data lake table formats (e.g., Apache Iceberg, Delta Lake, or similar), including schema evolution and compaction
- Familiarity
with data governance / cataloging tools (e.g., DataHub or similar)
- Familiarity
with lakehouse management systems (e.g., Apache Amoro or similar)
- Familiarity
using AI tools for development and debugging (Claude, Cursor, Codex)
Similar roles
-
Senior Data Engineer, Vice President
State Street Boston, Massachusetts, United States · $110K–$208K/yr
-
AI Data Engineer II (Business Data Analyst II)
UKG Bengaluru, Karnataka, India
-
Databricks Data Engineer
Elevate Government Solutions Washington, District of Columbia, United States
-
Senior Data Engineer
Iris Software Noida, Uttar Pradesh, India
-
Data Engineer II, GenAI
Travelers Hartford, Connecticut, United States · $126K–$209K/yr
-
Data Engineer - BODS
Weekday AI Bengaluru, Karnataka, India