About the role
Design, build, and maintain scalable data pipelines and cloud-based data solutions within an AWS ecosystem. Manage data ingestion, processing, storage, and delivery to support Analytics and BI teams.
What they look for
Requirements
Requires advanced proficiency in Python and SQL with hands-on experience in Apache Spark, Kafka, and Airflow. Candidates must have strong AWS experience and a background in data modeling and ETL/ELT processes.
Full description
Role Overview
We are seeking an experienced Senior Data Engineer to design, build, and maintain scalable data pipelines and cloud-based data solutions. The role covers data ingestion, processing, storage, orchestration, and delivery to Analytics and BI teams within an AWS-based data ecosystem.
Responsibilities
- Design and maintain batch and real-time data pipelines.
- Integrate data from APIs, databases, files, and external sources using Apache Kafka and AWS services.
- Develop and orchestrate workflows with Apache Airflow, including monitoring, retry mechanisms, and alerting.
- Build and optimize distributed data processing solutions using Apache Spark, PySpark, Spark SQL, and AWS EMR.
- Manage Data Lake and Data Warehouse environments using Amazon S3 and Amazon Redshift.
- Design analytical data models, including Star and Snowflake schemas.
- Prepare and optimize datasets for Analytics and Business Intelligence use cases, particularly for Qlik.
- Develop high-quality solutions using Python and SQL, following best practices for Git, code reviews, testing, and documentation.
- Ensure data quality, security, governance, and compliance, including AWS IAM, encryption, and access control.
- Monitor production data pipelines and proactively address operational and performance issues.
Mandatory Requirements
- Proven experience as a Senior Data Engineer or in a similar role.
- Advanced proficiency in Python and SQL.
- Hands-on experience with:• Apache Spark / PySpark
- Apache Kafka
- Apache Airflow
- Strong AWS experience, particularly with:• Amazon S3
- Amazon Redshift
- AWS EMR
- Experience with Data Lakes, Data Warehouses, ETL/ELT processes, and data modeling.
- Strong understanding of batch and streaming architectures.
- Experience with performance optimization, Git, testing practices, and production support.
Technical Stack
- Python
- SQL
- Apache Spark / PySpark
- Apache Kafka
- Apache Airflow
- AWS EMR
- Amazon S3
- Amazon Redshift
- Git
- Qlik
Similar roles
-
Data Engineer
Central One Federal Credit Union Shrewsbury, Massachusetts, United States · $96K–$124K/yr
-
Data Engineer III
Playlist Brazil
-
Data Engineer - DataOps (Position located in Bengaluru, India)
KnowBe4 Bengaluru, Karnataka, India
-
Lead Data Engineer (12 Month FTC)
AND Digital London, England, United Kingdom
-
Full Stack Data Engineer
Ford Motor Company India
-
Staff Data Engineer
Sandisk Bengaluru, Karnataka, India