Weekday AI

Data Engineer

Weekday AI Mumbai, Maharashtra, India

Technology, Information and Internet · 11-50 employees

6 h ago
data-engineer Mid (2-5 yrs) Full-time India
Create a free account to apply — email only, no card. You can also save this posting or score it against your profile with AI.

About the role

Design, develop, and maintain scalable end-to-end data pipelines for high-volume batch and real-time data processing. Collaborate with cross-functional teams to build reliable data lakes and warehouses while ensuring data quality and pipeline performance.

What they look for

Python Apache Spark SQL AWS Kafka Data Engineering ETL Data Modeling Airflow Amazon Redshift Amazon Glue Amazon S3 PySpark Streaming Architectures Data Pipelines Linux

Requirements

Requires 3+ years of professional experience in data engineering with strong expertise in Python, Apache Spark, and the AWS data ecosystem. Candidates should possess a solid understanding of SQL, data modeling, and modern software engineering practices.

Full description

𝗧𝗵𝗶𝘀 𝗿𝗼𝗹𝗲 𝗶𝘀 𝗳𝗼𝗿 𝗼𝗻𝗲 𝗼𝗳 𝘁𝗵𝗲 𝗪𝗲𝗲𝗸𝗱𝗮𝘆'𝘀 𝗰𝗹𝗶𝗲𝗻𝘁𝘀

𝗦𝗮𝗹𝗮𝗿𝘆 𝗿𝗮𝗻𝗴𝗲: 𝗥𝘀 𝟭𝟴𝟬𝟬𝟬𝟬𝟬 - 𝗥𝘀 𝟮𝟯𝟬𝟬𝟬𝟬𝟬 (𝗶𝗲 𝗜𝗡𝗥 𝟭𝟴-𝟮𝟯 𝗟𝗣𝗔)

Experience: 3+ yrs

Location: Mumbai, Maharashtra, India

Job Type: Full-time

We are looking for an experienced Data Engineer to design, develop, and maintain scalable data platforms and pipelines using AWS, Apache Spark, Python, SQL, Kafka, and modern data engineering frameworks.

The role will focus on building reliable data solutions for high-volume batch and real-time workloads, including data ingestion, transformation, processing, orchestration, storage, and delivery. The ideal candidate will have strong hands-on experience with AWS data services and distributed data processing, along with a solid understanding of data modelling, streaming architectures, data quality, and pipeline optimisation.

KEY RESPONSIBILITIES

  • Design, develop, and maintain end-to-end data pipelines for high-volume data ingestion, transformation, processing, and delivery.
  • Build scalable Spark-based ETL/ELT workflows for both batch and real-time data processing.
  • Develop data ingestion solutions using Kafka, Amazon Kinesis, and other streaming technologies.
  • Build and manage data lakes, warehouses, and lakehouse solutions using AWS S3, Glue, Redshift, Athena, and EMR.
  • Develop efficient data models using dimensional modelling, star schemas, partitioning, and other data engineering practices.
  • Implement data quality checks, validation rules, monitoring, and error-handling mechanisms.
  • Develop automated workflows using Airflow, MWAA, AWS Step Functions, or similar orchestration tools.
  • Collaborate with Data Analysts and Data Scientists to deliver clean, structured, and analytics-ready datasets.
  • Optimise data pipelines for performance, scalability, reliability, and AWS cost efficiency.
  • Integrate data from multiple internal and external systems while maintaining data consistency and reliability.
  • Develop Python-based automation and data-processing solutions.
  • Monitor production pipelines, troubleshoot failures, and perform root-cause analysis.
  • Follow modern software engineering practices including Git, CI/CD, testing, documentation, and code reviews.
  • Contribute to data platform architecture, engineering standards, and continuous improvement initiatives.
  • Support data governance, cataloguing, lineage, and metadata management practices where required.
  • Work with Linux/Unix environments and efficiently process large datasets.

WHAT MAKES YOU A GREAT FIT

  • 3+ years of professional experience in Data Engineering, Big Data, Analytics Engineering, or a related field.
  • Strong hands-on programming experience with Python for data processing, automation, and pipeline development.
  • Strong expertise in Apache Spark, particularly PySpark and/or Spark SQL.
  • Deep working knowledge of the AWS data ecosystem, including S3, Glue, Redshift, Athena, EMR, Kinesis, Lambda, and IAM.
  • Hands-on experience with Kafka, Kinesis, Flink, or similar real-time streaming technologies.
  • Strong command of SQL and experience with data modelling, dimensional modelling, partitioning, and large-scale data processing.
  • Experience with Airflow, MWAA, Step Functions, or comparable workflow orchestration tools.
  • Strong understanding of batch and real-time data processing architectures.
  • Experience working with high-volume datasets and distributed data processing environments.
  • Familiarity with Git, CI/CD, testing, and modern software development practices.
  • Comfortable working in Linux/Unix environments.
  • Strong troubleshooting, analytical, and problem-solving skills.
  • Experience with Delta Lake, Apache Iceberg, Hudi, or other lakehouse technologies is an advantage.
  • Knowledge of data governance, cataloguing, metadata, and lineage tools such as Glue Data Catalog, DataHub, or Amundsen is a plus.
  • Familiarity with Docker, ECS, or EKS and containerised deployments is desirable.
  • Basic understanding of AI/ML data requirements and workflows is an advantage.
  • Bachelor's degree in Computer Science, Information Technology, Engineering, or a related technical disciplineis preferred.

Similar roles